Break it on purpose. Run the same violating request through the real enforcement path. Prove that your test sees the prohibited effect when one controlled defect is present, then restore the control and prove the effect disappears.

Do this only in an isolated non-production environment, using synthetic or canary resources and a contained effect sink. Here, “real enforcement path” means the actual enforcement components and integration topology—not live customer data, real recipients or an irreversible external effect. Production mutation testing needs a separately approved method that proves no disclosure, transfer, message or other irreversible effect can escape containment.

If you cannot represent the violation, reach the enforcement point or observe the effect, the result is UNEVALUABLE. A green result would be fiction.

This method gives you evidence about one declared control, path and failure mode. It does not prove that the control is universally effective. One killed mutation does not create statistical power out of thin air.

A green test can be doing nothing

Imagine a policy that prevents a member of one research project from reading another project’s documents. Your test sends a cross-project request. The API returns 403. Green.

What did you prove?

Perhaps the policy denied the request. Perhaps the request failed in a gateway before it reached the policy. Perhaps the test account never existed. Perhaps the endpoint returned an error while a downstream export still committed. The colour alone cannot tell you.

NIST defines a control test as exercising a mechanism under specified conditions and comparing the actual state with expected behaviour.1 Both halves matter. The mechanism has to run. The expected state has to be observable.

The test earns trust when it can detect the control’s absence.

Start by declaring the population

“We tested access control” is not a population. It is a label.

A useful declaration names:

  • the subjects and roles in scope;
  • the resources and relevant attributes;
  • the actions being attempted;
  • the entry paths and protocols;
  • the enforcement points;
  • the policy and attribute versions;
  • the downstream effects you can observe.

NIST separates assessment coverage from assessment depth.1 Coverage describes the number and types of assessment objects. Depth describes how rigorously they are examined. A deep test of one endpoint can still leave every alternate path untouched.

That distinction keeps one success from quietly becoming a system-wide claim.

For this example, the declared population is deliberately small: one document-read endpoint, its direct API path, two project identities, one deployed policy version and the resource-side enforcement point. The guide makes no claim about other endpoints.

Express one relevant violation

The policy is plain:

A member may read a project document only when the member’s project identifier equals the project identifier of that document.

The request is equally plain:

subject.id        = "researcher-B"
subject.project   = "project-B"
action            = "read"
resource.id       = "document-A-17"
resource.project  = "project-A"

The expected authorization result is DENY. The expected protected effect is also explicit: no document-A-17 content is returned and no downstream export event is committed.

This sounds obvious. It is easy to get wrong.

Attribute-based access control evaluates facts about a subject, an object, an operation and the surrounding environment.2 The predicates have to stay attached to the same identified object. A loose implementation might discover that one document has ID document-A-17 and some other document belongs to project-A, then treat both predicates as satisfied. Every value was present. The relationship was lost.

XACML names an attribute by category, attribute identifier and optional issuer, and its selectors can address resource content by path.3 Those mechanisms provide context for evaluation. They do not prove that two values came from the same resource instance.

The code that assembles the authorization request has to keep those values attached to one resource record. The test fixture should show where resource.id and resource.project came from and prove they describe that same record.

Watch the enforcement point and the effect

A unit test of a policy function is useful. It says little about whether production traffic reaches that function.

The watched test sends the cross-project request through the real document endpoint. It records:

  • the request identity and resource identity;
  • the policy version and relevant attributes;
  • the evaluator’s result;
  • the enforcement point’s action;
  • the returned document state;
  • the downstream export state.

OWASP’s authorization-testing guidance uses direct requests under unauthorized identities and checks the actual response and exposed data.4 NIST’s access-enforcement assessment procedures make the same architectural distinction: examining a policy document and testing the mechanism that implements it are different activities.1 AI Governance Without Runtime Is Performative explains why documentation and execution remain different artifacts.

The effect check matters because denial can arrive late. NIST’s Zero Trust Architecture separates policy decision, policy administration and policy enforcement.5 An enforcement point may terminate access. That does not roll back a disclosure, write or message that already completed. Runtime Is the Moment of Consequence in AI covers the wider safety case for watching this boundary.

For this test, 403 is necessary. The stronger evidence check is that no document bytes appeared and no export event committed.

Make the test earn a red result

Now introduce one controlled defect:

Keep the mutation inside that isolated environment. The permit defect must make a synthetic prohibited effect visible to the test without exposing real data or reaching a real recipient.

original rule: subject.project == resource.project
mutated rule:  permit

Run the identical request through the identical endpoint with the identical observations.

The test must go red because the prohibited effect becomes visible. If it stays green, the test did not observe the defect. The likely causes are worth finding: the mutant never loaded, another layer denied first, the request missed the path, or the evidence check watched the wrong state.

Mutation testing calls a mutant “killed” when at least one test distinguishes its behaviour from the original program.6 A surviving mutant may expose a weak test. It may also be equivalent, meaning the change creates no observable behavioural difference under the defined semantics.

That is why the mutant has to be controlled and plausibly behaviour-changing. Randomly deleting code can create noise rather than evidence.

One killed mutant proves one useful thing: this test was sensitive to this seeded behaviour under these recorded conditions. Mutation adequacy is measured over a declared mutant population.6 Statistical power needs a population, sampling design, sample size and uncertainty treatment. One red bar supplies none of those.

Keep the claim small enough to survive contact with the evidence.

Restore the control and repeat the same case

Restore the equality predicate. Run the same cross-project request again.

The test should now show:

evaluation         = deny
enforcement        = blocked
document content   = absent
export event       = absent

The request, observation points and expected effect stay fixed. Only the controlled defect changes. This same-condition comparison is the heart of the red/green evidence.

Changing the request after restoring the control weakens the result. You would no longer know whether the repair or the new fixture produced green.

Preserve the unknown state

Remove resource.project from the request context.

What should happen?

XACML distinguishes Permit, Deny, Indeterminate and NotApplicable.3 A missing mandatory attribute can produce Indeterminate. A deny-biased enforcement point can still block every result except Permit.

Those are two separate facts:

policy evaluation  = indeterminate: missing resource.project
external effect    = denied

Flattening both into DENY throws away the reason. Treating the case as a clean pass is worse. The control did not evaluate the intended rule.

The assurance harness also needs its own UNEVALUABLE state. Use it when:

  • the relevant violating request cannot be represented;
  • the mutant may be equivalent or did not load;
  • the real enforcement path was not exercised;
  • the protected effect cannot be observed;
  • asynchronous work remains unresolved inside the observation window.

Harness UNEVALUABLE and runtime Indeterminate are different concepts. One describes insufficient assurance evidence. The other describes a policy evaluator that could not return a normal authorization decision.

Both prevent an unknown from being painted green.

Fail-safe still needs proof

Saltzer and Schroeder described two principles that remain useful here: fail-safe defaults grant access only when permission is established, and complete mediation checks every access to every object.7

They are design principles. A client error does not prove either one.

The test must show what happened at the enforcement point and effect boundary. It should also name the untested paths: caches, alternate protocols, recovery jobs, maintenance interfaces, asynchronous consumers and stale attribute stores.

Consequential systems accumulate paths. A defensible assurance claim needs an inventory of the paths it covers.

The evidence record

Another engineer should be able to reconstruct the result from a compact record:

control population
  subjects, resources, actions, paths, policy version, enforcement points

violating request
  exact subject, resource, action, attributes, expected result

observation
  evaluator result, enforcement action, returned state, downstream effect

mutation
  exact seeded defect, proof it loaded, expected behavioural change

red run
  unchanged request, observed prohibited effect, failure reason

green run
  restored control, unchanged request, denied effect

unknown handling
  indeterminate and unevaluable conditions, observation window

limits
  untested paths, equivalent mutants, stale state, races, irreversible effects

That record does not certify the system. It gives the next reviewer something real to challenge.

What remains unproven

This method leaves hard problems intact.

The path inventory can be incomplete. The expected outcome can be wrong. A mutant can be equivalent or unrepresentative. Caches and races can change behaviour. A denied request can arrive after an external effect. A policy can be correct against stale attributes. A green case can coexist with an untested bypass.

Formal completeness needs a defined model and state space. In a distributed system, one identifier must tie related events together, the test needs a defined period to wait, and it must check whether completed effects were undone or compensated. The standards cited here do not supply one check that works for every downstream effect.

So the defensible result is precise:

For the declared control instance, request, path and policy version, the test detected the seeded permit defect and verified denial after restoration. All unobservable conditions remained unevaluable.

That is smaller than “this control can never fail.” It is evidence an engineer can use.

If you want to inspect one public runtime-policy project after the general method, explore the 5D runtime-policy project. Its existence does not validate this guide or prove any production deployment.

The reader test is simple: can you name the population, the violation, the seeded defect, the watched effect and the condition that returns UNEVALUABLE? If one is missing, the assurance chain is incomplete.


References

Footnotes

  1. NIST Joint Task Force, SP 800-53A Revision 5: Assessing Security and Privacy Controls in Information Systems and Organizations, January 2022, §§2.4–2.4.2, §3.2.3.2, AC-03(04), AC-03(13), glossary “coverage,” and Appendix C. Government assessment standard. https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-53Ar5.pdf. Verified 2026-10-01. Scope: test conditions, expected behaviour, coverage, depth and access-enforcement assessment. Limitation: no universal product-specific evidence check or proof of mathematical completeness. ↩ ↩2 ↩3

  2. Vincent Hu et al., NIST, SP 800-162: Guide to Attribute Based Access Control Definition and Considerations, January 2014 with updates through 2019, executive summary, §2 and §3.2. Government guidance. https://nvlpubs.nist.gov/nistpubs/specialpublications/NIST.SP.800-162.pdf. Verified 2026-10-01. Scope: subject, object, operation and attribute model. Limitation: no guarantee of correct correlation in a particular policy language or adapter. ↩

  3. OASIS, XACML Version 3.0 Plus Errata 01, 12 July 2017, glossary, §§2.5–2.6, §§5.29–5.30, §§7.2.1–7.2.3, §§7.3.3–7.3.5 and §7.19. Formal standard. https://docs.oasis-open.org/xacml/3.0/errata01/os/xacml-3.0-core-spec-errata01-os-complete.pdf. Verified 2026-10-01. Scope: named attributes, selectors, decision states, missing attributes and enforcement bias. Limitation: Indeterminate is not automatically fail-closed; implementations must select and test their behaviour. ↩ ↩2

  4. OWASP Foundation, Web Security Testing Guide: Testing for Bypassing Authorization Schema, WSTG-ATHZ-02, current page. Community technical guidance. https://wstg.owasp.org/latest/4-Web_Application_Security_Testing/05-Authorization/02-Bypassing_Authorization_Schema/. Verified 2026-10-01. Scope: direct unauthorized, horizontal and vertical access tests. Limitation: response similarity can produce false positives. ↩

  5. Scott Rose et al., NIST, SP 800-207: Zero Trust Architecture, August 2020, §3.2. Government guidance. https://nvlpubs.nist.gov/nistpubs/specialpublications/NIST.SP.800-207.pdf. Verified 2026-10-01. Scope: policy decision, administration and enforcement roles. Limitation: no promise of rollback or compensation for completed effects. ↩

  6. Yue Jia and Mark Harman, An Analysis and Survey of the Development of Mutation Testing, IEEE Transactions on Software Engineering 37(5), 2011, pp. 649–651 and equivalent-mutant discussion including p. 657. Primary peer-reviewed survey. https://crest.cs.ucl.ac.uk/fileadmin/crest/sebasepaper/JiaH10.pdf. Verified 2026-10-01. Scope: killed, surviving and equivalent mutants and mutation adequacy. Limitation: evidence depends on selected operators and observable behaviour; equivalent-mutant detection is generally undecidable. ↩ ↩2

  7. Jerome H. Saltzer and Michael D. Schroeder, The Protection of Information in Computer Systems, Proceedings of the IEEE 63(9), 1975, §III.A. Primary peer-reviewed research. https://www.cs.virginia.edu/~evans/cs551/saltzer/. Verified 2026-10-01. Scope: fail-safe defaults and complete mediation. Limitation: design principles do not prove a particular implementation. ↩