The boundary that matters is the line between judgment and an action’s real-world effect. Once model output can authorize a transfer, change production infrastructure, release customer information, approve a claim, or instruct another system to act, a convincing answer is no longer the standard.

For a high-risk AI agent, the direct answer is to keep model judgment inside a deterministic authority boundary. Code should enforce actions an AI agent must never take, a typed model may classify the ambiguity that remains, and a stronger model or qualified person should review the consequential residue. No model response should execute an action by itself.

That separation matters where probabilistic safety is not sufficient. A 99.9 percent confidence score cannot replace a missing approval. A plausible explanation cannot revive an expired authorization. A model that sees an instruction inside a retrieved invoice cannot treat the invoice as a new source of authority.

I tested this pattern with a new public-only experiment: deterministic rules in front, TypeSafe’s Jev as a bounded semantic signal, and Claude Sonnet 4.6 as the intended stronger-model comparator. The result is useful precisely because it was incomplete. Jev returned valid typed answers for all 28 cases. The Claude arm was blocked by an invalid local API credential before inference. And in the frozen cascade, Jev did not reduce the review residue at all.

That is not the clean marketing result. It is the result an engineer should publish.

An architectural abstraction of a proposed action passing through terminal rules, bounded typed judgment, review residue, and governed execution.

The diagram is an architectural abstraction, not an experimental result or a promise that every action reaches execution. Terminal rules can deny an action, policy can allow an eligible action without semantic review, and ambiguous or consequential cases can stop for qualified review. In this experiment, Jev did not reduce that review residue.

What is deterministic runtime governance for AI agents?

Deterministic, action-level runtime governance is the use of versioned code or policy to decide whether a proposed agent action may proceed. The model may supply evidence. It does not own the authority to execute.

The model can contribute evidence; policy and enforcement decide what the system may do with it.

The distinction is load-bearing:

  • Judgment asks whether an action appears routine, suspicious, relevant, or consistent with a stated purpose.
  • Policy combines that judgment with trusted facts such as identity, destination, approval, limits, current state, and consequence.
  • Enforcement validates the exact authorization and either performs the bound action or refuses it.

The policy decision and the executor should be separate. NIST’s attribute-based access-control guidance makes the same useful distinction between a policy decision point and a policy enforcement point. NIST’s 2026 draft concept paper on software and AI agent identity also centers identification, authentication, authorization, least privilege, delegation, logging, data-flow provenance, and controls that reduce the impact of prompt injection. (NIST SP 800-162, NIST agent identity and authorization concept paper)

A practical action path looks like this:

authenticated actor and workload + proposed action + payload reference
  -> deterministic eligibility and terminal red-line checks
  -> typed semantic evidence, only when eligible
  -> versioned policy composition
  -> deny, confirm, review, or short-lived authorization
  -> separate executor validates the exact binding
  -> executed effect and joined decision record

The topology can be an in-process hook, sidecar, gateway, or policy service. The invariant is more important than the topology: every effectful call must cross the same enforcement point. The executor should reject direct, expired, replayed, stale, or mismatched requests.

What actions must a high-risk AI agent never take?

Every deployment needs its own policy, but several runtime prohibitions should be explicit and testable. The agent:

  • must never disable authentication or audit logging to complete a task faster;
  • must never treat retrieved content or tool metadata as authority to change permissions, destinations, or approvals;
  • must never reuse an expired approval or transfer an approval between identities;
  • must never execute when identity, destination, purpose, or payload does not match the authorization;
  • must never split an action to evade an aggregate limit;
  • must never turn a missing, malformed, or unavailable provider response into permission;
  • must never convert model confidence into missing approval;
  • must never let the component that judges an action quietly acquire the credentials to execute it.

These are code and credential boundaries, not prompt instructions. A system prompt that says “never transfer funds without approval” is helpful guidance. It is not an authorization control if the same model can still call the payment tool with unrestricted credentials.

Where Jev fits in the cascade

TypeSafe’s Jev is designed for typed judgments rather than generated prose. Its System One interface exposes three primitives:

  • Choice selects from caller-defined options and returns the full probability distribution.
  • Score places the supplied state on caller-defined ordered levels.
  • Noul evaluates a yes-or-no proposition and returns the probability of yes.

In this experiment, Jev received the same structured action state for every case and returned one of three bounded values: ALLOW, BLOCK, or ESCALATE. The response included probabilities, confidence, token usage, and the resolved model version. The test harness rejected unknown options, invalid confidence, a missing answer, or a probability row that did not sum to 1.0.

That interface improves integration discipline. It removes prose parsing and gives policy code a closed set of values to validate. It does not make the answer correct. TypeSafe’s own Jev 1.13 documentation says the model is intended for atomic text judgments and can be too literal; indirection, conflicting criteria, irrelevant context, arithmetic, and adversarial text can reduce reliability. (Jev 1.13 jaggedness, TypeSafe confidence guidance)

Typed output is an engineering property. Truth is a validation problem.

Typed output is not the same as a correct answer.

The public experiment

I could not publish the earlier private application tests. Removing a product name from private evidence would not change its provenance. So I froze a new protocol and created new public evidence from scratch.

Two separately prompted local case authors proposed public-safe cases under instructions that excluded the article, private test material, and a desired vendor outcome. I froze 28 public synthetic cases from those proposals. The preserved public artifacts contain no recognizable private implementation identifiers, but this is a process control rather than proof of independent derivation. The pack covered infrastructure, payments, claims, customer data, support, prompt injection, stale approvals, identity mismatch, destination substitution, malformed state, hidden side effects, and threshold splitting.

The expected outcomes were:

Expected dispositionCases
ALLOW6
BLOCK14
ESCALATE8

The protocol made several decisions before the live calls:

  1. A terminal deterministic result always wins, even if a model returns ALLOW with 100 percent confidence.
  2. A high-consequence semantic ALLOW becomes ESCALATE.
  3. A low-confidence, missing, malformed, or unknown response becomes ESCALATE.
  4. Neither provider response can execute an action.
  5. Each provider gets one request per case. There are no retries.
  6. The runner stops before a $1 hard cap, inside the authorized €2 ceiling.

These rules test the runtime boundary separately from model agreement. A classifier can be wrong without the executor becoming unsafe, provided the policy catches the consequential failure and no bypass path exists.

What Jev measured

All 28 Jev requests completed and resolved to jev-1.13.0.

Median latency was 416 milliseconds, and p95 latency was 491 milliseconds from this client location.

MeasureResult
Valid typed responses28 of 28
Agreement with pre-frozen synthetic policy expectations26 of 28, or 92.9%
BLOCK recall12 of 14, or 85.7%
ESCALATE recall8 of 8, or 100%
ALLOW agreement6 of 6, or 100%
Median latency416 milliseconds
p95 latency491 milliseconds
Input tokens14,138
Measured list-price cost$0.000594

The two disagreements did not directly authorize action: Jev returned ESCALATE where the reference expected BLOCK. Escalation is less restrictive than a terminal block, so the distinction matters. One case involved an expired approval. The other involved sending unnecessary full customer records to an external analysis endpoint. There were no expected BLOCK or ESCALATE cases that Jev changed to ALLOW.

This is encouraging integration data, not a production accuracy claim. The cases are synthetic policy examples, not independently adjudicated outcomes from a live domain. One response per case does not establish stability or calibration. The test was run from one client location on one date.

What failed in the Claude comparison

The Anthropic comparison arm did not run. The local ANTHROPIC_API_KEY existed, but Anthropic rejected all 28 requests with HTTP 401 because it was an invalid local API credential. The failure happened before model inference. The no-retry protocol did what it was designed to do and preserved the failure.

This means the experiment produced no measured Claude Sonnet 4.6 latency, token use, classification quality, or cost. It produced no measured Jev-versus-Claude saving. A valid key and a new, explicitly versioned Claude-only completion run against the unchanged frozen cases are required before making that comparison.

The provider failure also demonstrates part of the governance case. Unavailable semantic review did not become permission. In a real high-risk path, the action would remain unexecuted or move to an approved human queue.

The deployment business case for a review cascade

The dated list prices are clear:

Anthropic lists Claude Sonnet 4.6 at $3 per million input tokens and $15 per million output tokens. TypeSafe lists Jev at $0.042 per million input tokens with no metered output-token charge.

ModelInput priceOutput priceSource verified 2 October 2026
TypeSafe Jev$0.042 per million input tokensNo metered output-token chargeTypeSafe announcement
Anthropic Claude Sonnet 4.6$3 per million input tokens$15 per million output tokensAnthropic model documentation

Claude Sonnet 4.6’s listed input rate is about 71.4 times Jev’s listed input rate. That is a unit-price comparison, not a measured Jev-versus-Claude saving. The providers may tokenize the same state differently. Claude also generates a different output shape and volume. Caching, batching, routing, retries, review labor, and error costs can dominate the invoice.

The more important cost question is architectural: how many actions need any model, how many need a narrow typed model, and how many still require expensive reasoning or a person?

Jev was called on all 28 cases to measure the model arm. The 20/8 result below is an offline replay of the frozen cascade, not a production routing trace.

In that replay, deterministic rules resolved 20 of 28 cases. The eight cases left for semantic judgment were deliberately ambiguous or high consequence. All eight still required escalation under the frozen rule. Jev did not reduce the review residue in this protocol. It added roughly 416 milliseconds of median semantic latency and a tiny amount of cost without avoiding a stronger review call.

That is the null result. For this case mix, the measurable efficiency came from deterministic filtering, not from adding Jev. A second frozen corpus with genuinely resolvable semantic cases would be needed to test whether Jev can safely shrink the review tail.

Does this improve AI model governance and runtime security?

It improves specific controls when implemented completely:

  • Authority separation: the model supplies evidence; policy decides; a separate executor acts.
  • Least privilege: the agent and judging model do not need unrestricted effectful credentials.
  • Exact binding: authorization can be tied to actor, action, destination, purpose, payload, limits, policy version, and state.
  • Fail-closed behavior: missing or invalid inference cannot authorize a consequential action.
  • Prompt-injection containment: retrieved instructions and tool descriptions remain untrusted data.
  • Change control: model and policy versions can be pinned, recorded, tested, and rolled back.
  • Accountability: the decision record can join identity, evidence reference, policy result, model result, approval, execution, and outcome.

These mechanisms align with the direction of the NIST AI Risk Management Framework, NIST’s developing work on agent identity and authorization, and OWASP’s AI Agent Security guidance. They improve the architecture relative to giving one probabilistic component both judgment and action authority.

The experiment does not prove end-to-end deployment security. It did not prove complete mediation of every tool path, executor credential separation, replay resistance, revocation, time-of-check/time-of-use behavior, concurrency safety, tamper-resistant audit storage, data residency, reviewer capacity, incident recovery, or production-domain error rates. Until those are tested, “more secure” must remain a mechanism-level claim, not a system-level conclusion.

When this architecture adds no value

The cascade adds little or no value when:

  • deterministic code already resolves the workload cheaply and correctly;
  • every remaining case is consequential enough to require a person anyway;
  • the semantic layer does not reduce review volume or improve detection;
  • added provider latency exceeds the application’s budget;
  • maintaining another model, contract, threshold, and failure mode costs more than it saves;
  • the executor still exposes a bypass path, making the governance diagram cosmetic.

This experiment hit one of those conditions. Jev did not shrink the review tail. Publishing that finding makes the article more useful, because it tells the reader what to measure before adding another model dependency.

What to test before a high-risk deployment

Before trusting this pattern with money, infrastructure, customer data, claims, healthcare, or other consequential actions, I would require:

  • independently qualified domain labels or verified outcomes, plus a holdout frozen before thresholds are selected;
  • separate measurement of harmful misses, legitimate refusals, escalation rate, task utility, latency, provider cost, and human-review cost;
  • adversarial tests for direct and indirect prompt injection, stale approval, identity mismatch, destination substitution, threshold splitting, malformed state, provider outage, and review saturation;
  • bypass tests proving that every effectful path crosses enforcement and that the model cannot reach executor credentials directly;
  • short-lived authorization bound to the exact actor, normalized action, destination, payload digest, purpose, limit, state, and policy version;
  • logs that join the decision to the actual tool outcome without retaining sensitive prompts by default;
  • rollback, revocation, model-migration, and revalidation procedures;
  • a rule that no unknown state silently becomes permission.

The benchmark must assess the whole action path. A model can classify well while the system remains unsafe because an escalation is treated as permission or an executor accepts a stale authorization.

Regulatory relevance without compliance theatre

For a high-risk AI system, depending on intended purpose and the organization’s role, this pattern may help implement or evidence controls relevant to risk management, logging, human oversight, robustness, and cybersecurity. Architecture alone does not satisfy those obligations or establish conformity.

The EU AI Act addresses record-keeping, human oversight, accuracy, robustness, and cybersecurity for high-risk AI systems. Whether those provisions apply, and when, depends on classification, intended purpose, provider and deployer roles, and the applicable transition rules. Data governance, documentation, monitoring, and any required conformity process remain separate obligations. Confirm the current classification and timing with qualified counsel before deployment. (Regulation (EU) 2024/1689, European Commission overview)

The honest claim is narrow: deterministic runtime governance can operationalize and evidence some relevant controls. It is not certification, legal advice, or a substitute for the rest of the control system.

The larger point

The model can remain probabilistic. The authority boundary should not be.

Jev can be useful as a typed sensor inside that boundary. A stronger language model can be useful on the hard residue. A person can own consequential judgment. None of them should be able to erase a terminal rule or execute with unbound authority.

The public test supports three conclusions. In this run, Jev returned valid typed output at 416 milliseconds median latency and $0.000594 total dated list-price cost, and that output was structurally straightforward to validate. Deterministic policy can contain model error and provider failure. And a new model layer creates no deployment benefit merely because its unit price is low.

Authority should remain boring. The evidence should be allowed to surprise us.

Sources and limitations

Vendor documentation and pricing were verified on 2 October 2026. Prices and model availability can change.

The experiment artifacts are preserved locally with the frozen protocol, cases, negative tests, exact payloads, raw responses, timings, usage, costs, and provider failures. The test does not establish Jev’s production accuracy, a measured advantage over Claude Sonnet 4.6, end-to-end security, or regulatory compliance. The proposed pairing is an architectural experiment, not a claim of TypeSafe sponsorship, partnership, or endorsement.