PreFlight sits in front of your agents' actions, blocks the ones that break your rules, and issues a proof the other side checks without a login, your policy, or your data.
Pre-action blocking, deterministic enforcement, verification by anyone, privacy-preserving, and cloud and hardware neutral, graded from each vendor's public documentation as of September 2026.
| Pre-action blocking |
Deterministic enforcement |
Verifiable by anyoneno vendor or hardware trust |
Privacy- preserving |
Cloud and hardware neutral |
|
|---|---|---|---|---|---|
| ICME PreFlightpre-action · formal · ZK receipt | |||||
| Hardware attestationEQTY Lab Verifiable Runtime · NVIDIA Sentry | |||||
| Signed gateway evidenceTraefik Sovereign Trust Plane | |||||
| NVIDIA OpenShellsandbox · formal policy prover | |||||
| MCP gatewaysLasso · Zenity · Pillar | |||||
| LLM-as-judge guardrailsDatadog AI Guard · NeMo Guardrails | |||||
| Runtime screeningLakera · F5 AI Guardrails · Fiddler | |||||
| AI governance platformsCredo AI · Monitaur · Holistic AI | |||||
| Audit-log toolsobservability and logging |
Partial means the capability is statistical rather than formal, covers authorization but not the action's content, or depends on trusting a signer, log operator, or chip vendor. Privacy-preserving means verification reveals nothing about the policy or the data. Cloud and hardware neutral means it runs on any cloud with no specific chip. Categories are graded on what their products typically provide, from public documentation as of September 30, 2026. Hardware attestation relies on Intel or NVIDIA keys, which WireTap, Battering RAM, and TEE.fail broke in October 2025. Traefik's transparency log proves a decision record was not altered, not that the check was correct; it covers authorization, not action content, and is graded on its announced design, listed as Early Access. OpenShell checks actions before they run but produces no evidence a third party can verify.
AWS AgentCore Policy + Automated Reasoning · Google AP2 · Visa Trusted Agent Protocol · Mastercard Verifiable IntentThese authorize the action. PreFlight proves the action was checked against your rules before it ran.
Each wants evidence of an automated action that a third party can check without access to your systems.
Federal judges now require a signed certificate that AI-assisted filings were independently verified. The Eastern District of Texas Standing Order MC-11 is one of dozens; sanctions in 2026 run to $110,204 in a single case.
E.D. Tex. Standing Order MC-11; Couvrette v. Wisnovsky (D. Or.)When a regulator stops believing an operator's own monitoring, it installs its own verifier. One financial regulator ordered a payments company to appoint an external auditor reporting directly to it, within 180 days, at the company's expense.
AML/CTF external audit order, January 2026Corporate clients are writing AI clauses into outside counsel guidelines: disclose generative AI use, name the tool in the time entry, have a lawyer check the output. The Hartford, UBS, Zscaler, and Microsoft publish their own.
The Hartford panel counsel guidelines; UBS billing guidelines (Feb 2026)Carriers writing affirmative AI coverage moved underwriting from whether a firm has an AI policy to whether its controls operate. ISO's January 2026 exclusions ended silent AI coverage, and professional liability policies do not pay fines or sanctions.
Counterpart (Nov 2025); Beazley (Apr 2026); ALPS and Berkley Select (Aug 2026); ISO CG 40 47, CG 40 48, CG 35 08Visa Intelligent Commerce and Mastercard Agent Pay bind the cardholder, the agent, and the mandate into a token and enforce the mandate before authorization. The acquirer verifies the agent; the operator's word is not the check.
Visa and OpenAI announcement; Mastercard Agent Pay for MachinesChecked September 2026. These are the demands as published; no court or regulator has named a cryptographic receipt as satisfying them, and this page does not claim one has.
Each action is checked before it runs. The check returns a verdict and a zero-knowledge receipt that can exist only if that check ran on that action against that policy. Altering it means breaking the cryptography.
There is no signature or chip to trust. A signed log proves the record was not changed, but the verifier must trust the signer's keys and read the record. Hardware attestation asks the verifier to trust the chip vendor. A receipt is checked against public math and reveals only the verdict.
The policy_hash identifies the exact ruleset checked. valid: true means the proof passed.
A court, examiner, card network, insurer, or client verifies a receipt at a public endpoint, with no account and no access to your systems. The verdict returns in under a second, before the action runs; the receipt finishes proving in the background in tens of seconds.
Each receipt commits to the policy version in force and the SAT or UNSAT result. It proves the rule was checked against what the action stated, not that those statements are true.
The verifier sees the verdict, the policy hash, and the proof, and nothing else. Probing with many actions returns only verdicts, never the rules. The proof stays about the same size however complex the policy.
All three check the action before it runs. The difference is who decides that the check happens: your code, the agent, or the platform. A tool an agent can call is a tool an agent can skip, so only the hook makes the check impossible to avoid.
A plain-English action description and a policy ID go in. No documents.
The fastest way to adopt PreFlight inside an agent that already speaks MCP.
When the platform invokes the check as a mandatory step, the action cannot proceed without a verdict. The MCP server makes PreFlight easy to adopt; the hook makes it impossible to bypass.
One call checks an action and returns a verdict plus a proof. A second call verifies that proof, and it takes no API key and no account.
// check an action
POST /v1/checkIt # X-API-Key
{ "action": "Send envelope: MSA, $240k, no CFO approval",
"policy_id": "vendor_v2" }
// under a second
{ "result": "UNSAT", "proof_id": "90484dac...",
"policy_hash": "3f2a91cc..." }
// anyone verifies it
POST /v1/verifyProof # no API key
{ "proof_id": "90484dac..." }
{ "valid": true, "verify_ms": 408,
"policy_hash": "3f2a91cc..." }
Same rule, same action, same answer. An API you can write tests against, unlike guardrails that drift between runs.
One verifiable artifact per action, bound to the exact rule version, that attaches to your own record. Anyone verifies it with no API key and no account.
Versioned policy IDs for production, staging, and draft, with scenario and test tools.
A plain-English description of the action goes out. No documents, no client data, and the receipt reveals no rule text.
How PreFlight fits your stack, what it costs, and the answers to the questions a security or risk reviewer sends before a pilot.
Not by anything written in a document or a prompt. The solver has no context window and reads no instructions; it evaluates the proposed action against compiled logic and returns SAT or UNSAT. There is nothing in it to inject. Skipping the check is a separate question, which is why placement matters. As a pre-execution hook, the platform runs the check and the agent never chooses whether to be checked.
The LLM does one job: translating your policy into formal logic, once, at compile time. That output is inspected, checked for contradictions, and battle-tested against generated scenarios before anything goes live. At decision time a solver decides, and a solver cannot be talked around or fed a trick instruction. The probabilistic component never makes the allow-or-block call.
A locked workflow makes the steps repeatable, then hands your auditor a trace of what ran. That trace is a log. To trust it, the auditor has to trust that it was not edited, that nothing was omitted, and that your system produced it honestly.
PreFlight returns a proof instead of a trace. A third party verifies it without trusting you, and without seeing your policy or your data. Repeatable is good. Provable, and provable without disclosure, is the part a locked workflow does not give you.
Keep them; your buyers require them. But certifications verify the operator's controls, at the organization level, on an annual cycle. They say nothing about whether a specific agent action complied with a specific rule today. A certificate vouches for the company. The receipt vouches for the agent action.
There is also an evidentiary hierarchy here. The audit profession's own evidence standards rank independent external evidence above anything produced by the entity under review. A receipt verified by outside math sits in that top tier. A control report produced inside the system it vouches for does not.
It depends on how you deploy, and in two of the three models the question does not arise.
Pilot. A pilot uses your published policy and test actions. No client data is involved, so no vendor security review is needed to start.
Production in your own cloud account. The checkpoint runs inside your AWS account. Actions are checked and proofs are generated within your environment, and only the proof and verdict metadata reach ICME. ICME verifies your receipts and never holds client data, so SOC 2 is not applicable to us. Your cloud provider's certifications cover the infrastructure and your own controls cover the data.
ICME-hosted. SOC 2 becomes relevant only if you prefer that we host. In that case we run a single-tenant environment with private connectivity and encryption keys you hold and can revoke, and we meet your vendor security requirements before any client data moves.
Policies are built on facts present in the action itself: amount, recipient, destination. Text slipped into an email the agent read (a prompt injection) cannot change what the proposed action actually says. When PreFlight gates at the point of execution, the checked action is the executed action. The agent cannot move $10,000 while describing a $100 transfer, because the description is the transfer.
Correct, and we lead with it. SAT means the action satisfied your policy. It does not mean the policy was wise, or that the facts the action asserted are true. The receipt proves the rule fired against what the action claimed, which is accountability, not a safety guarantee.
Then they still get the enforcement. The blocking decision comes from the solver consensus, not from the proof layer. So the worst case for a skeptic is "the guardrail still blocked the bad action, and the receipt is evidence I am choosing not to rely on yet."
The receipt adds verifiability on top of the safety. It does not carry the safety; that is the solver's job. A reviewer can adopt the evidence on their own timeline without the protection ever depending on it.
For the attorney or compliance owner who uses it day to day, there is nothing technical about it: they write plain-English if-then rules the way they would brief a new hire, and that is the entire authoring experience. Compilation and battle-testing run behind it. There is no separate interface to learn, so rules can be authored wherever your team manages policy today.
Two timings. Rules an attorney writes in their own words take minutes to compile, because the solver checks the whole set for contradictions and edge cases before anything governs a live action. Rules that range over fixed, known fields auto-generate in about two seconds. The first buys a version an attorney signed off on; the second is why per-customer setup at platform scale does not wait on a compile.
PreFlight is an API. By default there is no infrastructure to stand up and no server to run. For production in a regulated firm, the checkpoint can instead run inside your own cloud account so no client data leaves it. Either way it is cloud-neutral, and the receipt verifies on any stack, not just the one that produced it. A third party checks a decision against a single public endpoint, with no API key and no account.
No. Every check runs two independent verifications and reconciles them before returning a result: a custom LLM check that evaluates the action against the compiled SMT model of your policy, and a cloud reasoning engine that independently evaluates the same action against the same policy. A SAT result requires both to agree. If either returns UNSAT, the action is blocked.
So a single extraction error or a single reasoning gap cannot produce a false clearance, and no path, including the LLM, can talk the system out of a proven violation. The solver underneath is Z3, which came out of Microsoft Research and has been a standard formal verification engine for close to twenty years; the translation method is adopted from AWS Automated Reasoning checks (arXiv:2511.09008). Policies are also checked for internal contradictions at compile time, before they are saved.
The system surfaces ambiguity rather than quietly resolving it. At compile time, battle testing generates scenarios that expose a vague or wrong rule first. Before production, you can pull every extracted variable, its type, and its rule, and tighten a loose phrase. At runtime, an ambiguous term returns an uncertain result instead of a clean verdict, and the check resolves conservatively.
The runtime uncertain result is a backstop, not a workflow. If it appears in production, treat it as a signal to return to the policy, define the term, and re-run battle testing. A well-tested policy produces clean SAT and UNSAT verdicts.
Fully, as the customer. You can retrieve the original policy text, the compiled SMT-LIB, and the parsed rules at no charge, and export the SMT-LIB for your own tooling or an independent reviewer. Every check also returns the extracted variable bindings and the per-path results alongside the proof.
Outsiders get the other kind of transparency: the zero-knowledge receipt lets anyone confirm the decision was made correctly without seeing your rules and without trusting us. Your clients' reviewers verify the check ran; your policy stays private.
An audit trail is a record the operator keeps about its own conduct; when a reviewer asks how they know the log is honest, the only answer is to trust the operator. The receipt is a piece of math an outside party checks, and if anyone tampered with the result, the check fails. The verifier code is public and the protocols are peer reviewed.
Signing the log does not change that. Signed or transparency-logged records (the approach behind products like Traefik's Sovereign Trust Plane) prove the record was not altered after the fact, and some of these products also block before the action runs. But a signature proves who vouched for the decision, not that the check was computed correctly, and the verifier has to read the record to check it. A receipt proves the check ran correctly on that action against that policy, and the verifier learns the verdict and the policy hash without the record, the policy, or the data.
Products like EQTY Lab run the guardrail inside secure hardware and publish an attestation that it ran. That is a real improvement over a log. But the verifier has to trust the chip vendor's attestation keys, and in October 2025 researchers extracted or bypassed them in Intel, AMD, and NVIDIA confidential computing (WireTap, Battering RAM, TEE.fail). A PreFlight receipt has no hardware root to steal. The two approaches can be combined: hardware can vouch that the action description came from the real system, and the receipt proves the check on it was correct.
JOLT is a16z Crypto's open-source zkVM, and its README says it is in alpha and not yet audited. We disclose that plainly. The audit question is narrower than it looks, though. In a zero-knowledge system the property that protects you is verifier correctness, not whole-codebase perfection. A correctly implemented verifier rejects a cheating prover, so a bug or bad actor on the proving side produces a rejected proof, not a false one.
The verifier is the small, inspectable, open-source component, and it implements peer-reviewed protocols (JOLT, IACR ePrint 2023/1217, with our zkML layer in JOLT Atlas, arXiv:2602.17452). You do not need to inspect it personally; the assurance comes from the fact that anyone can. A dishonest verifier would be caught by the first cryptographer who looked, and cryptographers do look, which is the same reason banks trust TLS without reading its source. Compare that to a proprietary log format, where nobody outside the vendor can check anything.
Two scoping points: enforcement does not depend on JOLT at all, since the fail-closed blocking comes from the solver consensus above, and the upstream verification work is public and trackable at github.com/a16z/jolt rather than promised, including its Z3 verifier component. We hold ourselves to the same disclosure discipline we recommend to buyers: verify the small thing that matters, disclose the maturity of the rest. If a16z publishes audit results, we will surface them.
Developers can be checking agent actions in minutes. Enterprise teams can start with a scoped pilot on the policies your risk team is asking about.
Compile a policy, check an agent action, and get a verdict and a receipt. Public proof verification needs no API key.
Read the docsBring the policy stuck in review. We will count the rules that are checkable facts versus genuine judgment calls and scope a pilot from there.
Book a call