AI action verification

Verify what AI agents actually did.

Start with refunds, credits, or cancellations. Veridatum checks agent/vendor claims against your systems of record, then shows what held, failed, or needs review.

Read-only · export-first One workflow No production writes No live customer replies
Claim → evidence

An agent's claim is not a business outcome.

Your system of record is wherever the truth actually lives: Stripe, your order database, your billing system. Veridatum checks each reported action against that record and shows whether it held, meaning the change actually took effect and stayed true, not just that the agent reported it done.

What the agent reports
What your system of record shows

Vendor dashboards report their own status, counts, and completions. Veridatum's job is the independent comparison: does that reported status match the systems you already trust?

Where reports break

“Done” is what the agent reported. Not what the record shows.

01 · payment state

Refund issued?

An agent says it refunded the customer. Did the payment record post, at the right amount, without reversal or duplication?

02 · account state

Cancellation completed?

An agent says it cancelled an account. Did billing actually stop, and did the entitlement change?

03 · ticket state

Ticket resolved?

An agent says the ticket was resolved. Did the backend state change, or did a human quietly reopen and fix it later?

Verified action ledger

Every claim, its proof, and a verdict.

The first sprint focuses on one state-changing support workflow, usually refunds, credits, cancellations, order changes, or entitlement changes. The output is a buyer-readable ledger of what held, what failed, and what should be accepted, disputed, escalated, or excluded.

SAMPLE DELIVERABLE PREVIEW · support actions
We only ever read your data, via an export you share or read-only access
Claimed actionSystem-of-record proofHold statusVeridatum statusDecision
Refund issued Stripe refund posted, £42.50 Held 7 days, not reversed Verified Accept
Cancellation completed No cancellation in billing record Still billing at 24h Disputed Escalate
Duplicate refund requested Already refunded once No change Verified restraint Accept
Order address changed Shipping address changed to new recipient before fulfilment Held Verified Monitor
Credit applied Wallet balance updated, £15.00 Held 3 days Verified Accept
Ticket closed Ticket marked resolved Reopened in 2h Disputed Dispute
Design-partner program

Bring one workflow. Leave with a verified ledger.

Veridatum is early. We're opening a small design-partner cohort for teams with one AI-agent workflow to verify. You bring a redacted sample and a short debrief. You get the kind of buyer-readable ledger previewed above, direct input into the method, and design-partner terms.

How a sprint runs

  1. 1Describe a workflowNo data or credentials to start: just one workflow where an agent changes backend state, and the decision attached to it.
  2. 2Sample & source-of-truth mappingA redacted export or scoped read-only sample, mapped to the records that actually hold the truth, with clean joins and messy ones noted.
  3. 3Verified ledger & recommendationAn evidence-backed ledger of what held, what didn't, and what to accept, dispute, escalate, or exclude.
See what a sprint involves
1–2 WEEKS · ONE WORKFLOW · SMALL COHORT

The Verified Action Sprint.

Start from a redacted export or a scoped read-only sample, as small as a few dozen actions. Leave with a verified action ledger showing what the agent claimed, what your system of record shows, whether it held, and what should be accepted, disputed, escalated, or excluded.

What you bring
  • One workflow where an AI agent changes backend state: refunds, cancellations, credits, order or account changes, entitlement changes, or ticket resolutions.
  • A redacted sample of agent/vendor claims and the matching system-of-record records. No production access or credentials.
  • A named decision the result would inform, and roughly the exposure attached to it.
What you get
  • A verified action ledger and an evidence-backed register of disputed, failed, and corrected actions.
  • Source-of-truth mapping notes and sample evidence packets you can inspect.
  • A recommendation (trust, gate, dispute, monitor, or stop), plus design-partner terms and direct input into the method.
Public method demos

On public benchmarks, Veridatum caught what dashboards can miss.

These examples demonstrate the verification method. They do not claim production customer results, buyer demand, or real-world error rates. The next proof point has to be buyer-authorized data.

Reported vs. verified · τ-bench retail
If every reported action is trusted100%
Actually held in the system of record74%

26% didn't match. Across 456 GPT-4.1 retail simulations, 118 disagreed with the database on final state. That's the gap a vendor dashboard wouldn't surface on its own.

Veridatum analysis · public MIT-licensed benchmark · not real customer data
Public method demo · τ-bench retail

Agent promised a refund. The record showed the wrong order was cancelled. Final database check failed.

Status: disputed
METHOD DEMO · public benchmarks
Public benchmark data, not real customer data. Each row cites its source task
Claimed / attempted actionBackend (benchmark) proofHold statusVeridatum statusSource
Cancel order after price drop Order + item status = cancelled; refund £249 to card Held Verified STATE-Bench · 40
Cancel customer's order, refund promised Wrong order cancelled; db_match = false Mismatch Disputed τ-bench · 51
Second courtesy refund requested Already refunded; no mutation performed No change Verified restraint STATE-Bench · 134
τ-bench retail · Sierra Research

A quarter of runs didn't match the business state.

Across 456 GPT-4.1 retail simulations, 118 failed the final-state check: the agent's transcript and the database disagreed. In one, the agent cancelled the wrong order and told the customer a refund was on its way; in another it acted past an explicit "don't cancel anything else" instruction, inflating the refund. 104 of 114 tasks involve state-changing actions.

Microsoft STATE-Bench · enterprise support

The refund math, and the refusals that should happen.

Across 150 support tasks with 952 explicit state assertions, the same method checks more than "resolved": refund amounts after clawbacks, waived fees, and multi-leg compound actions. It also verifies restraint: an agent correctly refusing a duplicate refund is a control worth confirming, not just successful actions.

Public synthetic benchmarks demonstrate the verification method and the shape of the artifact. They do not, by themselves, prove buyer demand, real-world error rates, or that production data is this clean. Buyer deployments run on buyer-authorized systems of record.

Who it's for

For teams letting AI agents change business state.

When an agent can move money, close accounts, or change entitlements, someone has to confirm it actually happened. Veridatum gives buyers source-cited evidence for contested AI-agent outcomes: what held, what failed, what should be paid, and what needs review.

Support / Ops
WorkflowRefunds, cancellations, ticket closures
EvidenceTicket trail plus system-of-record state
DecisionFix, gate, or monitor silent failures
Finance
WorkflowRefunds, credits, true-ups
EvidencePayment record plus agent or vendor claim
DecisionAccept, dispute, or exclude
Procurement
WorkflowVendor-reported agent outcomes
EvidenceBuyer-owned ledger with source citations
DecisionRenew, challenge, or renegotiate
A good moment to start

You do not need a platform rollout. You need one scoped verification sprint.

An AI vendor renewal, true-up, or review coming up A new agent action type about to go live An incident that needs reconstructing A workflow that moves real money
Trust & data handling

What we touch, and what we don't.

Verification means looking at sensitive records, so the handling matters as much as the method. Here is how a sprint is scoped.

0No dataStart with a workflow description and the decision it would inform.
1Redacted exportShare agent or vendor claims plus matching records when needed.
2Scoped read-onlyOne workflow, limited fields, no writes, no live customer replies.
3Optional repeatOnly after the first sprint proves useful and the scope stays clear.
Not needed for a first sprint
  • Production writes
  • Live customer replies
  • Broad warehouse access
  • Secrets or credentials
  • Training data
Used for a first sprint
  • Workflow description
  • Redacted claim export
  • Matching system-of-record sample
  • Source-of-truth rule
  • Deleted by default
FAQ

Questions buyers ask.

Independence
Isn't this just our vendor's dashboard, or something we could build ourselves?

Neither is independent, and independence is the point. A vendor's dashboard reports what the vendor counted; your own team can reconcile it for internal use, but in a dispute the vendor contests your numbers first, and risk or compliance sign-offs are built to distrust a check run by the same side it's checking. A neutral, buyer-owned ledger both sides can reference is harder to wave away. The hard part isn't the reconciliation. It's knowing which record holds the truth, and what "actually held" means when two systems disagree, per action type and across messy systems.

Scope
Is this invoice reconciliation or AI-spend management?

No. Veridatum checks whether a business action actually became true in your systems, not whether usage arithmetic adds up. It can inform what's payable or what to take into a renewal, but it is not invoice reconciliation or spend management.

Access
Do you need access to our production systems?

No. A sprint runs on a redacted export or scoped read-only access. The first step needs no data at all, just a description of the workflow.

Messy records
What if our logs and records are messy?

Expected. Part of the sprint is mapping agent/vendor claims to the right source-of-truth records and noting where the join is clean and where it isn't. You get those mapping notes either way.

Cost
How much does a design-partner sprint cost?

Design-partner terms are deliberately light, and scoped to the workflow and the decision attached to it. We'll quote it on a short call once we've confirmed the workflow is a fit.

Starting point
What do you need from us to start?

One workflow where an agent changes backend state, a redacted sample of the agent's claims and the matching records, and a named decision the result would inform. That's it to begin.

Small design-partner cohort

Start with one workflow.

We're opening a small design-partner cohort for teams with one AI-agent workflow to verify. No data or credentials to start: just a description, and we'll tell you what evidence would verify it.

or email hello@veridatum.io