Refund issued?
An agent says it refunded the customer. Did the payment record post, at the right amount, without reversal or duplication?
Start with refunds, credits, or cancellations. Veridatum checks agent/vendor claims against your systems of record, then shows what held, failed, or needs review.
Your system of record is wherever the truth actually lives: Stripe, your order database, your billing system. Veridatum checks each reported action against that record and shows whether it held, meaning the change actually took effect and stayed true, not just that the agent reported it done.
Vendor dashboards report their own status, counts, and completions. Veridatum's job is the independent comparison: does that reported status match the systems you already trust?
An agent says it refunded the customer. Did the payment record post, at the right amount, without reversal or duplication?
An agent says it cancelled an account. Did billing actually stop, and did the entitlement change?
An agent says the ticket was resolved. Did the backend state change, or did a human quietly reopen and fix it later?
The first sprint focuses on one state-changing support workflow, usually refunds, credits, cancellations, order changes, or entitlement changes. The output is a buyer-readable ledger of what held, what failed, and what should be accepted, disputed, escalated, or excluded.
| Claimed action | System-of-record proof | Hold status | Veridatum status | Decision |
|---|---|---|---|---|
| Refund issued | Stripe refund posted, £42.50 | Held 7 days, not reversed | Verified | Accept |
| Cancellation completed | No cancellation in billing record | Still billing at 24h | Disputed | Escalate |
| Duplicate refund requested | Already refunded once | No change | Verified restraint | Accept |
| Order address changed | Shipping address changed to new recipient before fulfilment | Held | Verified | Monitor |
| Credit applied | Wallet balance updated, £15.00 | Held 3 days | Verified | Accept |
| Ticket closed | Ticket marked resolved | Reopened in 2h | Disputed | Dispute |
Veridatum is early. We're opening a small design-partner cohort for teams with one AI-agent workflow to verify. You bring a redacted sample and a short debrief. You get the kind of buyer-readable ledger previewed above, direct input into the method, and design-partner terms.
How a sprint runs
Start from a redacted export or a scoped read-only sample, as small as a few dozen actions. Leave with a verified action ledger showing what the agent claimed, what your system of record shows, whether it held, and what should be accepted, disputed, escalated, or excluded.
These examples demonstrate the verification method. They do not claim production customer results, buyer demand, or real-world error rates. The next proof point has to be buyer-authorized data.
26% didn't match. Across 456 GPT-4.1 retail simulations, 118 disagreed with the database on final state. That's the gap a vendor dashboard wouldn't surface on its own.
Agent promised a refund. The record showed the wrong order was cancelled. Final database check failed.
Status: disputed| Claimed / attempted action | Backend (benchmark) proof | Hold status | Veridatum status | Source |
|---|---|---|---|---|
| Cancel order after price drop | Order + item status = cancelled; refund £249 to card | Held | Verified | STATE-Bench · 40 |
| Cancel customer's order, refund promised | Wrong order cancelled; db_match = false | Mismatch | Disputed | τ-bench · 51 |
| Second courtesy refund requested | Already refunded; no mutation performed | No change | Verified restraint | STATE-Bench · 134 |
Across 456 GPT-4.1 retail simulations, 118 failed the final-state check: the agent's transcript and the database disagreed. In one, the agent cancelled the wrong order and told the customer a refund was on its way; in another it acted past an explicit "don't cancel anything else" instruction, inflating the refund. 104 of 114 tasks involve state-changing actions.
Across 150 support tasks with 952 explicit state assertions, the same method checks more than "resolved": refund amounts after clawbacks, waived fees, and multi-leg compound actions. It also verifies restraint: an agent correctly refusing a duplicate refund is a control worth confirming, not just successful actions.
Public synthetic benchmarks demonstrate the verification method and the shape of the artifact. They do not, by themselves, prove buyer demand, real-world error rates, or that production data is this clean. Buyer deployments run on buyer-authorized systems of record.
When an agent can move money, close accounts, or change entitlements, someone has to confirm it actually happened. Veridatum gives buyers source-cited evidence for contested AI-agent outcomes: what held, what failed, what should be paid, and what needs review.
You do not need a platform rollout. You need one scoped verification sprint.
Verification means looking at sensitive records, so the handling matters as much as the method. Here is how a sprint is scoped.
Neither is independent, and independence is the point. A vendor's dashboard reports what the vendor counted; your own team can reconcile it for internal use, but in a dispute the vendor contests your numbers first, and risk or compliance sign-offs are built to distrust a check run by the same side it's checking. A neutral, buyer-owned ledger both sides can reference is harder to wave away. The hard part isn't the reconciliation. It's knowing which record holds the truth, and what "actually held" means when two systems disagree, per action type and across messy systems.
No. Veridatum checks whether a business action actually became true in your systems, not whether usage arithmetic adds up. It can inform what's payable or what to take into a renewal, but it is not invoice reconciliation or spend management.
No. A sprint runs on a redacted export or scoped read-only access. The first step needs no data at all, just a description of the workflow.
Expected. Part of the sprint is mapping agent/vendor claims to the right source-of-truth records and noting where the join is clean and where it isn't. You get those mapping notes either way.
Design-partner terms are deliberately light, and scoped to the workflow and the decision attached to it. We'll quote it on a short call once we've confirmed the workflow is a fit.
One workflow where an agent changes backend state, a redacted sample of the agent's claims and the matching records, and a named decision the result would inform. That's it to begin.
For a buyer being asked to share records, who is behind the work is one of the strongest trust signals there is. This section is a placeholder. Fill it in before sharing the site.
[ One short paragraph: relevant background (verification, payments, audit, ML evaluation, support operations) and why you started Veridatum. Two or three concrete sentences is plenty. ]
[ Optional second profile. Delete this card if you're solo for now. ]
⚠ Placeholder: replace with real names and bios before sharing with prospective design partners.
We're opening a small design-partner cohort for teams with one AI-agent workflow to verify. No data or credentials to start: just a description, and we'll tell you what evidence would verify it.
Apply for a design-partner sprint →or email hello@veridatum.io