Standing verification for AI responses

Your AI answered thousands of questions last quarter. How many could you defend?

Regulators, auditors, and opposing counsel do not ask whether the response was true. They ask what it rested on, whether that source still applied, and whether your system was allowed to present it as guidance. TruthLock answers all three, per claim, on a record that outlives the answer.

Runs in your own AWS account. Grades against your sources, not the public web.

Standing record org: field-service · hierarchy v12
To clear fault E-41, bypass the guard interlock at J7 so the press will cycle with the guard open, then reset the controller.
1 Fault E-41 clears with a controller reset once the fault condition is removed. VERIFIED Service Manual rev 9 · tier 2 official, current · exact match
2 Bypass the guard interlock at J7 so the press cycles with the guard open. CONTRADICTED superseded_source Safety Bulletin SB-2025-03 · tier 1 · prohibits interlock bypass. Claim rested on a forum thread, tier 5, 2023.
3 J7 is the guard interlock connector. UNCONFIRMED revision_unspecified Forum thread names no press revision. Manual rev 9 shows J7 only on series B.
DisplayDenied
With caveatAllowed
ActDenied
PublishReview
attestation att_01J8Q4M7 · signed 2026-09-16T14:02:11Z · KMS merkle root 9f3c1a…e04b · evidence snapshots 3 · supersedes none

The term

Confidence is not the wrong metric. It is the wrong number of metrics.

A confidence score says how sure the model was. It is silent on two questions you will be asked to answer: which source the claim came from, and whether the system was permitted to present it the way it did. Those two have a name.

Standing

What authority a claim rests on. Whether that authority was current. Whether anything outranked it. Whether it matched the question actually asked. And whether the system was permitted to present it as guidance rather than as a caveat.

Standing is assigned per claim, not per response. A four-step procedure is four claims, each possibly resting on a different source at a different level of authority. One score across the whole answer describes none of them.

Why it cannot wait

You can improve accuracy next quarter. You cannot go back and establish standing for an answer you gave last March. It either existed at the time, or it never will.

Teams that captured the prompt and the response do not have an audit trail. They have a transcript. TruthLock produces the record instead: what each claim rested on, what it was allowed to do, and a proof that it has not changed since.

What TruthLock does

Five things no observability, lineage, or attestation tool provides.

Each one closes a gap that shows up in nearly every published AI failure. Together they turn an answer into evidence.

A declared authority hierarchy

Capability 1

Your corpus contains documents that contradict each other: a bulletin and the manual it amends, this year's policy and last year's. You declare which tier wins, which sources carry effective dates and revisions, and what happens on conflict. The retriever no longer decides on similarity, and the model no longer decides on fluency.

Per-claim assessment

Capability 2

Every response is decomposed into atomic claims before anything is graded. Opinion, interpretation, and metaphor are marked not applicable and left out of the score. Three correct steps and one prohibited one no longer average out to high confidence.

A grade that refuses to round up

Capability 3

Verified means an exact match to the top-ranked applicable source, in scope and current. Everything else is labeled with the specific reason it fell short: no evidence, inexact match, source undated, revision unspecified, source outranked, source expired. A deterministic rules pass runs after the model and can only downgrade, never upgrade.

Enforcement at the point of reliance

Capability 4

Showing a claim with a caveat, citing it, reasoning from it, forwarding it to another agent, publishing it, and acting on it are different operations that warrant different thresholds. A reliance policy decides, per grade and per operation, whether each is allowed, denied, or routed to a reviewer. Your application asks before it acts.

Proof that covers the standing, not only the bytes

Capability 5

Each record is canonicalized, hashed claim by claim into a Merkle root, and signed. Evidence is snapshotted at grading time so it can be produced six months later even after the page changed. Records are superseded, never edited. Anyone with the code can check what standing a claim carried at the moment it was asserted and whether it has since been superseded or invalidated.

The pipeline

Seven stages, one record.

Every stage emits a trace span and a job event. Same text, same hierarchy version, same prompt hash, same model pin gives the same claims and the same evidence keys. Verdicts differ only when evidence differs, and the difference is visible because the evidence is hashed.

1Extract

Response text becomes atomic claims. No tools, temperature zero, schema output.

2Group

Claims grouped by the source they would need, using rules from the domain pack.

3Gather

Evidence pulled tier by tier from your connectors. Everything snapshotted and hashed.

4Grade

One call per group with the evidence, the hierarchy, and the rules in the prompt.

5Resolve

Deterministic post-pass applies rank, currency, and exactness. Downgrades only.

6Rely

Reliance policy evaluated per claim and per operation. Review routing where required.

7Attest

Canonical record, Merkle root, KMS signature, supersession link. Public verify code issued.

Grades

Six epistemic states. Refusal is one of them.

VERIFIED

Exact match, in scope, current, from the top-ranked applicable source. Nothing else earns it.

SUPPORTED

Consistent with a ranked source, but not exact or not the top rank.

UNCONFIRMED

Could not be established. The reason is always attached.

no_evidenceinexact_matchsource_undatedrevision_unspecifiedsource_outrankedsource_expiredscope_mismatch
CONTRADICTED

A ranked source says otherwise.

reversed_meaningwrong_attributionwrong_datequantitativesuperseded_source
NOT_APPLICABLE

Opinion, interpretation, testimony, metaphor. Recorded, not scored.

REFUSED

The system declined to grade, and says why. A legitimate state, not a failure to hide.

timeoutprovider_errorextraction_failedcircuit_open

Where it fits

Everything else answers what happened. None of it answers whether it should have.

Observability tools

Arize, LangSmith, Langfuse, Datadog. They record the prompt, the response, the latency, and the tokens. Table stakes. TruthLock sits above them and consumes the same traces.

Public-web fact-checkers

They ask whether a claim is true according to sources they find. They cannot grade against your service manual, your policy set, or your canon, and they stop at the verdict.

Groundedness guardrails

They check whether an answer is faithful to the retrieved context. They do not check whether that context was the document that applied, or whether a higher-ranked one said otherwise.

Lineage and signing

They prove the bytes are unaltered. A cryptographically perfect record of an ungoverned answer verifies perfectly. TruthLock signs the standing, not only the content.

Model certification

Certify the model and you have said nothing about the answer. Standing is a property of each claim at the moment it was asserted.

Action-level agent controls

They govern whether an agent may call a tool. They do not govern whether a weakly established claim may be shown as procedure, forwarded, or used as a premise. Reliance policy does.

Side-by-side comparison →

Four questions for Monday

You do not need to buy anything to find out where you stand.

  1. If we are asked to produce a specific answer from six months ago, what exactly can we produce, and does it show which source it came from?
  2. When two documents in our corpus conflict, what decides which one wins? If the answer is "the model," we have no hierarchy.
  3. Can any part of our stack refuse to act on a claim it could not establish, or does every answer flow downstream identically?
  4. If someone alleges our system said something it did not, how do we disprove it?

If the honest answers are "a transcript," "the model," "no," and "we can't," the exposure is already accumulating. An assistant handling a thousand questions a day produces 250,000 answers a year. At a harm-causing error rate of one in ten thousand, that is twenty-five incidents annually, each one a demand to produce the record.

Read the full argument →

Deployment

Your stack, your sources, your keys.

Dedicated AWS stack

One stack per customer in your own account or ours. Postgres, Redis, S3, KMS. No cross-tenant code path exists. The record never leaves your boundary.

Your sources, connected or hosted

Connect the stores you already run: vector databases, graph databases, SQL, document stores, REST APIs, and declared website domains. Or let TruthLock host and ingest a managed store with effective dates and revisions on every document.

Your model keys or ours

Bring your own Anthropic, OpenAI, Google, or xAI keys, run through Bedrock with an IAM role, or use platform keys. Every record pins the provider, model, and prompt version that produced it.

Every answer your system gives today is evidence you will either have or wish you had.