THE VERY GOOD GUYS / R02 / WORKING PAPER
Is the answer based on the right business records?
An AI assistant can sound sure even when records are missing or it reads your numbers the wrong way.
Read the full paper ↓- What the evidence shows
- The review found four things to check: which records the assistant can read, what the numbers mean, what it remembers and which changes a person must approve.
- The limit
- This looks back at saved work. It does not show that the assistant can work on its own or how often it gives a right answer in daily use.
- The practical decision
- Check which records the answer used, what the numbers mean and who can approve a change.
FULL WORKING PAPER
Evidence Obligations in Operational AI: From Conversational Answers to Verifiable Workflow Transitions
Retrospective evidence and proposed evaluation. Historical checks, observed results and future tests are distinguished in the paper.
Abstract
An operational assistant can produce a plausible answer while reading an incomplete population, applying the wrong business definition or lacking authority to act. We compare a September 2026 observer-system audit with later reconciliation and policy implementations. The audit reports capped reads, misleading event definitions and discarded scheduled output. Later artifacts introduce coverage states, account-scoped identity, transactional reconciliation and policy checks that deliberately stop before execution. A saved cache records 47 test-file outcomes, not a complete newly reproduced test run. From these materials we propose an evidence-obligation graph: a task-specific account of what must be established before a conclusion or workflow transition is accepted. The contribution is a framework and falsifiable protocol, not proof of autonomous deployment or causal improvement. This TVGG working paper is a retrospective systems analysis and has not been peer reviewed. [A1] [A2] [A3]
Keywords
- operational agents
- evidence coverage
- authorization
- reconciliation
- distributed systems
- agent evaluation
Contributions
- An artifact-grounded distinction between source coverage, business meaning, authority and confirmed completion across separate implementation versions.
- A proposed evidence-obligation graph binding operational claims to their scope, source versions, policy and required observations.
- A synthetic fault-replay protocol measuring unsupported conclusions, withheld useful work, authority violations and recovery burden separately.
1. Introduction
A question such as which orders require attention appears conversational, but its answer depends on a chain of operational assumptions. The system must retrieve the relevant population, distinguish a shipping label from a physical handoff, resolve account identity and understand what attention authorizes. If it then schedules a message or proposes a change, completion introduces further requirements. Language quality is only one component of the result.
The case examined here contains a historical observer audit and later source artifacts. These reveal a progression in engineering questions, from access to provider data toward explicit treatment of missing evidence, identity, replay and policy. They do not establish a continuous production deployment or prove that the later code was caused by any particular incident. The legacy assistant, newer observer and later control system remain distinct implementations. [A1] [A2]
Our question is therefore narrower than whether an agent can run a business: what evidence must be available before a particular operational conclusion or transition can be accepted? We propose a task-specific representation of those obligations. The framework is intended to make evaluation and failure diagnosis more precise, while preserving the possibility that additional controls impose costs or withhold useful work.
3. Materials, version boundaries and method
The first source is a technical audit dated September 20, 2026. It describes the newer observer application and attributes its own runtime checks. The second is a later implementation snapshot covering reconciliation, polling and action-policy behavior, including regression history. The third is the retained test material and runner cache inspected in October. No provider write, live-system probe or full test suite was rerun to produce this paper. [A1] [A2] [A3]
We classified each claim by what its source can establish. An audit supplies dated observations and reviewer conclusions. Source establishes implemented logic in a snapshot. A test definition establishes an intended assertion, while cached outcomes supply limited execution evidence. These categories cannot be substituted for one another. In particular, a source-level guard does not prove every deployed request passed through it, and a passing file-level cache cannot identify all assertions, skips or environment conditions.
The comparison asks which obligation a documented defect left unsatisfied and which later mechanism addresses that class of problem. This is analytical correspondence, not causal attribution. We avoid pooling the observer's capabilities with the control system's tests or describing policy checks as executed actions. Anonymous descriptions preserve the technical relationship without releasing customer records, account identifiers or private operational values.
4. Analysis of the retained evidence
The audit reports that provider reads were capped and that partial source failures could remain embedded in an otherwise successful combined response. A correct calculation over retrieved records therefore could not establish a complete operational total. The limitation is epistemic: the system might have accurate values for observed rows while lacking evidence about the requested population. A source failure and a genuinely empty population must produce different interpretations. [A1]
The same audit identifies semantic substitutions. Orders created since midnight contributed to a due-today measure, while label creation was used as a shipping signal. Those events need not coincide with promised shipment deadlines or carrier possession. The audit also describes a scheduled task whose generated output was discarded rather than delivered. These findings separate three possible failures: insufficient evidence, an incorrect business definition and an uncompleted communication. They are historical findings, not assertions about every current branch. [A1]
Later reconciliation artifacts implement explicit coverage information, account-scoped identities and not-evaluated outcomes for unavailable populations. Test definitions address ambiguity, partial coverage, replay and transactional completion. The saved cache contains 47 file entries with no failed flags, including eight database test-file entries. It does not establish an assertion count, absence of skips, commit binding or live integration health. The evidence supports implemented mechanisms and retained testing activity, not independently repeated certification. [A2] [A3]
The action-policy component checks request scope, credentials, nonces, expiry, allowances and replay behavior. Regression history includes cross-account idempotency and denial-handling corrections. Its documented endpoint is policy_checked: it cannot approve, enqueue or execute a provider write. This boundary is part of the engineering result. Describing it as completed autonomous operation would attribute a capability the implementation explicitly excludes. [A2]
| Artifact | Supports | Does not establish |
|---|---|---|
| Dated observer audit | Historical reviewed limits | Current defects in every deployment |
| Reconciliation and policy source | Mechanisms present in inspected snapshots | Live enforcement or autonomous execution |
| 47 cached test-file entries | Retained file-level outcomes | Assertion counts, skips or complete environment reproduction |
5. Proposed contribution: an evidence-obligation graph
We propose representing an operational task as claims and transitions connected to explicit evidence obligations. For a shipment-attention report, obligations might include a defined time window, known source coverage, the correct shipment event and an authorized audience. For a proposed mutation, further obligations would cover request identity, current authority and the expected effect. Delivery or external execution introduces an observation obligation after the action, not merely a successful request before it.
Each obligation records its scope, evidence version, observation time, verifier and status: pass, fail or unknown. The graph expresses dependency rather than a universal sequence. Read-only analysis does not need proof of a provider mutation; an approved action does. The required set must therefore be declared per task. Calling every absent obligation not applicable would defeat the framework, while requiring every possible check would make ordinary work impossible.
The proposed acceptance rule below applies only to the obligations declared for that task. Evidence is bound to the account, request and relevant source or policy versions so a successful check cannot casually transfer to another context. A coverage check may become stale after a source changes; authority may expire before execution. The graph must record invalidation as well as progress. It is an audit representation, not a proof that every verifier is correct.
Proposition one is that explicit coverage obligations reduce unsupported population-level conclusions under partial reads. A counterexample is urgent work where bounded evidence legitimately supports a narrower action: indiscriminate withholding could be worse. Proposition two is that scope-bound authority checks reduce incorrect reuse of approval across retries. It fails if enforcement paths bypass the verifier or if the underlying identity is wrong. Proposition three is that a separate completion observation detects generated-but-undelivered work. A counterexample is an unreliable receipt source that reports success without establishing the intended application effect.
The proposal differs from simply adding a checklist by making dependencies, unknown states and invalidation inspectable at runtime. It differs from a final-state benchmark by explaining which unfulfilled obligation prevents acceptance before a final outcome exists. Its distinctive claim is about a useful representation for these cases, not priority over existing workflow, provenance or assurance methods. A broader novelty review and comparative evaluation remain necessary.
Accept(q, t) = AND over j in J(q) of [Vj(q, t) = pass]
J(q) is the declared set of required obligations for task q. Each verifier Vj returns pass, fail or unknown at time t. Unknown blocks the corresponding acceptance claim; it does not establish a negative business fact. Requirements vary by task.
B = (account, request, source versions, policy version, intended effect)
A proposed binding for evidence and decisions. An obligation must be reconsidered when relevant components change. This record alone does not enforce authorization or guarantee an external effect.
6. Evaluation protocol and counterfactual tests
Begin with a synthetic replay environment whose source populations, business events and authorization rules are known. Freeze task definitions and acceptance conditions before running treatments. Include complete reads, truncated pages, unavailable providers, stale snapshots, duplicated events, ambiguous account references and concurrent requests. Keep external provider execution outside the environment so failure experiments cannot affect real operations.
Compare a baseline summary path with one that retains explicit coverage and semantic checks. Separately compare unbound replay handling with account- and request-bound policy decisions. The separation prevents a combined improvement from being assigned to the wrong control. Hold model configuration and source responses fixed for paired comparisons, then repeat stochastic components. AgentDojo motivates including hostile instructions in retrieved data, but its attack results must not be treated as measurements of this system. [3]
Measure unsupported conclusions, correctly withheld claims, missed actionable issues and recovery after missing evidence arrives. For authority, measure incorrect acceptance and incorrect denial separately, including revocation between checking and proposed execution. For completion, inject lost acknowledgments, delayed observations and repeated requests. Record duplicate intended effects and unresolved outcomes rather than claiming exactly-once behavior from a local idempotency key. This is an application of the distinction between lower-level exchanges and application-level correctness, not a new distributed-systems guarantee. [5]
A useful result must show its costs. Report latency, operator interventions and time spent on unresolved cases alongside error reduction. A system that refuses every task can avoid unsupported conclusions while accomplishing nothing. Require a declared minimum of useful task completion, examine failures by scenario and retain all exclusions. Only after these tests should a supervised field comparison evaluate operator-reviewed outcomes. Neither the synthetic study nor the field phase has been run for this paper.
7. Limitations
The evidence is selected and retrospective. The audit reflects one version, the later implementation reflects another, and source chronology does not establish causal learning from a specific defect. Saved test outcomes do not demonstrate deployed coverage, production reliability or an autonomous warehouse. The paper supplies no field accuracy distribution, labor-saving estimate, commercial return or controlled before-and-after comparison. [A1] [A2] [A3]
An obligation graph can also create false confidence. Incorrect business definitions, compromised verifiers or stale evidence may produce convincing pass states. More logging can obscure responsibility rather than clarify it. The proposal must therefore be assessed against simpler alternatives and include review of verifier quality. Its six practical concerns are a case-derived framework, not a validated universal taxonomy. Private source pointers remain outside the public paper, limiting independent reproduction of historical claims; synthetic replay would provide a separate reproducible test of mechanisms.
A simpler checklist may be sufficient for a low-volume process with one trusted source and direct human review. The proposed graph earns its additional complexity only where changing scope, multiple sources or asynchronous effects make evidence dependencies difficult to track. The practical recommendation is to begin with the smallest unresolved obligation in a bounded workflow, compare a simpler alternative and expand only when the added record changes a real acceptance or recovery decision.
8. Conclusion
Operational reliability depends on what an answer is entitled to claim and what evidence establishes the next transition. The retained audit and later implementation make those questions concrete: coverage differs from arithmetic correctness, business events differ from convenient proxies, and a policy decision differs from execution. Our proposed evidence-obligation graph preserves those distinctions and makes their costs testable. It should be accepted only if evaluation shows that it helps operators complete justified work with manageable recovery effort. The present contribution is that bounded proposal and its evidentiary basis, not a declaration that the business has become autonomous.
References
Numbered entries link to public literature. Entries beginning with A describe retained private materials; identifying records and archive locations are not published.
- [1]
ReAct: Synergizing Reasoning and Acting in Language Models (opens in a new tab)
ICLR 2023; arXiv:2210.03629, version 3
Cited for reasoning-action interaction; no benchmark run was performed here.
https://arxiv.org/abs/2210.03629v3
- [2]
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (opens in a new tab)
arXiv:2406.12045
Cited for explicit final-state evaluation and consistency across trials.
https://arxiv.org/abs/2406.12045
- [3]
arXiv:2406.13352
Cited for adversarial evaluation over tool-returned data, not evidence of an attack on the archived system.
https://arxiv.org/abs/2406.13352
- [4]
The Protection of Information in Computer Systems (opens in a new tab)
Proceedings of the IEEE 63(9), 1278–1308
Primary author-hosted text, particularly design principles in Part I.
https://web.mit.edu/Saltzer/www/publications/protection/
- [5]
End-to-End Arguments in System Design (opens in a new tab)
ACM Transactions on Computer Systems 2(4), 277–288
Primary author-hosted paper; cited for function placement and application-level correctness.
https://web.mit.edu/Saltzer/www/publications/endtoend/endtoend.pdf
- [A1]
Observer-system technical audit, September 20
Unpublished dated engineering assessment
Anonymized descriptive label. Historical findings and reported runtime checks are attributed to the audit, not newly reproduced.
Private source · Description only - [A2]
Reconciliation and action-policy implementation, September snapshots
Unpublished source and regression history
Anonymized description of distinct later components. Policy checking stops before approval or provider execution.
Private source · Description only - [A3]
Reconciliation test definitions and retained runner metadata
Unpublished test artifacts inspected October 1
47 cached file entries with no failed flags. Not a complete assertion-level or commit-bound execution report.
Private source · Description only