THE VERY GOOD GUYS / R01 / WORKING PAPER
Can the next tool use the AI answer?
An AI tool can say the job is done, yet give back an answer the next tool cannot use.
Read the full paper ↓- What the evidence shows
- Of 116 saved results, 112 were marked as a success. Only 28 put the answer in the exact place the next tool expected. The other 84 put it inside another part of the response.
- The limit
- This checks where the answer was placed in one group of results. It does not tell us whether the answer itself was right.
- The practical decision
- Set the format the next tool needs. Check each answer before passing it on.
FULL WORKING PAPER
Completion, Conformance and Evidence: A Retrospective Study of a Website-Classification Pipeline
Retrospective evidence and proposed evaluation. Historical checks, observed results and future tests are distinguished in the paper.
Abstract
A business classification pipeline can finish a request without satisfying the interface that makes its output usable. We examine a retained batch containing 116 distinct domain records, of which 112 were marked successful. The requested target field appeared directly in 28 successful payloads and nested within all 84 others; 26 payloads satisfied a retrospective literal interpretation of the six-field prompt. These are structural observations, not estimates of factual accuracy. Source inspection identifies acceptance of returned objects without schema validation and a restart rule that skips saved errors. From this case we propose a task acceptance ledger separating recorded return, direct conformance, normalized conformance, evidence acceptance and retry disposition. We specify counterexamples and an evaluation protocol rather than claiming demonstrated improvement. This is a TVGG working paper, not a peer-reviewed study or a new model experiment. [A1] [A2] [A3]
Keywords
- agent evaluation
- structured outputs
- schema conformance
- retrospective analysis
- task completion
Contributions
- A reproducible distinction between recorded success, direct field presence and literal contract conformance in one operational artifact.
- A proposed task acceptance ledger preserving transformations, unresolved evidence and retry eligibility without classifying nested answers as factual failures.
- A falsifiable evaluation protocol separating structural recovery from supported classification and operator effort.
1. Introduction
A company-research workflow has at least two audiences. A person may interpret an explanation flexibly, while a downstream program requires particular fields, types and locations. An answer can satisfy one audience and frustrate the other. A dashboard reporting successful calls may therefore conceal additional interpretation work without establishing that the underlying answers are wrong. Our research question is whether an operational success flag corresponds to the interface the workflow itself requested.
The case is a saved website-classification batch whose prompt requested company identity, product description, local-market status, a sales-team classification, supporting evidence and a directory flag. Its implementation saved results and marked requests successful when its processing path completed. The archive permits comparison between that flag and the requested structure. It does not preserve independently adjudicated factual answers for every domain, so the present analysis concerns acceptance behavior rather than employer truth or commercial value. [A1] [A2]
When unresolved structure is hidden by a success flag, another process must fail, guess or introduce a repair. Each response changes the system being evaluated. We argue that acceptance and repair should be inspectable parts of the workflow. This is a proposed framework derived from a bounded retrospective case, not a claim that output validation is a newly discovered problem.
3. Materials and retrospective method
The materials are a saved JSONL artifact, its associated Python implementation and an October 2026 aggregate audit. Each nonempty line was treated as one retained record. The 116 records represent 116 unique domains, so there are no repeated-domain rows. This does not establish how many underlying requests or internal retries occurred. We use the saved record as the denominator, without requesting new pages, invoking models or altering original records. [A1] [A3]
The first measure counts records whose original success flag is true. The second requires the exact target key directly within the returned object. The third requires exactly the six requested fields: three strings, two lowercase yes/no/unclear categories and one Boolean. This literal interpretation was defined after the run; it was not a production validator. A separate recursive check locates the target inside objects and lists without normalizing field names, parsing additional strings or judging truth. [A3]
Source inspection supplies a mechanism-level interpretation. The program unwraps a content field when its value is a string, attempting to decode JSON. Object and list wrappers can remain intact. Returned results receive a success flag without six-field validation. The restart procedure builds its completed-domain set from every saved row, including errors. File metadata suggests a June run, and a later committed snapshot retains both source and output; neither binds each response to an exact execution time or model revision. [A2]
4. Results and implementation analysis
There are 112 records marked successful and four marked as errors. Of the successful payloads, 28 contain the target field directly, and those same 28 contain all six requested fields. Only 26 pass the literal type and value check. Thus field presence and contract conformance differ even before nested results are considered. These are exact counts for a finite artifact, not a sampling estimate of performance across all company-research tasks. [A3]
All 84 successful payloads lacking the direct field contain it deeper inside a content wrapper. No successful payload lacks a structurally locatable occurrence of the target. Describing those 84 as missing answers, wrong classifications or necessarily unusable records would be incorrect. Their existence motivates normalization research, but recursive key presence cannot establish which nested object should be selected, whether supporting evidence justifies its label or whether every required value is well typed. [A1] [A3]
A separate recovery issue follows from the source. Appending and flushing records protects partial progress, but the restart rule makes every recorded domain ineligible for ordinary reprocessing. An error can therefore be durably preserved while disappearing from the usual retry queue. This is an inspected code behavior. The archive does not establish whether a restart occurred, whether an error remained unresolved indefinitely or whether an operator used another correction procedure. [A2]
Cflag = 112/116 = 96.6%; Cdirect = 28/116 = 24.1%; Cstrict = 26/116 = 22.4%
All rates use saved domain records. They measure different acceptance conditions, not factual accuracy.
| Condition | Count | Meaning |
|---|---|---|
| Saved unique domains | 116 | Retained tasks, not all underlying requests |
| Original success flag | 112 | Recorded processing result |
| Direct target and all six fields | 28 | Field presence |
| Target nested only | 84 | All marked successful; no normalization |
| Literal contract passes | 26 | Exact fields, types and categories |
5. Proposed contribution: a task acceptance ledger
We propose recording acceptance as a small evidence ledger rather than one terminal Boolean. Each task retains its raw response, versioned consumer contract, transformations, validation decisions and closure rationale. The aim is to expose hidden interpretation and recovery work. This is our proposed synthesis from the case, not a validated model of all agent workflows.
The ledger distinguishes direct conformance from conformance after an approved transformation. Evidential acceptance remains separate because a valid object may contain an unsupported claim. Unknown is substantive: an unreviewed answer is neither accepted nor rejected. Retry disposition records whether a task is eligible for another attempt, permanently blocked, awaiting review or complete under a declared contract. Changed sources or a revised contract can reopen a previously accepted task.
Proposition one is that a deterministic, versioned normalizer increases structural yield when wrappers preserve an unambiguous valid object. It fails when wrappers contain competing objects or invalid values. A normalizer that silently chooses the most convenient answer may increase apparent yield while introducing selection errors. Recovery should therefore preserve the original artifact and abstain when transformation is ambiguous.
Proposition two is that outcome-aware resume logic reduces premature closure under recoverable failures compared with skipping every saved domain. A counterexample is a permanently inaccessible source: retries add cost without evidence. Proposition three is that separating structural and evidential acceptance exposes unsupported but well-formed answers. An unreliable reviewer is a counterexample because the added judgment may contribute noise. These claims require measurements of correctness, effort and cost, not simply more elaborate status fields.
L(i) = (T(i), D(i), N(i), E(i), R(i))
T records return status; D direct validation; N validation after a declared normalizer; E evidence acceptance; R retry disposition. Validation can pass, fail or remain unknown. This proposed vector was not measured for every archived record.
6. Evaluation protocol
A follow-up should freeze a permitted source corpus and task specification. Saved payloads support structural recovery research without new browsing, but factual review requires retained source evidence or a separately dated collection. Reviewers should define the label rubric before viewing treatment results, adjudicate ambiguous cases independently and retain disagreement. The unit remains a domain task, with attempts grouped beneath it.
Compare the original acceptance rule, strict validation with bounded retries, and declared normalization followed by validation and bounded retries. Hold the model, source snapshots and attempt budget fixed when testing acceptance policy. Record first-attempt conformance, final conformance, supported accuracy, unresolved cases, operator time and total attempts. Repeated generation requires randomized treatment order and repeated trials so acceptance policy is not confused with stochastic variation. τ-bench motivates attention to consistency, but eventual success after retries is not equivalent to repeated-trial reliability. [4]
Use paired task comparisons and uncertainty estimates over tasks, rather than counting retries as independent observations. Include ambiguous wrappers, duplicate fields, invalid categories and persistent access failures. Revise the framework if accepted yield rises while supported accuracy falls, if operators spend more time maintaining ledger states than repairing baseline outputs, or if costs grow without reducing unresolved work. Predeclared failure criteria prevent the richer record from becoming a retrospective explanation for every possible outcome.
7. Limitations and research ethics
One retained batch cannot establish model-wide reliability or comparative framework quality. Selection into the original list was not randomized, the strict contract was retrospective, and complete source snapshots and request-level settings are absent. Structure supports repeatable counting but not a causal account of wrapper formation. No outreach, hiring, revenue or labor result was measured. [A1] [A2] [A3]
An intentional framework wrapper is an important opposing interpretation. If the documented consumer interface permits that wrapper, its presence is not an upstream defect. The problem would instead lie in an adapter or in treating the prompt as the complete application contract. Our measurements remain reproducible under the stated rule, but the operational importance of the mismatch depends on the actual consumer. A follow-up should inspect that consumer before deciding that regeneration is preferable to a small, deterministic adapter.
Private company records are unnecessary for explaining this mechanism and are excluded here. A public replication package should use synthetic payloads preserving relevant validation and wrapper conditions. That would reproduce the acceptance behavior without claiming to reproduce the original factual task. Normalization benefits, retry improvements and reviewer agreement remain hypotheses until the proposed experiment is executed and its negative results are retained.
8. Conclusion
The artifact demonstrates a gap between recorded success and the requested interface, while showing why that gap requires careful interpretation: all 84 indirect responses contain the target somewhere in their structure. The defensible finding is an acceptance mismatch, not an accuracy collapse. The proposed task acceptance ledger makes the distinction testable by preserving returns, validating contracts, documenting repairs, evaluating evidence and retaining unresolved work. Its value should be judged through supported completion and recovery burden, not through the number of states displayed on a dashboard.
Replication companion
Run a small, synthetic example of direct, nested and invalid payloads against the stated output rule. The companion includes example records, the validator, expected counts and run instructions. It reproduces the structural distinction, not the original company classifications or their factual accuracy.
Download the self-contained Python example
Download the published aggregate counts (CSV)
The aggregate CSV transcribes the retained-batch counts reported in this paper. The example records are invented and are not rows from the private source.
References
Numbered entries link to public literature. Entries beginning with A describe retained private materials; identifying records and archive locations are not published.
- [1]
JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models (opens in a new tab)
arXiv:2501.10868, version 3
Current title verified; earlier versions used a different title.
https://arxiv.org/abs/2501.10868v3
- [2]
ReAct: Synergizing Reasoning and Acting in Language Models (opens in a new tab)
ICLR 2023; arXiv:2210.03629, version 3
Cited for reasoning-action interaction, not evidence about the local pipeline.
https://arxiv.org/abs/2210.03629v3
- [3]
SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (opens in a new tab)
ICLR 2024; arXiv:2310.06770
Cited for execution-based evaluation; no benchmark score is transferred to this case.
https://arxiv.org/abs/2310.06770
- [4]
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (opens in a new tab)
arXiv:2406.12045
Cited for goal-state evaluation and repeated-trial reliability.
https://arxiv.org/abs/2406.12045
- [A1]
Retained website-classification responses
Unpublished operational artifact
Anonymized descriptive archive label for 116 saved records; not a published document title. Exact provenance retained privately.
Private source · Description only - [A2]
Website-classification implementation snapshot
Unpublished source artifact
Anonymized descriptive label for the prompt, result handling and restart procedure. Configuration is not a per-request receipt.
Private source · Description only - [A3]
Aggregate classification-structure audit, October 1
Unpublished retrospective audit
Artifact hash, post hoc contract definition and aggregate counts; no model rerun or factual labeling.
Private source · Description only