THE VERY GOOD GUYS / R03 / RESEARCH ANALYSIS
Does the video keep the room the same?
An AI video can look right at the start and end, yet change the room or building in the middle.
Read the full paper ↓- What the evidence shows
- A saved review passed three video shots and rejected three. The rejected shots changed key shapes or added things that did not belong.
- The limit
- This is a small review from earlier work. A new reviewer did not score the videos again for this paper.
- The practical decision
- Watch the whole shot against the real reference. Check that the room and its layout stay the same.
FULL WORKING PAPER
Endpoint Fidelity Is Not Route Fidelity: A Constraint Protocol for Generated Site Video
Retrospective evidence and proposed evaluation. Historical checks, observed results and future tests are distinguished in the paper.
Abstract
Generated video of a real place must preserve more than an attractive beginning and ending. This paper examines a saved review of six generated golf flyover shots: three passed and three failed the original geometry gate despite endpoint agreement. The videos were not independently rewatched for this paper. [A1] We develop a proposed route-constraint certificate that records which landmarks and spatial relations are supported at intermediate positions, distinguishes absence from occlusion, and prevents an aggregate quality score from concealing a critical terrain error. The contribution is a domain-specific acceptance representation and a falsifiable comparison of complete single-span and midpoint-split production strategies. It is not a new video model or an experimental demonstration of improved fidelity.
Keywords
- video generation
- spatial consistency
- reference fidelity
- production evaluation
- human review
Contributions
- A proposed route-constraint certificate connecting reference landmarks, visibility, temporal alignment, and release decisions.
- An analysis of the difference between saved review judgments and independently reproduced video measurements.
- A bounded evaluation design testing midpoint constraints against full-span generation while counting seams and operator effort.
Introduction
A generated flyover can be visually convincing while describing the wrong place. A bunker may disappear, water may become turf, or the camera may travel along a corridor that never existed. Such changes matter when the output represents a specific site. The production requirement is therefore relational: preserve recognizable features and their arrangement as the viewpoint changes. Exact endpoint images constrain two moments but leave the path between them substantially underdetermined.
The motivating artifact is a private stage-gate review inspected on October 1, 2026. Its reviewer accepted three shots and rejected three, then held thirty further planned generations. This was an operational decision within a concept project, not a benchmark sample. Its value lies in the documented mismatch between endpoint agreement and intermediate identity errors. We ask how that mismatch can become a reproducible acceptance test without pretending that a few reviewed shots establish model-wide reliability. [A1] [A2]
Method
The retrospective method reads the saved shot judgments, their stated review procedure, and the production hold. The original reviewer compared normalized positions at zero, twenty-five, fifty, seventy-five, and one hundred percent, then used denser contact sheets and half-speed proxies. This paper inspects those records and their reasoning. It does not independently decode, rewatch, or rescore the source clips, and it does not infer the model's hidden geometry representation from their descriptions. [A1]
For a future reproducible review, the unit should be one generated shot under one frozen reference package. That package must preserve authorized source images, route landmarks, expected visibility, intended camera corridor, generator settings, and all retained attempts. Reviewers should annotate a landmark before scoring its generated counterpart. Otherwise, an appealing output can subtly redefine what the reference was supposed to contain. Difficult or obscured landmarks remain explicitly unassessable.
Evidence and Results
The saved review describes three distinct rejected behaviors: a lateral camera swing away from the mapped corridor, an omitted water hazard accompanied by newly appearing carts, and a missing greenside bunker. Accepted shots retained their sampled course relationships, although two carried pacing notes. Those distinctions are useful: a timing mismatch and an identity-breaking omission should not automatically receive the same disposition. All six judgments remain attributed to the prior reviewer. [A1]
Three of six is the composition of this saved review, not an estimated failure probability. The cases were selected for a production gate, the independent sampling mechanism is unknown, and the review does not provide inter-rater agreement. The strongest supported result is procedural: the documented gate stopped additional generation despite matching endpoints. No completed full-site delivery, public release permission, or benefit from a midpoint intervention is established by these artifacts. [A2]
Proposed Framework: Route-Constraint Certificates
Define a route-constraint certificate as a versioned set of obligations linking reference landmarks to generated observations along a shot. A landmark has an identity, an expected visibility interval, and relations to other landmarks, such as water left of a bunker from an aligned viewpoint. Each obligation receives one of three states: supported, violated, or unassessable. A release requires every critical obligation to be supported; an unassessable critical feature triggers review rather than silently becoming a pass.
The certificate uses route progress instead of assuming that equal percentages imply equal viewpoints. Let phi map generated time to reference progress, constrained to move forward but allowed to vary in speed. Reviewers establish correspondence using stable landmarks and record ambiguous intervals. This permits harmless pacing changes while detecting an incorrect corridor. If no defensible correspondence exists, the certificate cannot establish fidelity, even when an unconstrained image similarity measure looks favorable.
The proposed contribution is this combination of progress alignment, explicit visibility, and non-compensatory critical constraints for site production. Aesthetic strengths cannot cancel a missing hazard. Unlike a full geometric reconstruction, the certificate describes only the reference obligations selected for the intended use. It predicts that midpoint-constrained generation will reduce some unsupported long-span changes, while potentially introducing seams, duplicated landmarks, or acceleration discontinuities near the join. These competing predictions make the proposal testable.
O_j = (landmark, relation, visibility interval, reference version, criticality)
Each obligation identifies exactly what must remain true and when it can reasonably be assessed.
G = 1 iff every critical O_j is supported; C = assessed obligations / required obligations
G is a release gate, while C reports evidence coverage. Neither is a population reliability estimate; unassessable obligations remain in the coverage denominator.
Evaluation Design
A bounded next study should freeze twelve licensed reference spans before generation. Include turns, occlusion, open corridors, repeated vegetation, and small hazards. Produce two retained attempts under each of two conditions: a single full span and two subspans joined around an independently verified midpoint. Forty-eight assembled candidate shots would result. Both conditions use the same target duration, output format, endpoint references, and model version. The split strategy uses two generation calls per candidate rather than one.
The estimand is the difference between these complete production strategies on the frozen spans, including their generation and assembly effort. It is not the isolated causal effect of midpoint guidance. Extra calls, shorter spans, and joining work change together; recording their costs does not remove that confounding. Report fidelity and total effort side by side. Isolating midpoint guidance would require another comparison controlling span length and generation effort, with a selection rule fixed before outputs are inspected.
Two reviewers, blinded to condition where practicable, should score the same frozen obligations and independently report visibility uncertainty. Adjudication should preserve original disagreements. Compare critical-constraint passes within reference span, false passes against an expert reference judgment, unassessable obligations, seam defects, and human review minutes per accepted shot. Report every attempt, including generation errors. A higher pass fraction purchased with substantially more editing may be operationally unattractive even if its geometry is better.
Before testing, agree the criticality rubric, time budget, and rule for advancing the method. A defensible proposed gate requires fewer critical violations without an increase in critical seam errors, plus a review-cost ceiling chosen by the operator. It is not yet an accepted threshold. Repeated attempts from one route are correlated, so raw frame counts must not masquerade as independent sample size. The primary comparison is paired by reference span.
Limitations
The certificate fails when reference coverage cannot disambiguate real disappearance from occlusion, when seasonal or construction changes invalidate landmarks, or when the camera follows an unverified route. Dense foliage and repeated textures can defeat correspondence. It can also reward a static or timid output unless intended motion remains a separate obligation. A reference package may itself contain errors, and two reviewers can share the same mistaken site assumption.
There is also contrary evidence within the motivating review: three shots passed without the proposed midpoint intervention. That finding weakens any blanket recommendation to split every shot. The failed cases might instead require better reference selection, a different allowed camera path, or restrained editing. The proposed comparison should therefore establish where added constraints help, rather than assume that more constraints always improve the result. [A1]
There is no claim that twelve spans would establish general reliability or that midpoint control is universally beneficial. The original six-shot evidence was not reproduced, and model versions may change. Public examples require rights review and de-identification. Technical credibility depends on preserving these limits while making future reference packages and scoring procedures inspectable.
Conclusion
The saved review shows why a production team can reject a video with correct anchors: the represented route matters between them. The proposed certificate turns that requirement into explicit landmark obligations, uncertain observations, and a release gate. Its value remains a hypothesis until controlled comparisons measure both geometry and operator burden.
For the next production decision, retain the failed spans and their reference packets, agree the geometry rubric before rerendering, and test midpoint control only through the paired comparison. If the extra boundary produces more seams or review work without better fidelity, retain the simpler method for that class of shot. The next defensible result is a reproducible evaluation, not a broader claim about generated video's readiness.
References
Numbered entries link to public literature. Entries beginning with A describe retained private materials; identifying records and archive locations are not published.
- [1]
VBench: Comprehensive Benchmark Suite for Video Generative Models (opens in a new tab)
CVPR 2024; arXiv:2311.17982
Primary author manuscript submitted in 2023; conference publication in 2024. Used for multidimensional evaluation.
https://arxiv.org/abs/2311.17982
- [2]
VideoPhy: Evaluating Physical Commonsense for Video Generation (opens in a new tab)
arXiv:2406.03520, version 2
Used for the distinction between visual quality and physical adherence, not as a measurement of the private footage.
https://arxiv.org/abs/2406.03520v2
- [3]
CVPR 2025; arXiv:2407.14505
Primary author manuscript. Used for compositional and relational evaluation.
https://arxiv.org/abs/2407.14505
- [4]
RAFT: Recurrent All-Pairs Field Transforms for Optical Flow (opens in a new tab)
ECCV 2020; arXiv:2003.12039
A possible measurement aid; optical flow is not geographic ground truth.
https://arxiv.org/abs/2003.12039
- [A1]
Six-shot geometry review and review-method record
Saved internal review
Prior reviewer judgments: three pass and three fail. No independent rewatch for this paper.
Private source · Description only - [A2]
Reference mapping and production hold
Internal production scaffold
Records the hold on thirty additional planned generations. Exact provenance retained privately.
Private source · Description only