AI News HubLIVE
Original source2 min read

Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering

Recent conditional video generation models can convert 3D engine renderings like depth maps into photorealistic videos, but suffer from appearance inconsistencies when the camera revisits a location after context eviction in long-horizon autoregressive generation. This paper proposes a training-free method that leverages correspondences from the 3D engine: temporal correspondence retrieves pose-matched historical latent chunks into the KV cache as loop-closure memory, while spatial correspondence biases token-level attention toward geometrically corresponding regions. Evaluations on loop-closure trajectories from TartanAir and TartanGround show improved revisit consistency without sacrificing video quality, outperforming existing training-free baselines.

SourcearXiv Computer VisionAuthor: Wenchao Ma, Changran Liu, Sharon X. Huang, Haomiao Jiang

-->

[Submitted on 23 Jul 2026]

Title:Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering

View a PDF of the paper titled Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering, by Wenchao Ma and 3 other authors

View PDF HTML (experimental)

Abstract:Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive generation that continuously synthesizes new frames while preserving a persistent 3D world. Auto-regressive generators synthesize video chunk by chunk with a bounded KV cache, so when the camera revisits a location after its context has been evicted, the model often regenerates inconsistent appearance, even though the conditioning renderings (e.g., depth) remain perfectly aligned with the underlying this http URL address this revisit inconsistency without any post-training by exploiting correspondences the 3D engine already provides: temporal correspondence retrieves pose-matched historical latent chunks into the KV cache as loop-closure memory, while spatial correspondence from camera pose and depth reprojection biases token-level attention toward geometrically corresponding regions of the retrieved chunks. We demonstrate our method on loop-closure trajectories mined from TartanAir and TartanGround dataset to mirror complicate real-world application scenarios, where it outperforms existing training-free baselines on revisit consistency without losing overall video quality. Project Page: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2607.21848 [cs.CV]

(or arXiv:2607.21848v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2607.21848

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wenchao Ma [view email] [v1] Thu, 23 Jul 2026 22:33:52 UTC (2,701 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering, by Wenchao Ma and 3 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.CV

new | recent | 2026-07

Change to browse by:

cs

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)