2026-06-29 04:00 UTCOriginal source2 min readUpdated: 2026-06-29 08:04 UTC

Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents

Large language model (LLM) agents struggle to update facts in long-term interactions. Replacing full context with bounded memory drops accuracy from 92% to 77% even on frontier models. The gap scales with conversation length, not memory size. The authors introduce Supersede, a reinforcement learning environment that trains agents to prioritize current facts over superseded ones. Fine-tuning Qwen2.5-3B in this environment nearly doubles held-out accuracy (9.0% to 16.7%).

SourcearXiv Computational LinguisticsAuthor: Vedant Patel

[2606.27472] Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents

[Submitted on 25 Jun 2026]

Title:Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents

View a PDF of the paper titled Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents, by Vedant Patel

View PDF HTML (experimental)

Abstract:Large language model (LLM) agents operate over long, multi-session interactions in which facts change: a user moves, a price updates, a plan is revised. Acting correctly requires using the current value of a fact and discarding values that have been superseded. We isolate this ability on real conversational data and show that it is a distinct, unsolved failure. On the knowledge-update subset of LongMemEval, replacing an agent's full context with a bounded, self-maintained memory drops accuracy from 92% to 77% even on a frontier model (gpt-5.4), a gap that is statistically significant (paired McNemar p

new | recent | 2026-06

Change to browse by:

cs cs.AI cs.LG

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)