跳到主要內容
AI News HubLIVE
站內改寫2 分鐘閱讀

待翻譯:Shared Selective Persistent Memory for Agentic LLM Systems

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use patterns that made previous sessions productive. Naively persisting entire conversation histories is both token-inefficient and counterproductive—irrelevant context degrades generation quality. We introduce shared selective persistent memory, a memory architecture for agentic systems that identifies and retains four categories of reusable context—task specifications, data…

待翻譯:Shared Selective Persistent Memory for Agentic LLM Systems
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

content type paperpublished September 2026 Shared Selective Persistent Memory for Agentic LLM Systems AuthorsSanjana Pedada, Aditya Dhavala, Neelraj Patil View publication Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use patterns that made previous sessions productive. Naively persisting entire conversation histories is both token-inefficient and counterproductive—irrelevant context degrades generation quality. We introduce shared selective persistent memory, a memory architecture for agentic systems that identifies and retains four categories of reusable context—task specifications, data schemas, tool configurations, and output constraints—while discarding session-specific reasoning traces. Crucially, this memory is shared: workspaces encapsulating selective memory can be transferred across users with role-based access control, enabling collaborative reuse of accumulated context without redundant specification. We implement this architecture in a deployed collaborative workspace platform where LLM agents produce, edit, and maintain git-versioned artifacts—including interactive dashboards, structured reports, and data-driven documents—from heterogeneous data sources accessed via multiple connector types (CSV upload, SQL, REST APIs, and MCP servers). Git-backed versioning with draft isolation enables users to explore modifications risk-free and restore to any prior state without re-invoking the model. A complementary zero-token data refresh mechanism decouples generated programs from runtime data, enabling artifact reuse without re-invocation. Across three enterprise deployment scenarios, shared selective persistent memory achieves 96% task completion (vs. 79% without memory and 71% with full history). A complementary zero-token data refresh mechanism eliminates LLM re-invocation entirely for recurring data updates (14×task time reduction), while summary-driven generation reduces per-invocation token cost by 97×versus raw data injection. A replication on four public datasets confirms generalizability, with zero-token refresh succeeding in 12/12 trials. Notably, naive full-history persistence actively degrades task completion by biasing the agent with stale reasoning traces, while selective memory outperforms both extremes. Agent Seer: Synthesizing Scenarios from Specification Understanding August 28, 2026research area Data Science and Annotation, research area Tools, Platforms, Frameworks Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, and produces static benchmarks that cannot track evolving APIs. We observe that tool specifications—function names, natural-language descriptions, and typed parameter… Read more Less Is More: A Unified Architecture for Device-Directed Speech Detection with Multiple Invocation Types June 8, 2023research area Methods and Algorithms, research area Speech and Natural Language Processingconference ICASSP Suppressing unintended invocation of the device because of the speech that sounds like wake-word, or accidental button presses, is critical for a good user experience, and is referred to as False-Trigger-Mitigation (FTM). In case of multiple invocation options, the traditional approach to FTM is to use invocation-specific models, or a single model for all invocations. Both approaches are sub-optimal: the memory cost for the former approach grows… Read more

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain…

技術影響

可能影響 Agent 架構、工具調用、工作流自動化和產品集成。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。