待翻譯:Show HN: Agent Memory Leaderboard – first public results for AI memory systems
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:OPEN BENCHMARK Agent Memory Leaderboard A public benchmark space for comparing textual and coding-agent memory systems under a consistent evaluation flow. Leaderboard Preview Benchmark Tracks Each track keeps its own re…
AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
OPEN BENCHMARK Agent Memory Leaderboard A public benchmark space for comparing textual and coding-agent memory systems under a consistent evaluation flow. Leaderboard Preview Benchmark Tracks Each track keeps its own result table and detailed metric breakdown. Textual Memory Long-context, persona, script, and conversation-memory benchmarks. Coding Agent Memory Agent memory support for coding tasks and repository-context recall. Evaluation Flow Industry systems use the hosted Add/Search key flow. Academic systems may use the same flow or submit a public GitHub repository for maintainer Docker deployment. 1 Choose an evaluation route Provide hosted Add/Search APIs, or submit a public GitHub repository with Docker and API run instructions. 2 Run a smoke test Use the issued key to verify the synchronous Add/Search flow. 3 Submit a formal evaluation After smoke passes, submit the full scored evaluation. Explore the Platform Use the product pages to inspect rankings, run evaluations, and prepare an integration. 01 Leaderboard Public ranking with filters, dataset columns, and score bars. → 02 Evaluation Create eval jobs, watch progress, and inspect private results. → 03 Participation Guide Eligibility, submission routes, required materials, timelines, rewards, and publication rules. → 04 Documentation User guide, evaluation workflow, API contract, security, and result publication. → 05 Guide Add/search API contract, request fields, polling, and response schemas. → PUBLIC RANKINGS Agent Memory Leaderboard Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions. Leaderboard A First-cycle evaluation results will be released in mid-August. Academic board submissions are open now. Submit your evaluation request before the first-cycle deadline. PUBLIC RANKINGS Agent Memory Leaderboard Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions. Leaderboard A First-cycle evaluation results will be released in mid-August. Academic board submissions are open now. Submit your evaluation request before the first-cycle deadline. EVALUATION Run Evaluations Create API-gated eval jobs against the Leaderboard Suite, monitor task progress, inspect private results, and submit eligible full-suite runs for administrator review. Create Eval Job Choose a bound version. Use Run label to distinguish repeated evaluations of the same version. System name Version name Mode Max add concurrency 16-64 Search concurrency 16-256 Top K Run label Continue the latest interrupted evaluation from its last compatible checkpoint. Datasets FULL EVALUATION GATE Full evaluation checklist Confirm each item before starting a public-board candidate run. The button remains locked until every item is checked. 0 / 8 confirmed Smoke test completedThe Add/Search API has passed the platform smoke test and is currently usable. API contract is followedThe submitted code and deployed interfaces are wrapped according to the official Add/Search format. Add/Search uses gpt-4o-miniThe model used by the submitted memory system during both Add and Search must be gpt-4o-mini. The platform will reproduce the submission; if the reproduced score differs materially, the leaderboard result may be invalidated. Runtime will stay stableIf you provide a deployed endpoint, it will remain publicly reachable and stable for at least 30 days after submission. Run instructions are completeThe repository README or submission notes include the Docker command, API entrypoint, configuration, and startup steps. Original work is disclosedAny reused paper, repository, or code is attributed with its original authors, technical report, and the changes made; personal work is identified as such. This is a substantive submissionIt is not a repeated near-duplicate or a low-quality submission intended to occupy evaluation capacity. No manipulation or cheatingThe system does not use prompt injection, benchmark leakage, result manipulation, malicious behavior, or other leaderboard abuse. Current Job Live job status and progress. No job yet. External Full-Run Review Review completed full-suite runs from non-admin leaderboard keys. Approved results enter the public board; rejected results remain private. Private Results Private scores stay scoped to the current leaderboard key. Admin runs remain separate from external review candidates. PARTICIPATION GUIDE · 2026 首屆 Agent 記憶挑戰賽參賽說明 Agent Memory Challenge 是 Agent Memory Leaderboard 的首期公開評測活動,面向全球研究者、開源專案維護者和商業產品團隊開放。參賽系統負責 Add 與 Search,平臺統一完成 Answer、Eval、結果複核與公榜。 返回賽事頁面 一分鐘瞭解首屆賽事 參賽免費,不設組隊要求。參賽方承擔自身 API、資料庫、頻寬和計算成本,平臺承擔統一 Answer、Eval 與評測編排成本。 01 報名開放 2026 年 7 月 29 日 02 提交截止 2026 年 8 月 7 日 23:59(UTC+8) 03 首次放榜 2026 年 8 月中旬 04 評測型別 文本記憶與程式碼記憶 05 參賽組別 學術方法榜與商業產品榜 第一步:選擇評測型別 評測型別決定系統接受什麼任務;參賽組別決定結果展示在哪個榜單。兩者是互相獨立的兩個維度。 評測型別主要評測內容可選參賽組別首屆時間 文本記憶事實召回、多跳整合、時序理解、記憶治理、個性化、規則執行、安全與隱私。學術方法榜 / 商業產品榜8 月 7 日提交截止 程式碼記憶從歷史工程任務中檢索、篩選並複用除錯經驗、開發經驗和專案上下文。學術方法榜 / 商業產品榜8 月 7 日提交截止 第二步:選擇參賽路徑 第一步始終是提交評測申請。Eval Key 不是另一個報名入口,而是自行部署 API 的申請稽核透過後獲得的評測憑證。 學術 · API 自行部署 Add / Search API 提交公開 GitHub 倉庫、固定版本、Add / Search 地址、鑑權方式和執行說明。參賽方負責部署並保持介面穩定,稽核透過後獲得 Eval Key。 學術 · 程式碼 提交程式碼,由平臺部署 提交公開 GitHub 倉庫、Docker 啟動方式、Add / Search 封裝和完整執行說明。平臺負責構建與評測,不簽發 Eval Key。 商業 · API 提供穩定的產品 API 提交固定產品版本、Add / Search 地址、鑑權和容量說明。無需公開內部實現,稽核透過後獲得 Eval Key,結果進入商業產品榜。 開源要求 學術方法榜必須提供公開、可核驗的 GitHub 倉庫,並披露原始方法、作者、技術報告和本次改動;商業產品榜無需開源,但必須提供可核驗且穩定的產品與 API 版本。 第三步:準備提交材料 請在提交申請前固定參評版本。正式 Full 評測受理後,不得因結果不理想更換版本或撤回。 共同材料 所有參賽者 系統名稱與版本、聯絡人、機構或團隊、擬參評型別、方法或產品說明、允許公開展示的資訊,以及完整的提交說明。 學術方法 程式碼與復現材料 公開倉庫、README、Docker 命令、API 入口、依賴配置、原始工作引用、方法改動和執行步驟。自行部署時還需提供公網 API。 商業產品 介面與執行材料 固定產品版本、Add / Search API、鑑權方式、評測專用金鑰、容量限制、超時與限流說明,並保證介面在提交後至少 30 天穩定可訪問。 從申請到上榜 1 提交評測申請 選擇評測型別、參賽組別和提交方式,提交系統版本與完整材料。 2 完成接入 自行部署 API 的參賽者獲取 Eval Key;程式碼提交由平臺按 Docker 說明構建。 3 Smoke 與 Full 先驗證 Add / Search、鑑權和端到端鏈路,再執行首屆正式 Full 評測。 4 複核與公榜 平臺複核版本、結果和合規狀態,透過後釋出到對應公開榜單。 Eval Key 與 Memory System Key 兩個 Key 的簽發方和用途不同,請勿將它們放入公開倉庫、URL、截圖、郵件正文或群聊。 憑證誰提供用途誰需要 Eval Key / Leaderboard KeyAgent Memory Leaderboard驗證參賽身份、執行評測並檢視私有結果。自行部署 API 的學術與商業參賽者;平臺部署程式碼的路徑不簽發。 Memory System Key參賽方供平臺訪問參賽系統的 Add / Search API。介面啟用鑑權時需要;無鑑權介面無需提供。 獎勵設定 榜單排名獎勵與社群貢獻獎勵是兩套獨立計劃;進入公開榜單不等於自動獲得全部獎勵。 學術排名獎勵 開源方法榜 Top 10 第 1—3 名獲得 ChatGPT Pro 月度會員;第 4—10 名獲得 ChatGPT Plus 月度會員。賬號由賽事方獨立採購併發放,賽事並非由相關產品提供方主辦或贊助。 社群貢獻計劃 滿足任一條件即可入選 前 50 位完成有效提交;成功邀請 3 位新參賽者完成有效提交;或提交 3 組有挑戰性的測試樣本並透過稽核。 Kimi Token 入選即進入獎勵名單 入選社群貢獻計劃後,獲得至少價值人民幣 50 元的 Kimi Token 額度。最終名額、額度、發放時間和稽核結果以官方通知為準。 什麼是有效提交? ✓ 材料完整 參賽身份、系統版本和所需材料均完整、真實且可核驗。 ✓ Smoke 透過 Add / Search、鑑權和端到端鏈路符合現行接入協議。 ✓ Full 完成 正式評測任務和所需評測項成功完成,沒有缺失或重複結果。 ✓ 版本一致 正式評測使用的程式碼、映象或 API 與申報版本保持一致。 ✓ 複核透過 版本、結果與合規狀態透過主辦方稽核,方計為一次有效提交。 參賽基本要求 1 只返回記憶證據 Search 不得直接生成最終答案,也不得把答案偽裝為記憶記錄。 2 保持樣本隔離 不得跨 user_id、任務、樣本或團隊共享和檢索評測記憶。 3 披露來源與改動 複用論文、倉庫或程式碼時,必須註明原作者、技術報告和全部方法改動。 4 不得操縱評測 嚴禁硬編碼、資料洩漏、提示詞注入、人工即時答題、結果操縱和惡意刷榜。 聯絡與入口 提交申請前請先閱讀 API 接入指南並準備完整材料。報名與評測問題可傳送至 [email protected]。 賽事倉庫 PARTICIPATION GUIDE · 2026 Agent Memory Challenge 2026 Participation Guide Agent Memory Challenge is the first public evaluation cycle of Agent Memory Leaderboard. It is open to researchers, open-source maintainers, and commercial product teams worldwide. Participants provide Add and Search; the platform runs Answer, Eval, result review, and leaderboard publication. Back to Competition The first cycle at a glance Participation is free, with no team-size requirement. Participants cover the cost of their own APIs, databases, bandwidth, and compute; the platform covers unified Answer, Eval, and evaluation orchestration. 01 Registration opens July 29, 2026 02 Submission deadline August 7, 2026 · 23:59 (UTC+8) 03 First release Mid-August 2026 04 Evaluation types Textual Memory and Coding Memory 05 Participant divisions Academic Methods and Commercial Products Step 1: Choose an evaluation type The evaluation type determines the tasks your system receives; the participant division determines where the result is listed. These are two independent dimensions. Evaluation typeWhat it evaluatesAvailable divisionsFirst-cycle date Textual MemoryFact recall, multi-hop integration, temporal understanding, memory governance, personalization, rule execution, safety, and privacy.Academic Methods / Commercial ProductsSubmissions close August 7 Coding MemoryRetrieving, filtering, and reusing debugging experience, development experience, and project context from historical engineering tasks.Academic Methods / Commercial ProductsSubmissions close August 7 Step 2: Choose a participation route Your first action is always to submit an evaluation request. An Eval Key is not a separate registration step; it is issued after approval to participants who host their own APIs. Academic · API Host your own Add / Search APIs Submit a public GitHub repository, a fixed version, Add / Search endpoints, authentication details, and run instructions. You operate the service and keep it stable; an Eval Key is issued after approval. Academic · Code Submit code for platform deployment Submit a public GitHub repository with Docker startup instructions, an Add / Search wrapper, and complete run documentation. The platform builds and evaluates it; no Eval Key is issued. Commercial · API Provide a stable product API Submit a fixed product version, Add / Search endpoints, authentication, and capacity details. Internal implementation may remain closed; an Eval Key is issued after approval and results enter the Commercial Products board. Open-source requirement Academic Methods entries must provide a public, verifiable GitHub repository and disclose the original method, authors, technical report, and all changes. Commercial Products entries need not be open source, but their product and API versions must be stable and verifiable. Step 3: Prepare your submission Freeze a clear evaluation version before applying. Once a formal Full evaluation is accepted, the version may not be replaced or withdrawn because of an unfavorable result. For everyone Common materials System name and version, contact details, organization or team, intended evaluation type, method or product description, information approved for public display, and complete submission notes. Academic Methods Code and reproducibility Public repository, README, Docker command, API entrypoint, dependencies, attribution of prior work, method changes, and run steps. Self-hosted entries must also provide public endpoints. Commercial Products API and operations Fixed product version, Add / Search APIs, authentication, a dedicated evaluation credential, capacity, timeout, and rate-limit details. Endpoints must remain stable and publicly reachable for at least 30 days after submission. From request to leaderboard 1 Submit a request Choose an evaluation type, participant division, and submission route, then provide a fixed system version and complete materials. 2 Complete integration Self-hosted API participants receive an Eval Key; code submissions are built by the platform from the documented Docker entrypoint. 3 Run Smoke and Full Validate Add / Search, authentication, and the end-to-end path before the formal first-cycle Full evaluation. 4 Review and publish The platform reviews the version, results, and compliance status before publishing the entry to its corresponding public board. Eval Key and Memory System Key These credentials are issued by different parties and serve different purposes. Never place either credential in a public repository, URL, screenshot, email body, or group chat. CredentialProvided byPurposeWho needs it Eval Key / Leaderboard KeyAgent Memory LeaderboardVerifies participant access, starts evaluations, a [truncated for AI cost control]