AI News HubLIVE
In-site rewrite7 min read

Show HN: Agent Memory Leaderboard – first public results for AI memory systems

OPEN BENCHMARK Agent Memory Leaderboard A public benchmark space for comparing textual and coding-agent memory systems under a consistent evaluation flow. Leaderboard Preview Benchmark Tracks Each track keeps its own re…

SourceHacker News AIAuthor: IreneAI

OPEN BENCHMARK Agent Memory Leaderboard A public benchmark space for comparing textual and coding-agent memory systems under a consistent evaluation flow. Leaderboard Preview Benchmark Tracks Each track keeps its own result table and detailed metric breakdown. Textual Memory Long-context, persona, script, and conversation-memory benchmarks. Coding Agent Memory Agent memory support for coding tasks and repository-context recall. Evaluation Flow Industry systems use the hosted Add/Search key flow. Academic systems may use the same flow or submit a public GitHub repository for maintainer Docker deployment. 1 Choose an evaluation route Provide hosted Add/Search APIs, or submit a public GitHub repository with Docker and API run instructions. 2 Run a smoke test Use the issued key to verify the synchronous Add/Search flow. 3 Submit a formal evaluation After smoke passes, submit the full scored evaluation. Explore the Platform Use the product pages to inspect rankings, run evaluations, and prepare an integration. 01 Leaderboard Public ranking with filters, dataset columns, and score bars. → 02 Evaluation Create eval jobs, watch progress, and inspect private results. → 03 Participation Guide Eligibility, submission routes, required materials, timelines, rewards, and publication rules. → 04 Documentation User guide, evaluation workflow, API contract, security, and result publication. → 05 Guide Add/search API contract, request fields, polling, and response schemas. → PUBLIC RANKINGS Agent Memory Leaderboard Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions. Leaderboard A First-cycle evaluation results will be released in mid-August. Academic board submissions are open now. Submit your evaluation request before the first-cycle deadline. PUBLIC RANKINGS Agent Memory Leaderboard Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions. Leaderboard A First-cycle evaluation results will be released in mid-August. Academic board submissions are open now. Submit your evaluation request before the first-cycle deadline. EVALUATION Run Evaluations Create API-gated eval jobs against the Leaderboard Suite, monitor task progress, inspect private results, and submit eligible full-suite runs for administrator review. Create Eval Job Choose a bound version. Use Run label to distinguish repeated evaluations of the same version. System name Version name Mode Max add concurrency 16-64 Search concurrency 16-256 Top K Run label Continue the latest interrupted evaluation from its last compatible checkpoint. Datasets FULL EVALUATION GATE Full evaluation checklist Confirm each item before starting a public-board candidate run. The button remains locked until every item is checked. 0 / 8 confirmed Smoke test completedThe Add/Search API has passed the platform smoke test and is currently usable. API contract is followedThe submitted code and deployed interfaces are wrapped according to the official Add/Search format. Add/Search uses gpt-4o-miniThe model used by the submitted memory system during both Add and Search must be gpt-4o-mini. The platform will reproduce the submission; if the reproduced score differs materially, the leaderboard result may be invalidated. Runtime will stay stableIf you provide a deployed endpoint, it will remain publicly reachable and stable for at least 30 days after submission. Run instructions are completeThe repository README or submission notes include the Docker command, API entrypoint, configuration, and startup steps. Original work is disclosedAny reused paper, repository, or code is attributed with its original authors, technical report, and the changes made; personal work is identified as such. This is a substantive submissionIt is not a repeated near-duplicate or a low-quality submission intended to occupy evaluation capacity. No manipulation or cheatingThe system does not use prompt injection, benchmark leakage, result manipulation, malicious behavior, or other leaderboard abuse. Current Job Live job status and progress. No job yet. External Full-Run Review Review completed full-suite runs from non-admin leaderboard keys. Approved results enter the public board; rejected results remain private. Private Results Private scores stay scoped to the current leaderboard key. Admin runs remain separate from external review candidates. PARTICIPATION GUIDE · 2026 首届 Agent 记忆挑战赛参赛说明 Agent Memory Challenge 是 Agent Memory Leaderboard 的首期公开评测活动,面向全球研究者、开源项目维护者和商业产品团队开放。参赛系统负责 Add 与 Search,平台统一完成 Answer、Eval、结果复核与公榜。 返回赛事页面 一分钟了解首届赛事 参赛免费,不设组队要求。参赛方承担自身 API、数据库、带宽和计算成本,平台承担统一 Answer、Eval 与评测编排成本。 01 报名开放 2026 年 7 月 29 日 02 提交截止 2026 年 8 月 7 日 23:59(UTC+8) 03 首次放榜 2026 年 8 月中旬 04 评测类型 文本记忆与代码记忆 05 参赛组别 学术方法榜与商业产品榜 第一步:选择评测类型 评测类型决定系统接受什么任务;参赛组别决定结果展示在哪个榜单。两者是互相独立的两个维度。 评测类型主要评测内容可选参赛组别首届时间 文本记忆事实召回、多跳整合、时序理解、记忆治理、个性化、规则执行、安全与隐私。学术方法榜 / 商业产品榜8 月 7 日提交截止 代码记忆从历史工程任务中检索、筛选并复用调试经验、开发经验和项目上下文。学术方法榜 / 商业产品榜8 月 7 日提交截止 第二步:选择参赛路径 第一步始终是提交评测申请。Eval Key 不是另一个报名入口,而是自行部署 API 的申请审核通过后获得的评测凭证。 学术 · API 自行部署 Add / Search API 提交公开 GitHub 仓库、固定版本、Add / Search 地址、鉴权方式和运行说明。参赛方负责部署并保持接口稳定,审核通过后获得 Eval Key。 学术 · 代码 提交代码,由平台部署 提交公开 GitHub 仓库、Docker 启动方式、Add / Search 封装和完整运行说明。平台负责构建与评测,不签发 Eval Key。 商业 · API 提供稳定的产品 API 提交固定产品版本、Add / Search 地址、鉴权和容量说明。无需公开内部实现,审核通过后获得 Eval Key,结果进入商业产品榜。 开源要求 学术方法榜必须提供公开、可核验的 GitHub 仓库,并披露原始方法、作者、技术报告和本次改动;商业产品榜无需开源,但必须提供可核验且稳定的产品与 API 版本。 第三步:准备提交材料 请在提交申请前固定参评版本。正式 Full 评测受理后,不得因结果不理想更换版本或撤回。 共同材料 所有参赛者 系统名称与版本、联系人、机构或团队、拟参评类型、方法或产品说明、允许公开展示的信息,以及完整的提交说明。 学术方法 代码与复现材料 公开仓库、README、Docker 命令、API 入口、依赖配置、原始工作引用、方法改动和运行步骤。自行部署时还需提供公网 API。 商业产品 接口与运行材料 固定产品版本、Add / Search API、鉴权方式、评测专用密钥、容量限制、超时与限流说明,并保证接口在提交后至少 30 天稳定可访问。 从申请到上榜 1 提交评测申请 选择评测类型、参赛组别和提交方式,提交系统版本与完整材料。 2 完成接入 自行部署 API 的参赛者获取 Eval Key;代码提交由平台按 Docker 说明构建。 3 Smoke 与 Full 先验证 Add / Search、鉴权和端到端链路,再运行首届正式 Full 评测。 4 复核与公榜 平台复核版本、结果和合规状态,通过后发布到对应公开榜单。 Eval Key 与 Memory System Key 两个 Key 的签发方和用途不同,请勿将它们放入公开仓库、URL、截图、邮件正文或群聊。 凭证谁提供用途谁需要 Eval Key / Leaderboard KeyAgent Memory Leaderboard验证参赛身份、运行评测并查看私有结果。自行部署 API 的学术与商业参赛者;平台部署代码的路径不签发。 Memory System Key参赛方供平台访问参赛系统的 Add / Search API。接口启用鉴权时需要;无鉴权接口无需提供。 奖励设置 榜单排名奖励与社区贡献奖励是两套独立计划;进入公开榜单不等于自动获得全部奖励。 学术排名奖励 开源方法榜 Top 10 第 1—3 名获得 ChatGPT Pro 月度会员;第 4—10 名获得 ChatGPT Plus 月度会员。账号由赛事方独立采购并发放,赛事并非由相关产品提供方主办或赞助。 社区贡献计划 满足任一条件即可入选 前 50 位完成有效提交;成功邀请 3 位新参赛者完成有效提交;或提交 3 组有挑战性的测试样本并通过审核。 Kimi Token 入选即进入奖励名单 入选社区贡献计划后,获得至少价值人民币 50 元的 Kimi Token 额度。最终名额、额度、发放时间和审核结果以官方通知为准。 什么是有效提交? ✓ 材料完整 参赛身份、系统版本和所需材料均完整、真实且可核验。 ✓ Smoke 通过 Add / Search、鉴权和端到端链路符合现行接入协议。 ✓ Full 完成 正式评测任务和所需评测项成功完成,没有缺失或重复结果。 ✓ 版本一致 正式评测使用的代码、镜像或 API 与申报版本保持一致。 ✓ 复核通过 版本、结果与合规状态通过主办方审核,方计为一次有效提交。 参赛基本要求 1 只返回记忆证据 Search 不得直接生成最终答案,也不得把答案伪装为记忆记录。 2 保持样本隔离 不得跨 user_id、任务、样本或团队共享和检索评测记忆。 3 披露来源与改动 复用论文、仓库或代码时,必须注明原作者、技术报告和全部方法改动。 4 不得操纵评测 严禁硬编码、数据泄漏、提示词注入、人工实时答题、结果操纵和恶意刷榜。 联系与入口 提交申请前请先阅读 API 接入指南并准备完整材料。报名与评测问题可发送至 [email protected]。 赛事仓库 PARTICIPATION GUIDE · 2026 Agent Memory Challenge 2026 Participation Guide Agent Memory Challenge is the first public evaluation cycle of Agent Memory Leaderboard. It is open to researchers, open-source maintainers, and commercial product teams worldwide. Participants provide Add and Search; the platform runs Answer, Eval, result review, and leaderboard publication. Back to Competition The first cycle at a glance Participation is free, with no team-size requirement. Participants cover the cost of their own APIs, databases, bandwidth, and compute; the platform covers unified Answer, Eval, and evaluation orchestration. 01 Registration opens July 29, 2026 02 Submission deadline August 7, 2026 · 23:59 (UTC+8) 03 First release Mid-August 2026 04 Evaluation types Textual Memory and Coding Memory 05 Participant divisions Academic Methods and Commercial Products Step 1: Choose an evaluation type The evaluation type determines the tasks your system receives; the participant division determines where the result is listed. These are two independent dimensions. Evaluation typeWhat it evaluatesAvailable divisionsFirst-cycle date Textual MemoryFact recall, multi-hop integration, temporal understanding, memory governance, personalization, rule execution, safety, and privacy.Academic Methods / Commercial ProductsSubmissions close August 7 Coding MemoryRetrieving, filtering, and reusing debugging experience, development experience, and project context from historical engineering tasks.Academic Methods / Commercial ProductsSubmissions close August 7 Step 2: Choose a participation route Your first action is always to submit an evaluation request. An Eval Key is not a separate registration step; it is issued after approval to participants who host their own APIs. Academic · API Host your own Add / Search APIs Submit a public GitHub repository, a fixed version, Add / Search endpoints, authentication details, and run instructions. You operate the service and keep it stable; an Eval Key is issued after approval. Academic · Code Submit code for platform deployment Submit a public GitHub repository with Docker startup instructions, an Add / Search wrapper, and complete run documentation. The platform builds and evaluates it; no Eval Key is issued. Commercial · API Provide a stable product API Submit a fixed product version, Add / Search endpoints, authentication, and capacity details. Internal implementation may remain closed; an Eval Key is issued after approval and results enter the Commercial Products board. Open-source requirement Academic Methods entries must provide a public, verifiable GitHub repository and disclose the original method, authors, technical report, and all changes. Commercial Products entries need not be open source, but their product and API versions must be stable and verifiable. Step 3: Prepare your submission Freeze a clear evaluation version before applying. Once a formal Full evaluation is accepted, the version may not be replaced or withdrawn because of an unfavorable result. For everyone Common materials System name and version, contact details, organization or team, intended evaluation type, method or product description, information approved for public display, and complete submission notes. Academic Methods Code and reproducibility Public repository, README, Docker command, API entrypoint, dependencies, attribution of prior work, method changes, and run steps. Self-hosted entries must also provide public endpoints. Commercial Products API and operations Fixed product version, Add / Search APIs, authentication, a dedicated evaluation credential, capacity, timeout, and rate-limit details. Endpoints must remain stable and publicly reachable for at least 30 days after submission. From request to leaderboard 1 Submit a request Choose an evaluation type, participant division, and submission route, then provide a fixed system version and complete materials. 2 Complete integration Self-hosted API participants receive an Eval Key; code submissions are built by the platform from the documented Docker entrypoint. 3 Run Smoke and Full Validate Add / Search, authentication, and the end-to-end path before the formal first-cycle Full evaluation. 4 Review and publish The platform reviews the version, results, and compliance status before publishing the entry to its corresponding public board. Eval Key and Memory System Key These credentials are issued by different parties and serve different purposes. Never place either credential in a public repository, URL, screenshot, email body, or group chat. CredentialProvided byPurposeWho needs it Eval Key / Leaderboard KeyAgent Memory LeaderboardVerifies participant access, starts evaluations, a [truncated for AI cost control]