待翻译:Hybrid-retrieval memory layer for AI agents
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Notifications You must be signed in to change notification settings Fork 1 Star 10 BranchesTags Open more actions menu Folders and files NameName Last commit message Last commit date Latest commit History 2 Commits 2 Co…
AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。
Notifications You must be signed in to change notification settings Fork 1 Star 10 BranchesTags Open more actions menu Folders and files NameName Last commit message Last commit date Latest commit History 2 Commits 2 Commits benchmarks benchmarks deepmem deepmem scripts scripts server server tests tests .dockerignore .dockerignore .env.example .env.example .gitattributes .gitattributes .gitignore .gitignore Dockerfile Dockerfile LICENSE LICENSE README.md README.md config.json.example config.json.example demo.gif demo.gif docker-compose.yml docker-compose.yml pytest.ini pytest.ini requirements-dev.txt requirements-dev.txt requirements.txt requirements.txt Repository files navigation Drop-in AI memory layer with 2× faster response and 10× lower cost. Fully compatible with Mem0 API. Migrate in 5 minutes - one import line. Self-hostable. No auth, no payment, no lock-in. Or use the managed cloud at deepmem.dev. Cloud · Self-host · Benchmarks · Reproduce them Migrate from Mem0 in one line - same MemoryClient, same method signatures: # Before - Mem0 from mem0 import MemoryClient client = MemoryClient(api_key="m0-...") # After - DeepMem (only the import changes) from deepmem import MemoryClient client = MemoryClient(api_key="dm_live-...") # get a key at deepmem.dev pip install deepmem-client Turn conversations into searchable long-term memory: a FastAPI HTTP API in front of a Qdrant vector store, with LLM fact extraction, hybrid retrieval (vector + BM25 + entity boost + time-decay), semantic caching, async batched distillation, GDPR controls, and a built-in MCP server. It runs in open mode - no API key, no user registration - so you can deploy it for your own agents in minutes. Multi-tenant isolation is driven by user_id in the request body. Prefer not to self-host? DeepMem Cloud is the managed version of this exact engine at deepmem.dev - same API, no infra. Sign up, grab a key (dm_live_...), point your base URL at https://deepmem.dev, done. The cloud and the open-source server speak the same Mem0-compatible API, so client code is identical. Migrate from Mem0 Already using Mem0? Switch to DeepMem cloud in one line. The deepmem-client package mirrors mem0.MemoryClient - same class name, same method signatures, same filters={"user_id": ...} style - so everything after the import stays untouched. pip install deepmem-client # before (Mem0) from mem0 import MemoryClient client = MemoryClient(api_key="m0-...") client.add(messages, user_id="alex") client.search("What can Alex cook?", filters={"user_id": "alex"}) # after (DeepMem cloud) - change one import line from deepmem import MemoryClient client = MemoryClient(api_key="dm_live_...") # key at https://deepmem.dev client.add(messages, user_id="alex") # identical calls client.search("What can Alex cook?", filters={"user_id": "alex"}) Behavioral notes add(infer=True) (the default) is asynchronous on DeepMem cloud - it returns pending=True with results=[] and extracted facts land a few seconds later. (Mem0 cloud's add is async too - it returns PENDING.) Pass infer=False for synchronous raw-text storage that's immediately searchable. No graph relations - DeepMem uses hybrid vector retrieval (vector + BM25 time-decay), so relations is always []. Mem0's graph features aren't replicated. reset differs - Mem0's is account-wide; DeepMem's is per-user_id with a confirm guard. Pricing DeepMem Cloud is 10x cheaper than Mem0 at every paid tier - the same shape of plans, a tenth of the price. Tier DeepMem Mem0 cloud Hobby Free Free Starter $1.9/mo $19/mo Growth $7.9/mo $79/mo Professional $24.9/mo $249/mo Self-host instead and it's $0 - you pay only your own LLM/embedding provider (the same LLM cost Mem0 charges on top of its plan price), with no memory-service markup. Batched distillation also cuts LLM calls ~80%, so even your provider bill is smaller than per-message extractors. Plans and limits: deepmem.dev · mem0.ai. Benchmarks No cherry-picked headline. The scripts and workload ship in /benchmarks - run them yourself. Here's what we measured and the exact config that produced it: Metric DeepMem self-hosted ¹ DeepMem cloud Mem0 cloud Search p50 73 ms 643 ms 653 ms Search p95 86 ms 811 ms 710 ms Search hits (40 queries) - 195 84 Add p50 (raw store) 899 ms ² 792 ms 695 ms ³ ¹ BGE-M3 on a GTX 1070 GPU (2016-era), local file Qdrant, infer=False, 100 ops, concurrency 1. ² Dominated by local-file Qdrant I/O - a Qdrant server cuts this sharply. ³ Mem0 has no raw-store mode; add always runs LLM extraction, so this row isn't apples-to-apples. Self-hosted is where intrinsic latency lives - no internet RTT, your embedder, your Qdrant. 73 ms p50 search on an old consumer GPU. DeepMem cloud beats Mem0 cloud on search p50 (643 ms vs 653 ms) and returns ~2.3x more candidates per search (195 vs 84 hits across 40 queries). Cloud latency is RTT-dominated - both cloud columns were measured through a proxy from mainland China; run-to-run jitter is ~±10%. Run /benchmarks from a low-RTT location for your own numbers. Why DeepMem Agent frameworks keep re-discovering that they need persistent, retrievable memory. The hosted options bill per call and send your data to someone else's cloud. DeepMem is the self-hostable alternative: the same Mem0-shaped API you can drop in, but it runs on your box, with your embedder, your LLM key, and your Qdrant - and the code is right here to verify it. How does DeepMem compare to other Mem0 alternatives? Most are hosted-only or layer memory on top of someone else's vector DB. DeepMem combines three things at once: it's self-hostable (your data stays on your box - $0 beyond your own LLM key), MCP-native (Claude Desktop / Cursor read and write memories directly as tools), and fully open-source - and the managed cloud runs the exact same engine, so cloud and self-host are one API, not two products. Without DeepMem With DeepMem Re-explain who you are and what you're working on every session The agent recalls identity, projects, and preferences automatically Lose debugging and research context between sessions Past root causes, dead ends, and findings are recalled, so work isn't repeated Manually restate preferences every session Preferences persist across sessions, agents, and projects Hosted memory services that bill per call and hold your data Self-host on your infra, or use the cloud - your call, same API What it is (honestly) Hybrid retrieval, not a knowledge graph. Search fuses vector similarity, BM25 keyword match, entity boost, and time-decay scoring. There is no temporal graph layer; if that's what you need, look at Zep. Stores preferences, not code dumps. Large fenced code blocks are stripped before LLM extraction, so the store fills with durable user/project facts instead of pasted implementations. BYOK, multi-provider. Bring your own LLM (OpenAI / Anthropic / any OpenAI-compatible endpoint) and embedding (BGE-M3 / Google / OpenAI-compatible). MCP-native. Ships an MCP server so Claude Desktop / Cursor can read and write memories directly. Quick start Three ways to run. All speak the same Mem0-compatible API. 1. Cloud (zero ops) export DEEPMEM_API_KEY=dm_live_... # from https://deepmem.dev curl https://deepmem.dev/v1/memories \ -H "Authorization: Bearer $DEEPMEM_API_KEY" -H "Content-Type: application/json" \ -d '{"messages":[{"role":"user","content":"I am Pat, I live in Lisbon."}],"user_id":"pat","infer":false}' curl https://deepmem.dev/v1/memories/search \ -H "Authorization: Bearer $DEEPMEM_API_KEY" -H "Content-Type: application/json" \ -d '{"query":"Where does Pat live?","user_id":"pat"}' 2. Docker (one command) cp .env.example .env # add an LLM key docker compose up --build # DeepMem (HTTP :8000 + MCP :8001) + Qdrant sidecar curl http://localhost:8000/health Or pull the published image: docker pull langdeepmem/deepmem:latest docker run -p 8000:8000 -p 8001:8001 -e DEEPSEEK_API_KEY=sk-... langdeepmem/deepmem:latest The image exposes :8000 (HTTP) and :8001 (MCP). The Dockerfile and docker-compose.yml cover the GPU variant (CUDA torch + BGE_DEVICE=cuda) and BGE-M3 model-download options (HF mirror, proxy, or local mount). 3. From source git clone https://github.com/deepmemteam/deepmem.git && cd deepmem pip install -r requirements.txt cp .env.example .env # add an LLM key + embedder config python server/start.py # HTTP :8000 + MCP :8001 Write and search in three lines: import httpx httpx.post("http://localhost:8000/v1/memories", json={"messages":[{"role":"user","content":"I'm Pat, I live in Lisbon."}], "user_id":"pat"}) print(httpx.post("http://localhost:8000/v1/memories/search", json={"query":"Where does Pat live?","user_id":"pat"}).json()["results"]) user_id is optional (defaults to "default"); send different user_ids to isolate end-users. infer: false stores raw text immediately (test-friendly); the default infer: true queues for LLM fact extraction. Feature highlights Multi-provider by config, not code. Both layers switch on env vars: Layer Options Selector LLM (fact extraction) OpenAI · Anthropic (native SDK) · any OpenAI-compatible (DeepSeek / vLLM / Ollama / Groq / LM Studio) LLM_PROVIDER + LLM_API_KEY / ANTHROPIC_API_KEY / DEEPSEEK_API_KEY Embeddings BGE-M3 (local, GPU/CPU) · Google Gemini · any OpenAI-compatible EMBEDDING_PROVIDER + BGE_M3_PATH / GOOGLE_API_KEY / OPENAI_API_KEY BGE_DEVICE=auto|cpu|cuda picks GPU when available, else CPU (force cpu on small-VRAM cards to avoid multi-process contention). BYOK overrides the LLM per-request. Hybrid retrieval. Vector similarity + BM25 keyword + entity boost + time-decay, fused into one score. Over-fetch, re-rank, return. Async batched distillation. Writes queue behind a silence window and are extracted in batches - ~80% fewer LLM calls than per-message extraction. Semantic cache. Repeat adds/searches hit a similarity-gated cache and return cached facts without re-embedding or re-querying Qdrant. Stores preferences, not code. extraction_filter strips large fenced code blocks before LLM extraction, so the store fills with durable facts, not pasted implementations. MCP server. deepmem_write / deepmem_search / deepmem_delete tools for Claude Desktop, Cursor, and any MCP client. GDPR. Soft-delete with retention window, hard-delete reset, SHA-256 id masking in logs, export/import for portability. API Method Path Description POST /v1/memories Write messages; LLM-extract facts (infer=false stores raw) POST /v1/memories/search Semantic search (vector + BM25 + entity + time-decay) GET /v1/memories List all for a user_id (paginated) GET /v1/memories/{id} Get one by ID PUT /v1/memories/{id} Update one memory's text DELETE /v1/memories/{id} Soft-delete one DELETE /v1/memories Soft-delete all for a user_id (GDPR) GET /v1/memories/{id}/history ADD/UPDATE/DELETE audit log POST /v1/reset Hard-delete all + history (needs confirm_user_id) GET /v1/export · POST /v1/import Portable JSON export / import GET /health · /ready Liveness / readiness probes agent_id / run_id optionally scope writes/reads (mirrors Mem0's three-level isolation: user -> agent -> run). Interactive docs at /docs. MCP integration { "mcpServers": { "deepmem": { "command": "python", "args": ["server/mcp_server.py"], "env": { "DEEPMEMORY_BASE_URL": "http://localhost:8000" } } } } Open mode needs no API key - DEEPMEMORY_API_KEY stays empty. Architecture client (HTTP / MCP) -> FastAPI (:8000) -> rate-limit middleware (per-IP token bucket on add/search) -> SemanticCache.check (return cached facts on similarity hit) -> AsyncBatchDistiller.enqueue (POST /v1/memories, infer=true) ↳ silence-window or max_batch triggers on_batch_ready ↳ VectorStore.process_batch (LLM extraction -> Qdrant upsert) -> VectorStore.search (vector + BM25 + entity + time-d [truncated for AI cost control]