AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。
Every tool you've ever set up has an onboarding screen you just straight click past. Usually the defaults are chosen by someone who thought about them harder than you have time to. Engram's setup has one of those screens too, and it takes a couple of minutes to get through: name a project → pick a template → click past topics that are already filled in for you → generate an API key. note Engram is our fully managed memory and context service purpose-built to help agents remember, learn, and improve over time. Then you call memories.add with your raw conversation data, memories.search before your model call, and it works. It keeps working, too. That's the interesting part, because nothing ever nudges you back to that screen. But somewhere between "it works" and "it works the way I meant it to", four decisions turn out to be yours rather than Engram's: What gets remembered (the topic descriptions) How many memories, and whose (bounded topics and scope) What comes back (the retrieval mode) Where it goes (placement in your prompt) In this post we walk through all four, with real outputs from a project mostly running the stock personalization template. How Engram turns raw data into memories Engram runs an asynchronous pipeline over whatever raw data you send it. By default, it: extracts the memories that matter, transforms them against what's already stored, commits the result. Victoria's walkthrough video covers that end-to-end, including the console setup and the SDK: memories.add hands back a run, not a memory, and nothing you send is searchable until that run finishes. The run status guide covers the states, and why you usually shouldn't wait on one. For how the pipeline itself works, the Engram: Memory by Weaviate blog goes through it step by step. Which memories it pulls out of the data you send, though, is decided by your topic descriptions, so that is where we start. How topic descriptions decide what gets remembered When you create a new Engram project in the console, the User Personalization template sets up UserProfile and UserKnowledge for you, and offers ConversationSummary as an optional third. I took all three, then sent it something about me: "I'm Prajjwal, a developer advocate at Weaviate. I write all my demos in Python, and I have a hard rule that they stay under 100 lines." Five memories came back across the three topics: [UserProfile] The user's name is Prajjwal. [UserKnowledge] Prajjwal works as a developer advocate at Weaviate. [UserKnowledge] Prajjwal writes all his demos in Python. [UserKnowledge] Prajjwal enforces a hard rule that his demos stay under 100 lines. [ConversationSummary] Prajjwal introduced himself as a developer advocate at Weaviate, mentioned that he writes demos in Python, and follows a rule to keep demos under 100 lines. note Other ready-to-use project templates exist as well, like Coding Assistant and Personal Claw Agent. And you can always start from a blank slate and fully customise the topics to your domain. UserKnowledge broke that one sentence into three atomic facts, each retrievable on its own. ConversationSummary kept it whole as a flowing narrative. Same input, same pipeline, same moment, and two completely different shapes, because the two topics describe themselves differently. We didn’t write any routing logic or a formatter. The descriptions did both, because a topic description is the memory extraction prompt, and it controls three things: What to leave out UserKnowledge is the template's catch-all topic, and anything personal about the user that might change over time belongs there. That makes it useful when you want broad coverage. But when you want a focused, niche memory, the description needs to do more than say what belongs - it also needs to define what doesn't. Take a throwaway message like: "Two hours lost to a Docker rebuild, my headphones died mid-call, and it started raining right as I stepped out. Anyway, I finally swapped the demo over to qwen3-embedding-8b." Four memories came back. Three were weather, hardware, and other passing events. The agent now knows it rained on Thursday, and short of a delete call, that fact can stick around indefinitely. To fix this, the broad topic can be replaced with a focused one by just specifying what you don't want to be remembered. I created a new topic called UserFacts to replace UserKnowledge and its description ends with a new rule: "Do not record events, incidents, or passing conditions." That one line made the difference as it gives the model general criteria for what to ignore. I tested it with three new messages: a late train, a cat on the keyboard, and a stolen lunch, each paired with one lasting decision. Across 24 runs, none of the noisy incidents reached memories created by UserFacts, and the decisions were captured in all of them. So, for a focused topic, a rule about what to exclude does more work than a list of what to include. Name the kind of thing you want left out, rather than every example you can think of, and the rule can generalise to cases you never anticipated. Also, UserFacts was created only for this experiment and from here on, I switched back to the stock UserKnowledge topic. What shape to write it in You're almost certainly going to read your memories into a prompt, so you should describe the form you want, not just the subject. For example, "Two or three sentences of plain prose, second person, no headings or bullet points" gives you something you can drop into context untouched. Or ask for atomic facts instead, and you can get rows you can retrieve one at a time. When something is worth remembering A description is applied to one input at a time, alongside whatever related memories the pipeline pulls in for it. That makes it good at judgements it can settle on the spot, but it cannot help with anything that depends on what came before. Accumulating is a separate step. A pipeline chains extract, transform and commit step by default, and a fourth kind of step, a buffer, can sit anywhere among them. It holds memories or raw inputs back until a trigger fires which could be a count, or a timer. A rule about something building up over time belongs there, not in the wording of your description. note A default project runs extract, transform and commit. Adding a buffer, or reordering the steps around it, is configured per project and is currently available on enterprise plans. Also, topics are editable in the console at any time, so they can be added, removed, or reworded without recreating the project. This makes it easy to iterate until they work as expected. The other config fields Description is one field. The rest are configuration rather than wording, and they carry their own effects. This is the form you see when you add a topic of your own in the console: FieldWhat it decidesIf you get it wrong Namehow you address the topic in code, topics=["UserProfile"]- Descriptionthe extraction prompt: what gets pulled out, and how it's writtenthe topic keeps everything, or nothing you can use User scopedmemories belong to one user_iduncheck it and every user shares one pool Property scopesextra partition keys like conversation_id or repoevery write must carry them, or nothing reaches the topic Boundedat most one memory per scopeleave it off and nothing guarantees a single memory to read back The topics docs cover all these concepts in more depth. What bounded topics and scope guarantee Bounded caps how many memories a topic may hold. Scope decides which of them a given read or write request can reach. Both are about what happens when a new message arrives for a topic that already holds memories. Continuing the 100-line demo example from earlier, let's say four days later I send: "Update: the demo is 400 lines now. The 100-line rule is officially dead." Zero memories created, two updated. UserKnowledge memories afterwards: [UserKnowledge] Prajjwal works as a developer advocate at Weaviate. [UserKnowledge] Prajjwal writes all his demos in Python. [UserKnowledge] Prajjwal no longer enforces his previous hard rule of keeping demos under 100 lines; as of 5 Sep 2026 his demos are 400 lines long. The memory about the 100-line rule was rewritten in place and the rule is gone. The other two were left alone, since nothing in the new message contradicted them. That is the transform step from the Engram pipeline doing its job, and every topic gets it, bounded or not. Note that deleted was zero, as reconciliation generally supersedes rather than erases. If you want a memory gone explicitly, you delete it with memories.delete(). Bounded What bounded adds is a promise about the count. A bounded topic holds at most one memory per unique scope. Engram derives the memory's ID from the topic name and the scope, so every later write lands on that same ID and updates it instead of adding another. The update above was one run. Send the introduction and then the update message to five fresh users, each starting from an empty store, and count how many memories land: run 0 run 1 run 2 run 3 run 4 UserKnowledge (unbounded) 2 4 4 3 3 UserProfile (bounded) 1 1 1 1 1 The unbounded topic landed anywhere between two and four memories: sometimes the retraction became one memory, sometimes two, and so on. UserProfile always held exactly one memory, five times out of five, because it is bounded. So if your code needs to read a single standing memory for a topic, you should always bound the topic rather than trusting the count. The template bounds ConversationSummary for the same reason. A conversation should have one running summary that gets rewritten, not a new one per message. Scope Scope partitions a topic, and there are two kinds. User scope is a hard wall as every write and every read has to carry a user_id. Property scopes like conversation_id or repo are required on writes but optional on reads, so you can read one partition or all of them at once. That optionality is the reason to use a property. The template scopes ConversationSummary by conversation_id, so each thread keeps its own summary and you can still ask for every summary a user has. If two partitions are never meant to be read together, don't use a property and give them separate user_ids instead. And if a fact should follow the user everywhere, leave it at user scope and add nothing. A write has to carry every key the topic declares, or it's rejected like this: insufficient scope: missing required scope properties [conversation_id] to write memories insufficient scope: missing required user_id to write memories invalid scope property: [repo] not configured on any topic (configured properties: [conversation_id]) An empty string counts as missing, and a property no topic declares gets an error naming the ones the project has. Reads only insist on user_id. An empty conversation_id is treated as absent and searches every conversation, even though a write would have rejected it, and a conversation_id that was never written simply returns nothing from the conversation-scoped topic. How to read memories back There are four retrieval modes. vector, bm25 and hybrid all are for search: you give them a query, they score every memory against it, and you get the best ones back in ranked order. hybrid is what runs if you don't pick one, while fetch is designed for direct, non-ranked memory retrieval. Search search ranks by relevance to a query. In a chat app, that query is usually the user's current message: from engram import HybridRetrieval hits = client.memories.search( query=user_message, retrieval_config=HybridRetrieval(limit=3), user_id=uid, properties=props, ) context = "\n".join(f"- {m.content}" for m in hits) Set limit to the number of memories you actually want in the prompt. The default is ten, and you get ten whether or not the tenth has anything to do with the question. Ask for too many, and yo [truncated for AI cost control]