AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:Show HN: When AI Decides What Matters

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Disclosure: These views are my own and do not represent my current or any former employers. Executive summary Email, calendar, meeting, and notification assistants are increasingly presented as a way to begin the day wi…

來源Hacker News AI作者: thepsychguy

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Disclosure: These views are my own and do not represent my current or any former employers. Executive summary Email, calendar, meeting, and notification assistants are increasingly presented as a way to begin the day with the right information already selected. They summarise long threads, identify action items, rank messages, group notifications, and prepare short accounts of what needs attention. This sounds like compression. It is also allocation. The system decides which inputs appear, how they are described, and what remains outside the user's immediate view. It influences attention before the user makes a decision. The distinction matters because errors are not equally visible. A false statement in a summary may be noticed and challenged. An important message that was not surfaced can leave no obvious sign that anything went wrong. The user sees a plausible account of the day without seeing the material that the account omitted. Microsoft describes Outlook's Prioritize feature as a system that selects important mail using past behaviour, organisational context, user instructions, and other signals. It leans towards messages that require action. It also excludes several categories of mail from evaluation, including messages delivered outside the Inbox, meeting mail, some encrypted messages, and mail marked as low importance by the sender.[1][2] Microsoft separately states that an email summarisation feature may "overlook important details or misinterpret the context" and advises users to check the original message. The same document says that precision, recall, and completeness are evaluated, but it does not publish the results.[3] The public record contains enough evidence to treat these limitations seriously: Apple Intelligence produced false notification summaries that appeared to come from established news organisations. Apple later paused summaries for news and entertainment applications while it worked on improvements.[4][5] A Microsoft 365 Copilot defect caused emails labelled confidential to be processed despite configured data loss prevention rules. Microsoft said the issue did not grant access to anyone who lacked permission, but acknowledged that the behaviour did not match the intended protection.[6] Security researchers demonstrated that email and calendar content could contain instructions that altered an assistant's output or caused private information to be copied into another record. These were controlled demonstrations, not confirmed attacks against customers.[7][8][9] A meeting transcription service reportedly sent an external participant a transcript that included private conversation held after the participant believed the meeting had ended.[10] A Canadian tribunal held Air Canada responsible when a customer relied on incorrect information supplied by its website chatbot. The case did not concern an inbox assistant, but it illustrated a useful principle: an organisation remains responsible for information supplied through its automated systems.[11] None of these cases proves that an inbox prioritisation feature has caused a missed filing, lost contract, or personal emergency. I found no verified public case that makes that direct causal link. That gap should be stated plainly. It should not be mistaken for evidence that the risk is absent. Mailbox failures are private, omissions are difficult to observe, products are relatively new, and I found no public incident register or published false-negative rate for the items these products classify as important. The proposal in this paper is not to reject AI assistance. It is to treat systems that allocate attention as decision support rather than neutral convenience: Measure whether important items are missed alongside whether users report saving time. Show what sources were evaluated, what was excluded, and which product and configuration produced the result. Keep legal, safety, financial, security, and other critical obligations on channels that do not depend on probabilistic ranking. Test assistants with novel, ambiguous, adversarial, and role-specific messages before relying on them. Require vendors and deploying organisations to define responsibility for omissions, changes, and incidents. AI assistants do not merely summarise information. They govern attention. Their most consequential failure may be the item they quietly leave out. 1. From information tool to attention system A traditional search tool responds when a person asks a question. A traditional mail rule follows an explicit condition. A person can usually inspect the query or rule and understand why a result appeared. An AI assistant can operate earlier in the process. It can decide which messages are important, which parts of a thread deserve inclusion, which meeting actions belong to the user, and which notifications should appear first. The result may be waiting before the user has looked at the underlying material. That changes the role of the software. The assistant is no longer only helping a user work with information. It is constructing the user's first view of that information. It becomes an attention system. An attention system performs at least four functions: Function Apparent user benefit Judgement being made Selection Removes noise Which inputs deserve consideration Ranking Shows important items first Which inputs have greater consequence Compression Reduces reading time Which facts can be removed Presentation Creates a coherent briefing Which interpretation should frame the day Each function can be useful. Each can also fail independently. A system may select the right thread but summarise it incorrectly. It may summarise a message accurately but rank it too low. It may identify a deadline but assign the action to the wrong person. It may process every visible message correctly while excluding a folder that contains the only legal notice. The finished briefing can still look coherent. This is the central difficulty. People judge the output they can see. They cannot easily judge an input that never reached the output. Compression changes meaning Summarisation is often described as shortening. In practice, a useful summary must decide what is material. That requires interpretation. Consider a long customer thread containing: a technical question; an apology from an account manager; a revised delivery date; an implied threat to terminate; a request for written confirmation by Friday. A summary can be factually correct and still fail if it omits the termination risk or the required response date. No sentence needs to be fabricated for the result to be misleading. The same applies to a daily briefing. A list of five real tasks does not prove that the sixth, omitted task was unimportant. Ranking changes behaviour Priority labels affect what a person reads first, what they defer, and what they may never open. A filter that shows only high-priority messages turns a classification into a practical visibility boundary. Microsoft's documentation states that Copilot can mark mail as high, normal, or low priority, and that users can filter or sort by those classifications.[1] The low-priority icon is off by default. A user may therefore see positive signals about what the system selected without receiving an equally prominent account of what it deprioritised. The email remains available. The user's attention may not. A briefing creates implied completeness A product does not need to claim that its output is complete for users to treat it that way. Labels such as "prepare for your day", "catch up", "what matters", and "priority" create an expectation that the system has reviewed the relevant field and selected the important parts. The interface may show a polished list rather than a coverage report. It may not say: how many records were evaluated; which folders or message types were excluded; whether an attachment could be read; whether the model changed since yesterday; whether a security filter removed part of the context; how often comparable important items are missed. The user receives an answer without a clear account of its boundaries. 2. Importance is not a property of a message An email does not contain a universal importance value. Importance depends on the recipient, their role, the time, the surrounding events, and the consequence of inaction. Several different ideas are often collapsed into one label: Concept Question Urgency How soon must someone act? Consequence What happens if nobody acts? Relevance Does this relate to the recipient's current work? Authority Does the sender hold organisational power? Actionability Is there a clear request or decision? Novelty Is this outside the normal pattern? Familiarity Has the user often engaged with this sender or topic? These concepts overlap, but they are not interchangeable. A routine request from a manager may be familiar, authoritative, and actionable without being consequential. A legal notice from an unknown sender may be unfamiliar and written in neutral language while carrying a strict deadline. A security alert may be automated, repetitive, and important only once. A customer may imply cancellation without using an urgent keyword. The system must infer which of these signals matters now. Personalisation does not remove the ambiguity Outlook allows a user to provide natural-language instructions such as "It's from my manager" or "It's about a customer complaint". Microsoft says the feature also uses past behaviour and organisational context, and adapts as communication patterns change.[1][2] This can improve relevance. It can also preserve assumptions that no longer hold. A person may change role. A project may enter an incident phase. A new regulator, customer, supplier, or manager may appear. A sender who was previously unimportant may become the source of the most consequential message in the mailbox. Past behaviour is useful when the future resembles the past. Critical events often matter because they do not. Product scope can be mistaken for judgement Microsoft documents several categories that Prioritize does not evaluate: mail delivered to a folder other than the Inbox; mail marked as low importance by the sender; out-of-office messages; meeting invitations and related meeting mail; mail sent only to oneself; messages with very little content; non-delivery receipts; quarantine messages; encrypted messages whose content cannot be read.[1] These are understandable implementation boundaries. They are also possible sources of false confidence. An item may be absent because the system judged it unimportant, because a rule moved it, because its content was inaccessible, or because the product did not support that message type. Those are different causes. The user may experience the same result: the item is not in the briefing. Exact weighting remains opaque Vendors disclose broad categories of signals and examples of custom instructions. They do not publish the exact weighting, thresholds, model behaviour, or error rates that produce a particular priority. There are valid reasons not to expose every internal detail. Full disclosure could create security problems and would become stale as models change. The absence of exact weights does not remove the need for evidence. A buyer does not need source code to ask: What percentage of predefined critical items does the system surface? How does performance change for new senders? Which content types are not assessed? How often does a product update alter the result? What happens when messages contain instructions aimed at the model? An opaque method can still be evaluated. At present, the public evaluation is limited. 3. The asymmetry of attention errors The common discussion of generative AI errors focuses on false content. Attention systems introduce a second problem: missing content. Fail [truncated for AI cost control]