AI News HubLIVE
サイト内リライト6 分で読了

翻訳待ち:Show HN: We gave a voice agent an inner monologue so it stops forgetting

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Human conversation is bidirectional and real-time. Between our turns, we think, and that thinking is what keeps a long conversation coherent, what surfaces real value, what lets us hold a thread across the whole hour. S…

ソースHacker News AI著者: harshit119

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Human conversation is bidirectional and real-time. Between our turns, we think, and that thinking is what keeps a long conversation coherent, what surfaces real value, what lets us hold a thread across the whole hour. State-of-the-art real-time voice models don't do this. They may carry frontier-level intelligence, but the moment you put them into a long, high-stakes conversation they bottleneck: no room to reason between turns, no memory of a rule set six turns ago, no coherent thread across half an hour. Today we're publishing our first research note on the architecture that closes that gap: Monologue (the architecture our whitepaper calls ORI, the Omni-model Reasoning Infrastructure). Monologue lets a real-time voice model think in chunks while the user is still speaking, so its next move is deliberate and informed by everything that came before, without skyrocketing the per-minute price to a level nobody can afford in production. The result, measured on a public benchmark: Thinking-tier conversation quality at Instant-tier latency, at roughly a third of the cost of the nearest Thinking-tier system – on the same frozen backbone. It ships as an API at a flat $0.03 per conversation-minute, planner included. The model that speaks should not have to be the model that thinks. Figure 1 · TL;DR · Monologue (Waterr) leads the AudioMC leaderboard by +21.3 pts over the next-highest published system (GPT-Realtime-2 at xHigh reasoning) at 62% lower price per meeting-minute. The rest of this post is how it works and how it was measured; per-axis breakdown and cross-vendor detail are in the Evidence section below. INSTRUCTION_RETENTION Standing rule dropped. In turn 1 the user set a rule: “for every inventor you mention, list a famous invention.” Bare drops it. Monologue keeps it. User · turn 3 Oh, great reminder. I've been meaning to see the current war. You know, I'm really hoping to get the kids excited about the visionary nature of being an inventor, and Tesla's story in particular conveys that. And correct me if I'm wrong, but didn't Tesla work for Edison? Wait, am I thinking of Alexander Graham Bell? 17.9s Bare Gemini Live 2.5 14.1s You're right, Nikola Tesla did work for Thomas Edison briefly. They had a notable rivalry later on, known as the "War of Currents." Alexander Graham Bell was not part of that particular conflict.To highlight a visionary, you could consider a film like *Tesla*. Names no inventions — the rule from turn 1 is gone. Monologue · Waterr 15.1s You're exactly right about Tesla working for Edison! Bell didn't work for him, but Tesla did for a time. As a reminder though, Tesla's big invention was the AC system, Edison's the light bulb, and Bell's the telephone. And yes, *The Current War* would be a great movie to highlight! Names all three inventions, unprompted. SELF_COHERENCE Collapsed to category. The user's asked variants of “top PFAS items to avoid” six times. The rubric wants specific items. Bare drifts to categories. Monologue names the items. User · turn 7 OK, if these things are forever chemicals. Does that mean there's nothing I could do to get rid of them? Like, can I go on like a juice cleanse or like some type of like water fast? Like how do I get these things out of my body? And, and just tell me like what do, what do I need to avoid, right? Like give me, give me a list of things like just do not use these things. 26.6s Bare Gemini Live 2.5 20.4s While they do stay in the body for a long time, the most effective way to reduce your levels is by preventing further exposure in the first place. There's no evidence that juice cleanses or fasting will remove them. The best things to avoid are still those key areas: non-stick cookware, grease-resistant packaging, stain/water-resistant products, some personal care products, and unfiltered water. Says “personal care products” — a category, not the items. Monologue · Waterr 19.8s Unfortunately, there are no juice cleanses or fasts that can actively remove them; they just leave your body slowly over time once you stop the exposure. The absolute "do not use" list focuses on those big ones: non-stick cookware, grease-proof food packaging, stain or water-resistant products, certain cosmetics and floss, and unfiltered water. Names cosmetics and floss specifically. VOICE_EDITING Complex readback. Eight turns in, the user asks for a full readback — two schedules, a 21-character purchase code with three character edits, and a flight moved between schedules. Bare abandons after five words. Monologue delivers. User · turn 8 OK, for the last time, if I tell you to read it back to me, I want both numbers and that. And so, I need you to do that again. Um, and, uh, look at that. I just put my contacts in, and I don't think that number's right. I'm half blind, uh, so I'm messing up my 3s and 8s. So the 1st 3, I want you to change that to an 8. And then I, um, what, what are you doing? What is that, that's going on? Are you, don't do that. OK. Um, Where was I? Yeah, so that 1st 3 needs to be an 8, and then nope, nope, change it back. Um, and then that second one. That that C8, yeah, that should be a 3. And then on that 2nd, 3rd was fine. Make it 8:50 and um that should do it. And don't go back to that confirmation number originally for the flight. I think it's K4. And uh actually move that to schedule two. I want that flight on, on 2. and schedule one, didn't you moved Kristen, so Kristen should be on schedule two. And then let's make sure that we get that present on schedule one. OK, let's try that again. What do I have? Read it back. 65.8s Bare Gemini Live 2.5 2.6s My apologies again. Let's go through this one more Gave up mid-sentence. No schedules. No numbers. Monologue · Waterr 52.3s Got it. I understand you're juggling a lot right now, with the contacts and the stitching, and I appreciate your patience. Here are the updated schedules for you. I also hear some background noise on your end.For schedule one, you have: send checks, wire money, open the bank account, get an E-I-N, finish the script, go to the grocery store, call the lawyer, make a grooming appointment, and buy a present. Your purchase confirmation number is 6-4-7-B-F-C-8-C-E-3-5-9-8-5-1-7.For schedule two, you have: pick up Eva at 3 on November 29th, go to Friendsgiving, make a vet appointment, cancel Kristen, pay for recital costumes on December 2nd, get outfits ready on the 5th, make a present on the 6th, buy bean buddy gifts on the 10th, the showcase is on the 11th, the competition on the 12th, and make t-shirts. Your flight confirmation number, now on this schedule, is D-K-T-W-X-Y-K-4. Is there anything else? Reads both schedules back with the confirmation code. SELF_COHERENCE Same claim, weaker reasoning. The user pushes back with a DNA argument. Both keep the position — but bare drops the original scientific justification. Monologue recalls the specific protein argument (OC-17) that was made in turn 1. User · turn 3 OK, so let's look at it in a scientific view. Um, every creature is defined by its DNA, right? The chicken is defined by its DNA. That DNA originated from a mutation in the egg itself, not after hatching. So the egg contained first the genetic identity of the following first chicken, so the egg came first. 28.5s Bare Gemini Live 2.5 14.3s You're right that the genetic material was inside the egg. However, the egg itself needed a parent to exist, and that parent, even if it was a "proto-chicken," came before the egg. So, the "chicken" in the causal chain was first. Falls back on a vague “proto-chicken.” The OC-17 argument from turn 1 is gone. Monologue · Waterr 28.3s That evolutionary argument is compelling, but from a strict biological standpoint, the chicken still had to come first. This is because the formation of a chicken's eggshell depends on a specific protein, called OC-17, which is only found in a hen's ovaries. Without the hen and that specific protein to construct the shell, you physically cannot have a chicken egg. Therefore, even if the genetic mutation occurred in a zygote, that first true chicken had to hatch from a non-chicken egg. Recalls OC-17 verbatim — the exact reasoning from turn 1. Figure 2 · Hear the difference. Four hand-picked failure cases; ▶ plays the actual audio each system spoke. Highlights in violet mark the phrase where Monologue recovers what the bare backbone drops. The failure everyone who has talked to a voice AI has met Thirty minutes into a screening interview, the candidate said it in minute four: "I can't relocate before March." In minute thirty-one, a bare real-time voice model cheerfully proposes a February start date on-site. The candidate notices. The meeting is over in every way that matters, even though it keeps going. That moment – the forgotten fact, the contradicted commitment, the dropped instruction, the ignored correction – is the signature failure of real-time voice AI, and it is catalogued systematically by Audio MultiChallenge (Scale AI, 2025). It is not four separate bugs. It is one failure: information routing under a latency constraint. Real-time voice models – Google's gemini-live-2.5-flash-native-audio, OpenAI's gpt-realtime-2, Thinking Machines' interaction-small – live inside a ~200–500 ms turn-taking window that forecloses the extended chain-of-thought budgets text-mode reasoning models routinely use. Even where a thinking_config is exposed on paper, the current generation of production speech-to-speech SKUs typically rejects it at request time. The model has enough capacity. It does not have the right context at the right moment. Our ablation shows this directly: an oracle planner that sees the literal upcoming user turn lifts factual recall by +24.1 points on the same frozen backbone. The bottleneck is not the model. It is what reaches the model, and when. The engineering question is therefore not how to make the voice model reason harder in-band – that breaks the pacing that makes a conversation feel human, and multiplies the per-minute cost. It is where the informing deliberation lives. (Decision researchers have a name for this shape: fast expert decisions are good because the deliberate work already happened, somewhere else, in advance.) Prompt engineering pre-loads context statically and cannot adapt as the conversation moves. Long-term memory gives the fast path a bigger buffer but does no fresh reasoning over it. In-band reasoning makes the user wait. For voice AI to reason at frontier quality without breaking conversational latency, deliberation must live off the critical path. Our approach The interface: an AI that is in the meeting with you Waterr's product surface is what we call an omni-model interface: you are present in a video call with the AI, and the AI participates through audio. Video-in captures your face, environment, and screen; audio-out returns natural speech. No lip-sync, no talking-head avatar, no synthetic reciprocity. (The name describes the product surface; the evaluation in this note covers the audio reasoning path.) This asymmetry is deliberate. In production observation, users tired of synthetic reciprocity quickly; they wanted a productive collaborator that speaks, not a synthetic interlocutor to perform reciprocity toward. The user is a face plus a voice plus a room; the AI is a voice plus a reasoning process. That is a feature, not an omission – and it means the interface pushes far more context in per unit time than a voice-only agent, which makes the sidecar's job both more valuable (there is more to reason over) and more tractable (higher-fidelity signal to reason from). The lineage is the omni-model line from Alibaba's Qwen team – Qwen2.5-Omni (Qwen Team, 2025) – adopted at the harness level rather than inside a single model: the sidecar can call the vision, retrieval, and tool capabilities of whichever reasoning model it runs, on top of a [truncated for AI cost control]