AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。
The following article originally appeared on Addy Osmani’s blog and is being republished here with the author’s permission. In the past year, the conversation around agentic engineering has moved to harnesses and loops, fleets and software factories. My 2 cents is engineers need to own the outer loop—the accountability for these systems. This only gets more true as powerful models like Fable and GPT-5.6 become available. Agents have leverage, and leverage creates obligations. Someone must be able to explain exactly what changed, why it was safe, and what will happen if they’re wrong. Otherwise, their actions can’t be justified. Which makes it unlikely their organization will ask for them in the first place. And so I want to talk about three terms. The first, Quality, refers to all the checks we install before we let the system loose. Those checks produce evidence, and from that evidence we derive a Verdict. The second, Verdict, refers to the final decision we make before work enters our dependent system: I’m the line-producer of this content. I run the team whose work is shipped under my name. The model may write the line, but the Verdict is mine. The work of my team will not enter our dependent systems without my decision. A Verdict is the production decision: Should we ship, block, redirect, narrow the response, add a guardrail, or reject outright? The third, Answerability, refers to the guarantee that if someone asks, I can explain why. To say this another way: Our agent (which I define as a model plus a harness of files, tools, memory, skills, sandboxes, permissions, observability, and recovery) is what runs our loop (which I define as investigation, implementation, verification, and repeat). And it’s what creates our software factory. The model is just the engine. The harness—tools, memory, permissions, sandboxes, tests—is the car you build around it so it can do real work safely. The loop is how one good run becomes a process you can trust to run again. Wrap that harness in a repeatable cycle—investigate, implement, verify, repeat—where an independent check, not the model’s own say-so, decides when the work is done. Now run many loops at once. A factory is loops at scale: The agents ship the work inside, while humans own the decisions at the boundary. And at the heart of that factory is a careful boundary between what’s inside the system and what’s outside it. Inside the system we collect inputs (from the product team’s intent, or knowledge of previously shipped work, or of recent incidents, or of specific feedback from users). The agent loop investigates the task, implements a plan, and verifies the result. Then, evidence crosses that boundary. A human, who owns the dependent system, sees the evidence and decides whether to proceed. And that, friends, is the shift we’re trying to make. Before, our agents were doing the inner loop of the execution loop. Now they run the inner execution loop. Engineers own the outer loop. Inside the system, there’s really just one kind of thing our agents are doing: capability. The capability to investigate tasks, implement plans, test their results, and report back. That’s the capability of a model. And as we’ve said, that future is already here. Outside the system, there’s a single kind of thing: agency. The agency to decide, verify, approve, and own. We’re still talking about code, you see. It just needs to live in a place and be performed by people who know what they’re doing. The potential for AI code is no longer marginal. In a Sonar 2026 survey, we asked teams about the share of their commits that were AI-assisted. It was small but nontrivial. And several of the respondents said they expect the share of AI-assisted commits to grow substantially. Sonar’s 2026 State of Code report found that 42% of committed code was AI-generated or significantly AI-assisted, with expectations for that share to keep growing rather than plateauing. Creation, in other words, is getting cheaper. Scarcer resources are review, validation, understanding, and maintenance. We moved the speed of generation faster than we moved the speed of control, and so we have a trust-verification gap. A lot of people we talk to still express some degree of distrust in AI code. Yet fewer of them seem to consistently build that distrust into their verification processes. That’s a dangerous place to be. We’re going to need cheaper, clearer ways to verify the trustworthiness of AI code. If you look at the GitLab June 2026 report, you’ll see that governance questions have shifted. GitLab’s June 2026 AI accountability research shows that review and validation are the current bottlenecks when using AI and, more worryingly, that governance usually happens after code creation, after we’ve accepted the risk and lost control over ownership. Today, it’s not just about control. It’s about what constraints we set on the system. It’s about how we’ll check the work with evidence, and how we’ll hold teams accountable. It’s about who will own what part of the AI lifecycle. So the final distinction in this series is between process and quality. Quality is the concept of backpressure. We mean it literally. We don’t want to grant our agents as much autonomy as they can possibly exercise. We want to grant them just enough autonomy that we have enough backpressure to stop them, regulate them, check their work, and ensure our humanity. Ordinary engineering holds up a lot of signals that indicate that the work being done is doing the right thing. Type checks, tests, hooks, sandbox limits, audit logs, monitors. Our engineering systems are full of these kinds of signals, and they’re designed to provide enough backpressure to keep the system honest. And so as long as our agents are emitting these same signals, we can trust our ordinary engineering to provide appropriate backpressure. Trusting our systems doesn’t mean we don’t want a human in the loop. It just means that the human doesn’t need to be in the inner loop. We want them in the constraints loop (What inputs, architectures, instructions, or invariants should we set?), the sampling loop (How much output should we sample and review?), the audit loop (What evidence should we keep, and how do we make sure our audit log is effective?), and the ownership loop (What part of the production boundary should we own?). But the human doesn’t need to be in the inner loop. The agent can ship more than you can review. And the scarce resource is your own core human judgment, informed by quality signals like logs or tests. The AI June 2026 report shows that, in the experimental setting, agentic delegation along hour-scale time horizons is essentially here. The work by OpenAI this year on agents and the future of work was a great source for these ideas. So we need to start thinking about how to establish this ownership boundary, as our systems start shipping more than we can review. And that’s where the answerability comes in. Because with long-horizon agents, the decisions made over hour-scale time horizons are just that—decisions. And not all the decisions are going to be recorded. You can’t trace them all back to input tokens. If all you’re doing is trusting that the output you get is the correct choice for the problem at hand, the hundreds or even thousands of human hours of work you’re going to need to reconstruct the chain of decisions that lead to it become impossible. And so, again, answerability becomes something that must be at the core of our system design. Three hidden costs And there are three hidden costs: Cognitive surrender ~ blindly accepting what AI gives you. When you delegate work to an agent, the work itself may appear to be the work of the agent. But it’s actually your work. It’s your reputation. It’s your responsibility. And it’s your software that suffered the defects in the output. And it’s your software that needs to be changed to reflect that output. So the agent’s output becomes your answer. And with it comes all the accountability. The Wharton study that put this together is reassuring when the AI is right. But when it’s wrong, the news isn’t great. When the AI was wrong, nearly three-quarters of people accepted it anyway, and felt more confident than they would have without the AI. Cognitive debt ~ erosion of your understanding and memory of how to solve problems. When you delegate work to an agent, you’re offloading all the thought work to the agent. And while thinking it all out yourself takes time and energy, thinking it out on a massive codebase takes resources that aren’t available when you’re trying to run up the learning curve. So the output you get is often unattainable by you. And the longer the time horizon of the agentic planning, the bigger the gap between the code the agent produces and your understanding of it becomes. The gap compounds. The debt accumulates. And the cost of climbing the learning curve grows almost exponentially. There’s a randomized controlled trial from Anthropic looking at whether engineers who lean on AI to write code understand it as well as engineers who write it themselves. The conclusion was gloomy: On a comprehension quiz, the engineers who worked through AI scored 17 percentage points lower than those who didn’t, 50% versus 67%. And then there’s the orchestration tax: It’s easy to spin up lots of agents now, but your cognitive bandwidth doesn’t parallelize in the same way. Steering your agent away from the worst behaviors, sorting the work the agent produces to identify the ones that need your attention, directing it to focus on the work you care about first, verifying your most important constraints and your most dangerous assumptions before you let it run. . . All of that takes work, and it can’t be automated. There’s no substitute for human judgment. Brownfield systems are especially dangerous here, because the system behavior you have to audit doesn’t live in the code. It lives in the scars. Fixes? Make attention the priority in your architectural decisions. Use worktrees, scopes, and evidence to reduce the coupling between your initial plan and the work that emerges from it. Time-box the effort to resolve unactionable steps. And make change in your software strictly an opt-in permission. Alpha, decay, and taste: These are the three core patterns that shape careers and performances across domains. Alpha is the lead part taken up by the highest achiever in the competition, when you’re playing your highest-value game move. Decays are established patterns that everyone learns through repetition and watching others (plateaus, if you like). Taste is the earliest we can sense the lead in an alpha or the change in a decay. It’s our judgment of what’s coming before we have any evidence that anything is happening. Paul Graham’s point is that when anyone can make anything, choosing what to make matters more, and Mitchell Hashimoto’s definition is the operational one: making high-quality qualitative judgments where no objective metric exists yet. From now on, taste drives everything. Alpha shifts are taste changes. And decays fade out because we start to taste something different. Next step? Operationalize your taste. How? Give it a name that reflects what you’re trying to move from limbic to conscious. Practice it in critique and examples. Make its rationale explicit. And keep making the move that delivers the most durable competitive advantage in your industry. What’s that? Keep moving the edge up from just doing the task to teaching it, systematizing it, deciding when it should be done, and owning the result. Everyone is a developer, but not everyone is an engineer. Engineering is what a developer turns into when they embrace a work discipline that is more strict: thorough and logically sound reasoning, consideration of constraints and tradeoffs, recognition of risk and exposure, and practical accountability. [truncated for AI cost control]