跳到主要內容
AI News HubLIVE
站內改寫7 分鐘閱讀

待翻譯:Human Judgment Doesn’t Leave the Software Factory, It Relocates

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The following article originally appeared on Elevate and is being reposted here with the author’s permission. A software factory is a repeatable loop around software work. If you’re building a software factory, code good enough to ship still needs human taste and ownership. We’ll discuss this including whether you need a factory just yet. If […]

來源O'Reilly AI & ML Radar作者: Addy Osmani
待翻譯:Human Judgment Doesn’t Leave the Software Factory, It Relocates
回報錯誤

更正管道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

The following article originally appeared on Elevate and is being reposted here with the author’s permission. A software factory is a repeatable loop around software work. If you’re building a software factory, code good enough to ship still needs human taste and ownership. We’ll discuss this including whether you need a factory just yet. If so: You’ll likely need humans in the loop upfront for deciding on product intent, system design (if you care) and your quality bar. Do review code (lights-on factory) but be intentional with where it’s needed the most. I’ve found you want to watch out for where automated back-pressure breaks. Or where maintainability trade-offs need to be made. Aim for quality checks to happen as early and continuously as possible. Not all of them have to, but this includes type systems, automated tests, mutation testing, security scanners and linting for architecture rules. Number of checks ≠ quality. You’ll likely need to experiment with what checks give you the best signal to noise ratio. Be ready to tighten or relax your constraints deliberately. You want to build your factory so some aspects of human taste get encoded in the environment, the agent gives you evidence of its work being right, and where a human still “owns” what ships to production. Sponsored by Sonar: Your agents write the code. Your gate decides if it ships. AI agents are writing more of my code, faster than ever—but fast isn’t the same as shippable. So I let a coding agent build an app, then put it through SonarQube. Every commit gets the same deterministic check: deep cross-file analysis, a clear map of where the risk actually lives, and a quality gate that holds every human and every agent to one bar. The screenshot? A PR that didn’t pass. That’s the gate doing its job. Do you really need a software factory? In my experience, you can get surprisingly far with your stock coding harness! i.e., Claude Code or Codex, multiple sessions, good SPECs with verification baked in and constraints. You can even throw a batch of GitHub issues at them with implementation and human-involvement criteria, but it’s when this system needs to be repeatable and event-driven that a factory is helpful. So I started off by saying a software factory is a repeatable loop around software work. We can actually look at a prompt that demonstrates a very small factory loop here: Read GitHub issue #123 and the repository instructions before changing code. Implement only the stated acceptance criteria. Do not modify authentication, billing, migrations, or existing test assertions. Work in a branch and keep the diff reviewable. Run npm run lint, npm test, and npm run build. If a required check cannot run, stop and explain why. Open a draft pull request with the checks you ran, the remaining risks, and any decision a human still needs to make. Do not merge. A goal can keep this moving until the checks pass and we can poll GitHub issues for any specific labels or review open pull requests each morning. Branch protection could enforce a merge boundary and the human can stay in the loop by choosing what becomes ready, reviewing and making the final merge calls etc. Add a software factory when you need an event-driven queue of work (e.g. Slack triggers, GitHub issues, Linear, a backlog) to run in an isolated cloud environment to handle triage, implementation and testing with some explicit human babysitting. Some end their loop with a monitor agent watching production and filing issues which triage again. In my experience, the factory becomes useful when the hard part is making your different runs behave consistently, handing work off between agents and avoiding different sessions from claiming the same issue, preserving evidence and stopping production when human review is falling behind. What solves this might sound a little boring. For example, Warp mentions triaging every incoming issue into one of four states—ready-to-implement, ready-to-spec, needs-info, wait-to-implement—and the label is what fires the next agent. This label does a few jobs in one go: it’s the queue, the lock, and since a session only picks up what’s marked ready, it’s where a human can park stuff without saying no permanently. Workflow wise, there are a few similarities and differences to just using Claude/Codex: Steering: Agent needs course-correction, you can give input and redirect it. Notifications: How the factory says it’s blocked. This can be because a requirement was ambiguous, it started something risky or it needs human input (steering). Handoff: Move the task, its state, and context between the cloud factory/another agent/human reviewer. Good handoffs will keep track of what happened, what’s left to be done and why the handoff is needed. In a good factory, the human isn’t limited to just reviewing and approving the final diff at the very end. They can shape the work early on, steer it during implementation, get it through a handoff, or stop it shipping to production. Verification is where a responsible factory spends a lot of its time. We’ll cover this more later. If you decide you do need a software factory, building it isn’t the only option. Standing up the infra to scale a factory can be a lot of work and you may want to consider buying verus building. Factory, Warp and HumanLayer are all working on this. What my day looks like now My day-to-day experience of software development has changed a lot over the past year. I’ve been talking about increasingly doing a lot of parallel work with agents, moving towards having a lights-on software factory. And a lot of people have been asking me, like, what do these things actually mean? What are you building? What are the kinds of projects that you’re using these things on? It’s a lot of this: So on a very average day, I have a simpler lights-on software factory. I can have tasks that are running in the cloud. Half of these tasks might be working on production client applications with smaller companies that I’m working with. They’re going to have real users. They’re going to have real authentication, payments, subscriptions, real beefy risks that you need to be careful with. You can’t just say, “Oh, agent, just go and do this stuff” without having tests and constraints and quality checks in place. I could work on my open source projects. I could be building out companion sites for my books. I could be working on tools. I could be building apps of my own. And these are all very, very different kinds of applications that I’m working on. And sometimes all they have in common is the tool that I’m using to work on them, right? Maybe I’m working on a migration. Others, I might be doing actual beefy feature work. And the blast radius of the work might also be very, very different. So as you begin to think about getting to a place where we’re increasingly doing a lot of parallel work, we’re trying to improve velocity, we’re trying to improve productivity, and we’re trying to improve autonomy, which means getting the system to a place where we trust it more, you do have to think about what are the places that absolutely require human code review, human input. And a lot of that’s going to be required up front, right? When you’re defining your specification, your requirements, what’s the design of the product going to look like? What’s the intent of the product going to look like? And then, how are you verifying that the agents have actually gotten the work done right? How are you verifying that they haven’t broken the existing system that’s been in place? How are you making sure that it’s meeting your quality bar? So generating code is not necessarily the part that you need to worry about the most. Given enough context, agents can write the implementation, run the tests, inspect failure, and revise code for us. We need to get to a place where we feel like there’s enough of human taste encoded in the environment that we can trust what’s being built, so that our human attention can be focused on the places where it’s needed most. Now, there’s pushback from folks saying, “Hey, well, I don’t buy that you can just automate away a lot of this stuff.” It’s not to say that we’re automating away all of it, right? But given the volume of code that’s being generated, I don’t think that it’s realistic for humans to be reading all of it, especially when we’re not building rockets a lot of the time, right? We’re building UI, we’re building full stack applications. Our judgment, our taste is best focused on the places where it’s needed the most. Like, what are the riskiest parts of the systems? Where do we need to apply human taste? And that can be in the frontend. That can be in how the system works. It doesn’t have to be 100% of it. My cognitive bandwidth does not scale with the agents The reality is, yes, we can now fire up dozens, hundreds, thousands of agents in parallel, but your own cognitive bandwidth does not scale in the same way. This can feed into cognitive or comprehension debt which I’ve talked about before. If you remember back to just five, ten years ago, there was a lot of discussion in the engineering community about context switching and the cost of it. We would talk about how people hated when a colleague or someone would walk up to your desk when you were in the middle of a task. It would then take you so long to get back into your flow state because you had to catch back up in terms of like, where was I? What was I doing? Even if you had a little bit of residue there, it still took you time. We’re now context switching even more than we did before. On any given day, if I’m working outside of a software factory, I can be working on five or ten different projects with agents at a single time, or five or ten different features on a single project at a time. I can have five or ten different sessions, you can effectively say. That means that I have to be able to stay on top of at least a few of those. It is possible that I’m going to be able to increase how much autonomy I give some tasks if I have trust that I’ve defined the task well enough, I’ve defined the outcome, how it’s going to verify that it’s done well enough. But then there are going to be tasks where maybe I don’t necessarily feel that way and there’s more risk involved or more nuance. I’m going to have to pay attention. Consider optimizing the software factory for your reviewer. Given every one of those approaches still routes its output to one person’s attention, you should ask how much cheaper the factory is making the decisions you still have to make. A wrong-project mistake I remember when I’ve been working on multiple parallel projects with my agents, and there have been times when I’ve accidentally done things like, maybe I was working on a web app where I wanted to add in a dark mode, and so I had in my head, okay, well, this is what the shape of this needs to look like. But I accidentally went to the session for a different project, and I started putting in that same prompt. So I began implementing dark mode for something that absolutely didn’t need it. And so I can make that mistake. I don’t want my software factory making that kind of mistake. You need to think about this really in terms of a system. You are effectively trying to encode a software engineering culture, a team culture, into a system so that it has those same kinds of behaviors, so that it has ownership that belongs somewhere, so that someone is still on the hook for what happens, and you’re being very explicit about how you think about those things. When green is misleading Even in these systems, you want to be very careful, right? Many of us have seen that when you have asked AI to help you pass a test, like we’re talking about a programming test, a unit test, it can change the unit test to satisfy that condition, or it can change the logic of the code to pass that condition. T [truncated for AI cost control]

展開要點與分析

文章情報

工程師中級

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級後設資料。
  • The following article originally appeared on Elevate and is being reposted here with the author’s permission. A software factory is a repeatable loop around software work. If you’…

技術影響

可能影響 Agent 架構、工具呼叫、工作流自動化和產品整合。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。