AI News HubLIVE
站内改写6 分钟阅读

待翻译:Human vs. AI vs. Human and AI: Who Does Better Work?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:There's a comfortable assumption running through most corporate AI strategy right now: a skilled person plus an AI tool will outperform either one alone. It's intuitive, it's reassuring, and it lets organizations roll o…

来源Hacker News AI作者: rafaelaziz

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

There's a comfortable assumption running through most corporate AI strategy right now: a skilled person plus an AI tool will outperform either one alone. It's intuitive, it's reassuring, and it lets organizations roll out AI everywhere without asking hard questions about where, specifically, it helps. It's also not what the evidence shows. We identified eight controlled empirical studies — through a structured literature search described below, not an arbitrary shortlist — that compare at least two of unassisted human performance, AI-alone performance, and human-plus-AI performance on a well-defined task, using real accuracy, speed, or quality measurements rather than survey sentiment. Only a subset of the eight contain all three conditions in a single design; the remainder provide controlled two-arm evidence (most often human-alone vs. human+AI, or human-alone vs. AI-alone) that helps test whether the broader pattern generalizes across tasks — which study includes which conditions is noted in each section below and in the summary table. They span clinical diagnosis, management consulting, customer support, professional writing, and software engineering. These are not eight versions of the same experiment. They differ in design, sample, and what they measure: most are randomized controlled trials, one is a staggered field rollout with a randomized pilot; some measure a single 20-minute task, one follows real production work across three companies over months. Read individually, each study answers a narrow question about its own task and population. Read together, a pattern emerges that is not "AI helps." It's closer to a fault line: AI's advantage is real, large, and reproducible on some tasks, and reverses into a disadvantage — or simply disappears — on others that look, to a human, more or less the same. A large independent meta-analysis of the broader literature finds a closely aligned pattern, which is the strongest evidence this isn't an artifact of which eight studies we happened to pick. Methodology: how these eight studies were selected To avoid presenting a hand-picked set as more authoritative than it is, we ran a structured — not a formal systematic-review-protocol — literature search across Google Scholar, Semantic Scholar, PubMed, SSRN, NBER, arXiv, the ACM Digital Library, and IEEE Xplore for controlled studies published in 2022 or later that met all of the following: (1) compare at least two of unassisted human performance, AI-alone performance, and human+AI performance; (2) use a real or realistic work task, not a survey or self-reported opinion; (3) report a quantitative outcome — accuracy, speed, a graded quality score, or an error/completion rate, not satisfaction; (4) are peer-reviewed and published, or are a working paper from an established research institution or lab, not a vendor blog post or marketing study; and (5) cover clinical diagnosis, knowledge work, customer service, writing, or software engineering, with a broader scan across education, forecasting, hiring, translation, legal work, and classic human-AI "centaur" research to check for major work we might otherwise miss. This is a transparent, criteria-driven search, not a formal systematic review — it does not follow a pre-registered PRISMA-style protocol, log exact search strings or per-database result counts, or use a second independent screener, and readers who need that standard of evidence should treat it accordingly. That search surfaced roughly two dozen candidates. Most were excluded for a specific, checkable reason: some measured self-reported time allocation or satisfaction rather than task performance (a large Microsoft 365 Copilot field study of over 7,000 workers, NBER Working Paper 33795, measured how workers reallocated their time, not whether their output improved); some were observational rather than randomized, which weakens causal claims (a study of 72,000+ GitHub pull requests found AI-assisted PRs merged faster, but without random assignment); and a few strong two-arm studies (AI-alone vs. human-alone, no combined condition) in medicine and forecasting were kept as corroborating context rather than counted as primary sources, to avoid overweighting any single domain. The eight studies below are the ones that survived that screen with the strongest designs available — eight primary studies in total, including one earlier preprint (Peng et al., 2023) that we retained specifically as a labeled historical predecessor to a later, stronger study on the same question, rather than as independent evidence in its own right. We also checked our pattern against the closest thing to independent confirmation available: Vaccaro, Almaatouq, and Malone's 2024 meta-analysis in Nature Human Behaviour, which pooled 106 prior experiments (370 effect sizes) comparing human-alone, AI-alone, and human-AI teams across many task types.1 Its headline finding — that human-AI combinations underperform the better of either alone on average, with the effect reversing between decision-making tasks (where combining hurts) and content-creation tasks (where combining helps) — is a closely aligned, task-dependent, non-uniform pattern to the one this piece finds in its own eight studies. That convergence is meaningful: it means the "it depends on the task" conclusion below isn't a quirk of which eight studies we selected, but matches what a systematic review of the wider literature already shows. The clearest three-way test: doctors, AI, and doctors using AI The most direct answer to "who does better work" comes from a single-blind randomized clinical trial published in JAMA Network Open on October 28, 2024, run across Stanford, Beth Israel Deaconess Medical Center, and the University of Virginia.2 Fifty US-licensed physicians (26 attendings, 24 residents; median 3 years in practice) in family medicine, internal medicine, and emergency medicine were randomized into two groups and given up to six clinical vignettes — patient histories, exam findings, lab results, adapted from a landmark 1994 diagnostic-systems study — to work through in a one-hour session. One group used conventional resources only (UpToDate, Google, and similar). The other used those same conventional resources plus ChatGPT Plus (GPT-4). A third, non-randomized condition ran the same vignettes through GPT-4 alone, three separate times, with no physician involved. Participants completed 244 cases in total (125 in the LLM group, 119 in the control group). The results, as published2: ConditionMedian diagnostic reasoning score Physicians using conventional resources74% (IQR 63–84) Physicians using conventional resources + ChatGPT Plus76% (IQR 66–87) — adjusted difference +2 points, 95% CI −4 to 8; not statistically significant, p=0.60 GPT-4 working alone (median of 3 runs)92% (IQR 82–97) — 16 points higher than conventional resources, 95% CI 2 to 30; statistically significant, p=0.03 On the clinical cases tested, the LLM working alone scored a median of 92% — clearly ahead of both physician groups, including physicians who had access to the LLM. Adding ChatGPT to a physician's workflow did not produce a statistically significant accuracy gain over conventional resources alone. Physicians using AI were modestly faster on average (519 seconds per case vs. 565), a difference the paper's own results text reports as not statistically significant (95% CI −195 to 31 seconds, p=0.20). It's worth being precise about what this study does and doesn't show, because the authors themselves are unusually direct about it. Their own words: "Results of this study should not be interpreted to indicate that LLMs should be used for diagnosis autonomously without physician oversight." The clinical vignettes were curated and summarized by human clinicians rather than experienced firsthand, the study excluded patient interviewing and data collection, and the setting was acontextual — none of which capture the full scope of real clinical reasoning. What the study does show, within that bounded scope, is that a physician's real-time judgment about when to trust or override an AI's suggestion did not reliably improve on the AI's unaided output — and on these cases, the human addition did not clear the bar of statistical significance in either direction. A second, larger trial by an overlapping author team complicates this picture in an important way. A follow-up randomized controlled trial published in Nature Medicine in 2025 — led by several of the same researchers, using a different task design — found that physicians using GPT-4 on patient care tasks did score significantly higher than physicians using conventional resources alone: +6.5 points (95% CI 2.7–10.2, p<0.001).3 In that trial, 92 physicians were split across three arms (conventional resources, GPT-4 plus conventional resources, and GPT-4 alone), and — as in the first study — GPT-4 alone was statistically indistinguishable from physicians using GPT-4 (difference −0.9%, 95% CI −9.0 to 7.2, p=0.8). GPT-4-assisted physicians also took meaningfully longer per case (+119 seconds, p=0.02). Put the two studies from the same research group side by side and the honest conclusion isn't "human+AI helps" or "human+AI doesn't help" — it's that the answer depends on task design in ways not yet fully understood, even to the researchers running near-identical experiments. Across both studies, GPT-4 alone was not significantly worse than the physician+GPT-4 condition; however, the studies differed materially in whether physician+AI improved on physicians using conventional resources — a significant gain in the 2025 trial, no significant gain in the 2024 trial. Where "human + AI" earns its keep — and where it stops The diagnosis studies aren't outliers so much as one end of a spectrum visible across the research base. The clearest articulation of why comes from a large field experiment run by researchers at Harvard Business School with Boston Consulting Group, since published in Organization Science as "Navigating the Jagged Technological Frontier."4 The paper's central idea — the "jagged frontier" — is that AI's competence doesn't fall off gradually as tasks get harder, the way a human's does. It's extraordinary on some tasks and unreliable on adjacent ones that look, to the person assigning the work, about the same. In that study, 758 BCG consultants completed 18 realistic consulting tasks, some with GPT-4 access and some without. On tasks the researchers classified as inside the model's competence zone, AI-assisted consultants completed 12.2% more tasks, 25.1% faster, at measurably higher quality than the unassisted group. On a task designed to sit outside that zone, AI-assisted consultants were 19% less likely to produce a correct solution than the unassisted control group — a statistically significant effect (p<0.01). The finding that should complicate any tidy "just train people better" narrative: consultants who additionally received explicit training in how to prompt and collaborate with the model did not do better on the out-of-frontier task — they did worse. The trained group's correctness dropped by about 24.5 percentage points relative to control, compared to a 13.9-point drop for consultants using GPT-4 without training (both significant; the gap between the two AI conditions was marginally significant, p≈0.08–0.10). Training did modestly improve output quality on inside-frontier tasks. But on the one task built to sit beyond the model's competence, teaching people to use the tool more fluently made them more confidently wrong, not less. That's the uncomfortable version of the "jagged frontier" finding: the fix for AI failing outside its competence zone isn't more prompting skill. It's knowing, task by task, whether you're inside the zone at all — something the paper doesn't show people learn to detect just by getting better at using the tool. The real, r [truncated for AI cost control]