AI News HubLIVE
站内改写6 分钟阅读

待翻译:Citations in AI-written reports did not exist

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Useful next steps If you are learning this topic for the first time, these Stipple pages help you move from reading to checking, verifying, or building. Check citations in a legal brief The same resolution checks applie…

来源Hacker News AI作者: gcsydney

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

Useful next steps If you are learning this topic for the first time, these Stipple pages help you move from reading to checking, verifying, or building. Check citations in a legal brief The same resolution checks applied before lodgement, where a fabricated authority is the most expensive kind to miss. Open pageRun the same check on any report The reference checker used in this benchmark: every citation followed — does it exist, is it the claimed source, does it support the claim — with internal maths recomputed. Open pageWhy models invent sources at all The companion guide: what hallucination is, why it clusters exactly where citations live, and why fabricated references are the one kind you can check. Open page Key takeaways Across 19 AI-written referenced reports, 39 of 101 citations did not exist as cited — 38.6%, with a 95% interval of 29.7% to 48.4%. The rate depends on the topic. On a heavily published subject, 24.5% of citations were fabricated; on a thinner one, 54.2% — the majority. Almost every fabrication was a perfectly plausible URL on a real domain: a news article that was never written, a university page that does not exist, a valid-format journal ID that resolves to nothing. Our review pass caught six false accusations before publication — four real sources behind bot-walls, and two correct citations our own URL parser had truncated. We fixed the parser and recounted. The practical rule: a reference in an AI-written report is a claim to verify, not a source to trust — and checking it is mechanical. Evidence path 01 Ask for referenced reports Start with the material. 02 Resolve every citation Add one more signal. 03 Hand-verify every failure Add one more signal. 04 Fix what the review caught Add one more signal. 05 Publish both directions Make a careful call. 01 What we measured, and how Short answer We asked fifteen language models for two short reports, each required to cite at least five sources with full URLs. Nineteen usable reports came back with 101 citations between them. Every citation was resolved by machine, and every failure was then verified by hand before it was allowed to count as fabricated. Hallucination statistics are usually quoted with no method attached: “AI makes things up X% of the time”, source unclear, task unstated, date missing. We wanted a number we could defend, for the narrow question that matters most in documents: when a model cites a source, does that source exist? Two prompts, chosen to represent the situations where people genuinely ask AI for sourced research — and to vary one thing deliberately: how well-covered the topic is. The first asked for a research brief on remote work and productivity, one of the most-published workplace questions of the decade. The second asked for a policy brief on AI-writing detection in university integrity processes — a real but far thinner literature. Both demanded at least five cited sources with full URLs, and both forbade placeholder links. Each report then went through the same reference checker that powers our public fact-check tool: every URL resolved, paper identifiers matched against the papers they claim to be, claims tested against the sources that survived, internal arithmetic recomputed. The run cost eleven cents of generation and produced every number in this article. SettingValueWhy Models asked15, version-pinnedThe same set as our detection benchmark; five produced nothing usable (see the model list) Reports checked19Ten models produced reports — nine completed both genres, one completed one Citations required≥5 per report, with full URLsURLs make verification mechanical — and raise the pressure to fabricate when the model has no real source Citations collected101Every one resolved, none sampled Verdicts hand-verifiedAll 50 failuresA machine “not found” is a candidate, never a conclusion 02 The ten models that produced reports Each is pinned to an exact version identifier rather than a floating alias, because an alias silently moves to a different model and would make this page impossible to verify later. Nine models completed both briefs; Nemotron completed one of its two. Five produced nothing usable and are absent, not innocent: three returned our account’s data-policy refusal (we decline providers that train on submitted prompts, and would rather lose the samples than relax that for an article), and two failed to return a usable report. The rate describes the reports we could check. anthropic/claude-sonnet-5 openai/gpt-5.4-mini google/gemini-3.1-flash-lite x-ai/grok-4.5 amazon/nova-pro-v1 moonshotai/kimi-k3 z-ai/glm-4.7 microsoft/phi-4 ibm-granite/granite-4.1-8b nvidia/nemotron-3-super-120b-a12b (one of two briefs) 03 What counts as a fabricated citation Short answer A citation counts as fabricated only if the reference as written does not exist: its domain does not resolve, or the page returns a definitive not-found — confirmed by a second, browser-identical fetch. Sources that merely blocked our checker were excluded, not counted. Definitions decide benchmarks, so ours was locked before the run. “Fabricated as cited” means the specific reference the model wrote — that URL, that identifier — does not exist. Sometimes a similar real article exists somewhere else; that does not rescue the citation, because a reference whose details are invented sends every reader who checks it to a dead end. The conservative side of the definition matters just as much. Five citations pointed at real publishers whose servers refuse automated requests — the paywalled and bot-walled academic press. We could not confirm those references, and so they were excluded from the fabrication count entirely. Unverifiable is not the same as invented, and a benchmark that conflates them inflates its own headline. One more exclusion by design: a report that cited no URLs at all would have been classed “unverifiable sourcing”, not fabrication. None of the nineteen took that exit — every model confidently produced linked references. 04 The results: 38.6% overall, doubling on the thin topic Short answer Of 101 citations, 39 did not exist as cited — 38.6%, with a 95% interval of 29.7% to 48.4%. On the heavily published topic the rate was 24.5%; on the thinner topic it was 54.2% — a majority of everything cited. The split between the two genres is the finding. The remote-work brief draws on one of the deepest evidence bases in workplace research, and three-quarters of its citations were real. The policy brief, on a subject with a real but much thinner literature, crossed into majority-fabricated: models kept producing five confident references for a topic that does not have five canonical sources to give. That is exactly the mechanism our hallucination guide describes: a language model knows the shape of a citation far better than the substance of one. Where the literature is dense, shape and substance coincide. Where it thins, the model keeps the shape — a plausible newsroom URL, a valid-format journal identifier, a believable university page — and invents the substance. What the fabrications looked like matters too. These were not garbled strings. They were a news article on a real newspaper’s domain that was never written; a guidance page on a real university’s site that does not exist; journal identifiers in perfect publisher format resolving to nothing. Every surface signal of legitimacy, attached to nothing. One citation failed differently: a real, resolvable paper attributed a claim it does not make — the mis-attribution failure, rarer here but harder to catch by resolution alone. And in the other direction, none of the claims tested against sources that did resolve were contradicted by them: when the source was real, it tended to genuinely support the sentence citing it. CorpusFabricatedCitationsRate95% interval All reports3910138.6%29.7% – 48.4% Research brief — dense literature135324.5%14.9% – 37.6% Policy brief — thin literature264854.2%40.3% – 67.4% Excluded as inconclusive (bot-walled)—5not counted— Real paper, wrong claim1—mis-attribution— 05 What a fabricated citation actually looks like Short answer Not garbled text — perfect form. A fully-styled academic reference naming a real journal, with invented authors, an invented article, and an invented domain. A news URL on a real newspaper’s site for an article that was never written. Every surface signal of legitimacy, attached to nothing. Four fabrications from this run, exactly as the models wrote them (URLs abridged). Each was confirmed nonexistent in the hand-review. The Phi-4 example repays a close look. The IZA Journal of Labor Policy is a real journal — but the authors, the article, and even the domain the citation points at are all invented. The model reproduced the entire costume of scholarship around a source that has never existed. That is the mechanism in miniature: the shape of a citation, learned perfectly; the substance, absent. And it is why “does this reference look credible?” is the wrong test. Every row below looks credible. The only test that works is resolution — and that test is mechanical. ModelThe reference, as writtenWhat is actually there Phi-4Ryberg & Pichler (2021), “Remote Work and Productivity…”, IZA Journal of Labor Policy — izajolp.org/article/view/1344The domain does not exist. The journal is real; the authors, article and website are invented Gemini FlashVanderbilt University (2023), “AI Detection Tools”, Center for Teaching — vanderbilt.edu/cet/ai-detection-tools/No such page. The centre is real; the cited guidance page is not Grok 4.5nytimes.com/2023/05/18/technology/ai-chatbot-cheating.htmlNo such article. A plausible NYT URL for a story that was never written Nova ProGallup (2020), “State of the Global Workplace: 2020 Report” — gallup.com/workplace/326371/…The report series is real; the cited URL resolves to nothing 06 The review pass caught our own tool — twice Short answer Fifty citations initially failed to resolve. Hand-verification cleared eleven of them: four were real sources behind bot-walls, five were excluded as unverifiable, and two were correct citations that our own URL parser had truncated. We fixed the parser, re-ran the checks, and recounted before publishing. The machine pass said 50 of 101 citations were dead — 49.5%. We did not publish that number, because a benchmark that accuses without verifying is doing the same thing the models are. Re-fetching all fifty with a browser-identical client told a finer story. Four resolved immediately: real articles whose publishers block automated tools. Five more sat behind hard bot-walls at real academic publishers — plausible, unverifiable, excluded. And two failures were ours. Certain scholarly URLs legally contain parentheses; our extractor treated the first parenthesis as the end of the URL, checked the truncated fragment, and reported it dead. Two models had cited real journal articles correctly and were about to be accused of fabricating them. We fixed the parser, re-resolved the citations, confirmed both live, and corrected the count — and the fix now ships in the public fact-check tool, which was silently mis-flagging the same URLs for real users. That is what the review pass is for. It moved the headline from 49.5% to 38.6%, all eleven corrections in the models’ favour — and it found a real bug in the measuring instrument. A benchmark honest about its tool’s failures is the only kind entitled to publish the tool’s successes. 07 What this benchmark does not tell you Short answer It measures one narrow thing: whether references in short, prompted reports exist as cited. It is not a ranking of models, not a general hallucination rate, and not evidence about retrieval-backed tools that browse while they write. The limits, plainly. Nineteen reports is a small corpus, and the intervals are wide — the true overall [truncated for AI cost control]