AI News HubLIVE
In-site rewrite3 min read

The AI 'Ghosts' Contaminating Academic Publishing

Advertisement • Go ad free · Aug 27, 2026 at 2:14 PM “The academic record is being quietly haunted” by researchers with names like Elena Vasquez and Marcus Chen. “Elena Vasquez and Marcus Chen have appeared as volc…

SourceHacker News AIAuthor: saikatsg

Advertisement • Go ad free · Aug 27, 2026 at 2:14 PM “The academic record is being quietly haunted” by researchers with names like Elena Vasquez and Marcus Chen. “Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived,” a new preprint research paper from Samsung and the University of Warsaw said. The paper identified a number of names that co-authored hundreds of AI-generated academic papers, articles, and books. The authors don’t actually exist, but are instead names that large language models repeatedly produce when tasked with generating experts in certain fields. The paper, titled “The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing,” utilized a known phenomenon where certain LLMs will keep coming up with the same names in certain contexts. For example, In June, Sam wrote about how ChatGPT, Gemini, and Claude were likely to use the name Elias Thorne in fiction they generated, and that the character was often a lighthouse keeper. Similarly, users noticed that if they ask ChatGPT to generate a software developer, their name will often be Marcus Chen. From The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing The researchers were able to show not only that LLMs often default to the same names, but that they produce “correlated character ensembles,” meaning some names were more likely to appear together. Other names that were consistently generated by AI models include Elena Amara Okafor from Claude, Aris Thorne and Lena Petrova from Gemini, and Elara Voss from ChatGPT. Earlier this month, I reported a story about Research Gold, a company that offered what it claimed was human medical research, but that was in fact entirely AI generated. The founder and lead methodologist for that company was named Elena Vasquez. Research Gold removed Elena Vasquez from its site after I published the story. Michał Brzozowskim, the lead author of the paper, told me that searching for these names on Google turned up other instances of AI generated personalties. For example, following the killing of Alex Pretti at the hands of U.S. Border Patrol agents in January, a rumor spread on Facebook that he was fired from his nursing job for misconduct allegations. Snopes reported that the false statement was attributed to “executive director Dr. Elena Vasquez,” who does not exist. Brzozowskim was able to search databases of academic papers for the names they knew LLMs often generated. “On Zenodo, a CERN operated repository that mints real DataCite DOIs, we identify 1,655 ghost-authored records claiming nonexistent journals with fabricated publication dates.” A DOI, or a Digital Object Identifier, is a string of characters and numbers used to identify academic papers. The researchers saw that many of the papers authored by these AI names were backdated, meaning their publication dates were different from the date they were uploaded to Zenodo. Anyone with a free account can create a DOI on Zenodo, but the existence of AI generated papers with DOIs has impacts on other parts of the web and academic publishing. “These [AI generated papers] carry real DOIs harvestable by any scholarly aggregator; the infrastructure for large-scale scholarly record contamination is already in place. Ghost names additionally appear on ResearchGate, forming synthetic research groups with collaborators drawn from multiple model families, and are indexed without verification by Google Scholar and Semantic Scholar [...] The academic record is being quietly haunted.” Research Gate and Google Scholar are both aggregators of academic publishing that are likely to come up in search results. The researchers say that different LLMs and versions of those LLMs generate certain names so consistently, they think they can use them to determine the provenance of AI slop. For example, the name Elena Vasquez was particularly common in content generated by Claude Sonnet 4, so papers that list her as an author were likely generated by that LLM. However, Brzozowskim said that this level of accuracy might not hold for long now that AI generated content with these names is flooding the internet and feeding back into all AI models that are scraping the internet for training data. The fact that AI academic publishing is struggling to deal with the load of AI generated content isn’t new. We’ve previously reported that scientific journals have published AI generated text, that AI is impacting the peer review process, and that the open-access repository for preprint academic research Arxiv will now ban authors for a year if they are caught submitting AI generated work. The potential upside of this research is that these names might allow us to detect this AI generated content more easily. Advertisement • Go ad free • Hide