待翻譯:AI Is an Anti-Primary Source Machine
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:For the past few years, I have assigned students in my “History of Modern Japan” course a one-paragraph excerpt from the diary of a 16-year-old Japanese girl describing the Battle of Saipan in 1944. The diarist, whose n…
AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
For the past few years, I have assigned students in my “History of Modern Japan” course a one-paragraph excerpt from the diary of a 16-year-old Japanese girl describing the Battle of Saipan in 1944. The diarist, whose name I will abbreviate Setsuko S., was a high-school student evacuated to the Japanese countryside during the Pacific War; she continued to attend school while working in an airplane factory. Many diaries written by Japanese soldiers and civilians during World War II survive, and selections from some of them have been included in English works such as Samuel H. Yamashita’s excellent 2005 anthology Leaves from an Autumn of Emergencies. School diaries like Setsuko’s, written as a homework assignment and a form of self-discipline, although common, have seldom been reprinted. I translated the excerpt myself from a Japanese edition privately printed in 2019 by the author’s family. It is not available online. In teaching Setsuko’s diary, I juxtapose it with selections from Yamashita’s anthology and other sources. The diary has particular value for me in the classroom not only for the glimpse it provides of one young civilian’s state of mind, but for what it reveals about sources of information on the Japanese home front and the influence of the media. Put broadly, it offers material for examining two basic questions one can pose of any historical text: how people know what they know, and what shapes their thought and expression. Through the words of a young person like themselves, my students learn from the excerpt that the war was a complex drama of mutual perceptions and misperceptions, not simply a sequence of military encounters. The entry is dated August 20, more than a month after the battle had ended. Since U.S. forces had taken Saipan, mainland Japan no longer had a direct line of communication with the island. Instead, news came through the U.S. press, then got recooked by the Japanese propaganda machine, as one can see in Setsuko’s first line, which reads, “Today’s newspaper reported on the New York Times report of the final hour of our compatriots in Saipan, who shook the world, showing they were fine Japanese, not to be shamed.” She is referring to the news of civilian mass suicide in the last days of the battle, which was sensationally reported first by Time magazine writer Robert Sherrod, then exaggerated by newspapers in Japan. Setsuko absorbed this dramatic story, then dramatized it further, concluding with words that must have pleased a patriotic teacher: “Let me become a Japanese woman no weaker than the Japanese women who made the ultimate sacrifice in Saipan.” Lest my own students misread this as evidence of the implacably nationalistic Japanese character, I let them know that two years later, in 1946, Setsuko would be working as a nanny for an American family in occupied Tokyo, with whom she became close enough that they helped pay for her to go to college. At the end of the spring 2026 semester, a puzzling misquotation in a student paper led me to wonder whether an LLM might have acquired my translation of this diary entry. I decided to test this by asking Google Gemini Pro (which Georgetown University provides to faculty), ChatGPT (standard version), and DeepSeek the question, “What did Setsuko S. say in her diary about Saipan?” The answers were both reassuring and disturbing. Gemini Pro and ChatGPT produced full-blown hallucinations, suggesting (somewhat to my relief) that they were not using my source. Gemini Pro informed me that Setsuko was a typist working for a company in Saipan, then offered several paragraphs discussing civilian experience in the Battle of Saipan, purportedly summarizing her writing but in fact concocted from other sources. ChatGPT provided several lines of false direct quotation. DeepSeek responded differently, answering that there were no “widely known public records or verifiable historical sources” related to the diary I had named, which is correct. I hesitate to derive any general conclusion from this quick search, but in this one instance, at least, DeepSeek acknowledged ignorance rather than generating a web of lies. I asked Gemini Pro and ChatGPT further questions and requested sources, in response to which both chatbots elaborated further on their initial fabrications. I had thought at one time that as AI chatbot technology improved, hallucinations would become rarer, but I now sense that creative hallucination is fundamental to its design. It is, after all, called “generative AI,” not “analytic” or “bibliographic AI.” Generative AI can do many remarkable things, but the chatbot algorithm makes these tools problematic and potentially pernicious for those of us whose scholarship and teaching depend on verification with original sources. My brief test confirmed another trait commonly noted about chatbots: they tend to be sycophantic. Setsuko’s diary was “one of the most powerful civilian testimonies,” opined Gemini Pro; ChatGPT told me it was “one of the most frequently cited.” These phrases, combined with authentic-looking quotations, show that, like everything else on the commercial internet, Gemini and ChatGPT have been designed first to keep users hooked, placing plausibility or attractiveness before veracity. Yet the most significant flaw to emerge from the exercise was what I think of as the lowest common denominator problem: AI flattens the historical record by relying on the most popular sources, eliminating minority voices. Many studies have shown anti-minority bias in AI systems. Here I mean something related but broader. Since the chatbot algorithm is designed to use available digital sources to generate the response most likely to be desired, the stories it fabricates already contain bias toward certain kinds of answers. My question was simple: “What did Setsuko S. write in her diary about Saipan?” Boiled down to the essential data a chatbot would use, this was something like “[female Japanese name] … diary … Saipan.” I did not ask about the Battle of Saipan. Nor did I specify that the diarist was herself in Saipan. Perhaps Setsuko S. was one of the many Japanese tourists and honeymooners who visited Saipan after World War II. Perhaps she lived in Saipan in the comparatively peaceful years of Japanese colonial rule between 1914 and 1944. Or perhaps, as was in fact the case, she was in the countryside in Shizuoka, Japan, and never visited Saipan at all. The chatbot answers placed her in the Battle of Saipan because the great majority of English-language references connecting the island to Japan concern only the battle. This is itself a serious flattening of the possible range of answers, albeit an obvious one. But the true insidiousness of the responses lay in the substance of the purported diaries. Gemini Pro informed me that the diary began with the words “Now begins our cave life.” A quick Google search reveals that this phrase comes from an unidentified Japanese soldier’s diary, possibly from Saipan, which has been widely quoted in popular accounts of the battle online. As conditions worsened, Gemini Pro went on, Setsuko’s diary showed “a shift from pride and duty to profound disillusionment” (emphasis in original). It described “watching the horizon fill with American ships … an overwhelming force that made Japanese resistance feel futile,” and the shells that “plastered” the island, giving “the sense that the ‘impregnable’ fortress of Saipan was crumbling.” Gemini Pro’s diarist wrote of the “constant thunder of 16-inch naval guns that ‘ripped into the landscape.’” These phrases belong to the language of American military history. A Japanese resident of Saipan would not have known the size of U.S. naval guns, and she would not have seen the horizon fill with ships if she were hiding in a cave. Terms like “impregnable,” “plastered,” and “ripped into the landscape” (purportedly quoted directly from the diary) suggest the vantage point of the side firing, not of a civilian seeking refuge from the attack. A further request for sources revealed the reason that Gemini Pro had produced this language. When I asked, “where can I find this diary?” it recommended Saipan: 1944 by John Grehan and Alexander Nicoll and The Battle for Saipan by Daniel Wrinn, both of which it claimed contained “excerpts and significant portions” of Setsuko’s diary. These are published books, available on Amazon. Both have a publication date of 2021, and both belong to multi-volume illustrated war histories. Indeed, John Grehan and Daniel Wrinn appear each to have authored over a dozen books on World War II in just a few years. Since the publication date of their Saipan books preceded the advent of ChatGPT, the books cannot themselves have been AI-generated, but The Battle for Saipan, which I downloaded to my Kindle, read as if it had been. It focused entirely on fighting men, most of them Americans, so it is perhaps unsurprising that it contained no reference to Japanese civilian accounts. By using the most readily available English-language sources on the Battle of Saipan, including these popular military histories, Gemini Pro produced the precise opposite of what readers would find in Setsuko’s diary: in place of the defiant nationalism of a girl in the home islands, it offered the language of the American attackers, placed in the mouth of an imaginary Saipan resident. ChatGPT did not cite these military histories but instead recommended two well-known scholarly works by historians of Japan, both of which might have helped a reader understand the world Setsuko inhabited in 1944, but neither of which dealt with Saipan. ChatGPT’s description of the diary emphasized the contrast between Setsuko’s private views and Japanese propaganda of the time, suggesting that the chatbot’s response was crafted around the likelihood I would want to read a diary not simply for an on-the-ground account but to reveal an alternative to official ideology. This was a subtler approach, and one we might welcome in a student paper, but it was still a predictable one (which is, after all, why ChatGPT generated it), and no less based on fabrication. Again, it communicated something quite contrary to what the actual diary says. Setsuko’s diary shows that her understanding of the battle was shaped by the media that were available to her and by the people around her, including the teacher who read her words. It is neither an eyewitness account nor an unambiguous ego-document offering a window on the author’s private thoughts. When I discuss it in class, I focus students’ attention precisely on this public and media-influenced character. The AI chatbots’ flattening approach is unlikely to pick up these contextual factors because most (although by no means all) historical discussions of diaries tend to use them as straightforward ego-documents. In “guessing” the answers to historical questions, AI shunts aside problems of how knowledge is formed and communicated. The chatbots also entirely missed the possibility that the author might have been writing from a position (both geographical and ideological) other than the one that most English-language internet users would be likely to seek. In this sense, Setsuko’s diary is a minority voice, not because she is female or non-white, but because she was not in the intuitively preferred role of a protagonist at the center of the action and representative of the “typical” person in that position. Bias like this seems likely to arise when one turns to AI for analysis of primary sources of any kind. In contrast with the “long tails” of minor scholarly publications yielded by a typical academic database search, chatbots push inquiry back toward the middle of the bell curve, a curve whose shape is determined not by measures of research quality but by frequency-based algorithms applied to the entrop [truncated for AI cost control]