How do I check whether AI-generated citations are real?
Three steps, in order. Search the exact title in a bibliographic index. If it resolves, confirm the authors, year, journal, volume, and pages match. Then open the paper and read the passage that is supposed to support the claim. A reference can exist and still not say what was attributed to it.
Updated
Verification is fast once you have a routine, and the routine matters because doing it ad hoc means skipping it under deadline. Budget one to two minutes per reference and work through the list in one sitting rather than checking as you write.
Step one is existence. Paste the exact title, in quotation marks, into Crossref, PubMed, Scopus, Web of Science, or your library catalog. Not into the chat window, and not into a general web search alone, since a fabricated title sometimes returns pages that merely quote the same fabricated reference.
Step two is identity. A title that resolves is not enough, because a common failure is a real paper attached to the wrong authors, year, or venue. Check every field against the record you found, and correct the reference from the record rather than from the generated version.
Step three is support, and it is the one people skip. Open the paper and find the passage. A reference that exists, is correctly formatted, and does not contain the claim is worse than a fabricated one, because nothing about it looks wrong until an examiner follows it.
Published measurements make the case for the routine. In one 2023 study of generated academic text, over half of the citations from one model did not exist, and among the real ones a substantial share carried substantive errors.
The rate varies by model and, more usefully for you, by source type. Across 636 references generated for 84 short literature reviews, 55 percent of the GPT-3.5 citations and 18 percent of the GPT-4 citations pointed to works that do not exist (Walters and Wilder, 2023, p. 4). For GPT-4 the failure concentrated in one format: 70 percent of cited book chapters were fabricated, against 18 percent of journal articles, 10 percent of websites, and 8 percent of books. Check the chapters in your list first, and verify both the chapter and the edited volume, since fabricated chapters often named books that do not exist either.
What does a fabricated citation usually look like?
Entirely plausible. Real authors who publish in the area, a real journal, a title describing the paper you wanted, and a volume and page range in the right format. The giveaway is that the title returns nothing in any index, or returns a real paper by different authors in a different year.
Existence is only the first field that fails. Among the references in that study that turned out to be real, 43 percent of the GPT-3.5 set and 24 percent of the GPT-4 set carried at least one substantive error, including wrong volume, issue, or page numbers in 34 percent of GPT-3.5 articles and chapters, wrong dates in 22 percent, and wrong author names in 14 percent (Walters and Wilder, 2023, p. 4). GPT-4 halved those rates without removing them. A real paper with wrong metadata is a correction job, so fix the fields from the record rather than dropping the reference.
The DOI deserves special caution. A generated reference sometimes carries a DOI that resolves to a completely different paper, which looks like confirmation if you only check that the link works rather than what it opens.
The pattern to watch for is a reference that is exactly what your argument needed. Models produce the paper you were hoping for, and that convenience is a reason for more scrutiny, not less.
How do I verify that the source says what the AI claims?
Search inside the full text for the specific terms of the claim, then read the surrounding paragraph. Confirm the population, the measure, and the strength of the finding, since the most common error is not invention but a real finding described more broadly than the paper supports.
Existence and support are separate failures, and the second one predates AI. Jergas and Baethge reviewed quotation accuracy in medical journal articles and put bibliographic accuracy explicitly outside their scope, treating whether a source says what it was cited for as its own question (Jergas and Baethge, 2015, p. 2). They found substantial quotation inaccuracy across more than two dozen studies, and the estimate held across their sensitivity and subgroup analyses (Jergas and Baethge, 2015, p. 5). Human authors misreport their sources at a serious rate, so a reference that resolves cleanly still tells you nothing about the sentence it is attached to.
Do not ask the model to certify its own list. Walters and Wilder report that ChatGPT often answered inaccurately when asked whether the works it cited were legitimate, and their dataset included one response that described its own bibliography as sample entries rather than actual sources (Walters and Wilder, 2023, p. 4). Those admissions are a warning, not a verification. The same authors recommend a second pass after existence and metadata: evaluate the quality of the works themselves, including whether a reference points to a publisher identified as predatory (Walters and Wilder, 2023, p. 6).
Where a tool gives you a page number, go to that page and check it. Where it does not, that absence is itself informative about how the claim was produced.
Record what you verified and when. A short column in your reference list marking each entry as checked, with the date, takes seconds per row and is exactly the evidence you want if anyone asks how the bibliography was assembled.
Sources
- Walters and Wilder, Fabrication and errors in the bibliographic citations generated by ChatGPT, Scientific Reports (2023) · checked 6 August 2026
- Smith et al., Hallucination or Confabulation? Neuroanatomy as metaphor in Large Language Models, PLOS Digital Health (2023) · checked 6 August 2026
- Jergas and Baethge, Quotation accuracy in medical journal articles, a systematic review and meta-analysis, PeerJ (2015) · checked 6 August 2026