Can I trust AI summaries of research papers?

Trust them to tell you what a paper is about, not what it established. Summaries are reliable for triage and unreliable for detail, because the parts most often dropped are exactly the parts that decide whether a finding supports your claim: the population, the conditions, and the hedging.

Updated

The useful distinction is between a summary that describes a paper and one that reports its findings. Describing is easy and reliable: this is a longitudinal study of teacher retention using a national sample. Reporting is harder and is where errors concentrate, because a finding is a claim with conditions attached and summarization works by removing detail.

That gives you a usable rule. Summaries are safe for deciding whether to read something and for reminding yourself what a paper covered. They are not safe as the basis for a sentence in your thesis, because the sentence needs the conditions the summary removed.

The failure is usually not fabrication. It is compression in the direction of confidence, which is what summarization does to any hedged text, and academic writing is unusually hedged. A finding reported as holding "in this sample, under these conditions, with a small effect size" becomes "the study found that X improves Y".

There is evidence that looking right and being right are separate properties. Gao and colleagues gave ChatGPT only the titles and journal names of 50 medical papers from Nature Medicine, JAMA, NEJM, BMJ and Lancet, and asked for abstracts in each journal's style. The model could not read the papers, and its knowledge cutoff predated them. Every output looked like a scientific abstract, and the invented cohort sizes tracked the real ones closely enough to correlate at r = 0.76, but the numbers were generated rather than recovered from the studies, and only 16% of the outputs followed the requested journal's formatting (Gao et al., 2023, pp. 1-2, 4). That is a test of surface plausibility, not of faithful condensation, so read it as a warning about how convincing wrong output can be rather than as a measurement of summary accuracy.

Your own judgment is a weaker check than it feels. Four blinded reviewers in that study caught 68% of the generated abstracts and correctly passed 86% of the real ones, which means they accepted roughly a third of the fabrications and wrongly flagged 14% of genuine abstracts as machine written. They described the texts they suspected as vaguer and more formulaic, but detector scores did not differ significantly between the fabrications they caught and the ones they missed, p = 0.45 (Gao et al., 2023, pp. 1-2). Reading a summary and asking whether it feels off is not a verification step.

There is a second effect worth knowing about. A summary of a paper you have not read gives you the feeling of having read it, and that feeling is durable. Six months later you will remember the claim and not remember that you never opened the source.

What do summarizers systematically get wrong?

Four things. They drop the conditions a finding depended on. They flatten hedged claims into confident ones. They report what a paper discusses as what it found. And they lose the authors' own stated limitations, which is usually the most useful section for a critical reader.

The third of those is the sneakiest. Papers spend most of their length discussing other people's work, and a summary that does not distinguish between what a paper argues and what it merely reports will attribute the literature to the author.

Methodological detail is the fifth casualty, and it matters because two studies with the same headline finding can differ in ways that make one usable for your argument and the other not.

How do I check a summary quickly?

Read the abstract and the first paragraph of the discussion, which takes about ninety seconds, and compare. Then check one specific thing the summary asserts against the results section. If the summary is accurate on the specific thing, it is usually accurate overall.

The abstract and discussion pairing works because the abstract states the finding as the authors want it stated, and the discussion opening states what they think it means. Between them you can see whether the summary is calibrated.

Check the references the summary hands you as well as the claims. Helmy and colleagues note that newer tools have improved on factual accuracy but that fabricated information persists, and fake references and citations are the form that persists most. Their sixth rule is to refine the prompt, look at more than one output, and cross-verify every AI-generated insight against empirical evidence, domain expertise or established knowledge before using it (Helmy et al., 2025, pp. 2-3, 9). Their framing is guidance rather than a measured accuracy trial, so treat it as a working procedure, not as a number you can quote.

For anything you intend to cite, skip the check and read the relevant section. Verification of a summary is worth doing when the paper is peripheral. When it is central, the summary was never the right tool.