Can I use AI for a literature review?
Yes for the mechanical parts, and only with verification. AI is useful for screening, extracting findings from papers you supply, and drafting summaries you then check. It is not a substitute for reading the sources you cite, and no journal or examiner accepts an unverified reference.
Updated
The question is usually asked as though it had one answer, and it has several, because a literature review is not one task. Searching, screening, extracting, judging, structuring, and writing are different activities with different risks, and AI is genuinely useful for some and dangerous for others.
The dividing line is whether the tool has the source in front of it. A model asked to recall literature from memory produces references that do not exist, at rates measured in published studies at anywhere from a fifth to two thirds of the output. A tool that reads documents you supplied and points at the passage it used is doing something different and much safer, though still worth checking.
The second dividing line is whether a human judgment is being replaced or supported. Extracting the sample size from a paper is a mechanical task and a machine does it well. Deciding whether that sample size makes the finding credible is a judgment your examiner will hold you to, and it has to be yours.
The practical upshot is a workflow, not a verdict. Use the tools for volume work, keep the reading for the papers your argument stands on, and verify everything that ends up in your reference list. That is more or less how experienced researchers use research assistants too, and the accountability rules are the same.
Before you rely on any of it, check the rules that govern your specific submission. Smith et al. advise consulting the target journal's current policy, because journals differ on which uses they accept and revise their guidance over time, and they recommend disclosing the system you used, the methods and tools, its role, its limitations, and its possible effects on the research (Smith et al., 2024, p. 3). They also warn against unattributed AI-generated content and against putting confidential material into these systems (Smith et al., 2024, p. 2).
Which parts of the process can AI genuinely help with?
Four. Screening abstracts against criteria you wrote. Pulling a specific finding out of a paper you already have. Translating a passage in a language you do not read. And drafting a first version of a paragraph from claims you supplied. All four produce output you check.
What these have in common is that the source is available and the check is cheap. You can confirm a screening decision by reading the abstract, and an extracted finding by opening the page. The work saved is real and the verification cost is small.
Screening is the part with the most published evidence behind it. Du et al. reviewed machine-learning models for abstract screening in health economics and outcomes research, a field where a single systematic review is estimated to cost an average of $141,194.80, and where major pharmaceutical companies run about 23.36 reviews a year and major academic centers about 177.32 (Du et al., 2024, p. 2). Their comparison found that transformer models consistently outperformed conventional machine-learning models when given only titles and abstracts, but that once all citation features were included the conventional models performed comparably at lower computational cost, which led the authors to name XGBoost and support-vector machines as workable screening choices (Du et al., 2024, p. 6). The same authors are cautious about the newer chat models, flagging bias and hallucination and calling for further evaluation of how such systems handle free-text eligibility criteria and explain their screening decisions (Du et al., 2024, p. 7).
Anything where verification would cost as much as doing the task yourself is not worth delegating. That includes asking for a summary of a literature you have not assembled, which is exactly the use with the worst error profile.
Which parts have to stay human?
Deciding what the question is, judging whether a study is any good, choosing which findings matter, and taking responsibility for every claim in the final text. Publishers are unanimous that accountability cannot be delegated, and a claim you did not verify is one you cannot defend.
Quality judgment is the one people most often try to hand over. A model can describe a study's design, and it cannot tell you whether that design is adequate for the inference your chapter needs, because that depends on what you are arguing.
Pautasso, writing from roughly 25 reviews produced during a PhD and postdoctoral career, treats a review as a compound task that combines searching, evaluating sources, synthesizing evidence, critical thinking, paraphrasing, and citation, and expects the finished text to identify methodological problems, areas of debate, and outstanding research questions rather than to collect summaries (Pautasso, 2013, pp. 1-3). That procedural advice is what keeps that work defensible under scrutiny. Record your search terms so the search can be replicated, track the papers whose PDFs you could not get, define exclusion criteria early, use a reference-management system, and search for existing reviews as well as primary studies (Pautasso, 2013, p. 1). Smith et al. add the verification half of the same discipline, which is to cross-reference anything a model tells you against reliable sources and to corroborate it with empirical evidence or expert judgment (Smith et al., 2024, pp. 1-2).
Structure is a second. The order of your review chapter is your argument, and an ordering generated from surface similarity between papers produces the topic-shaped review that supervisors ask you to rewrite.
Sources
- COPE position statement on authorship and AI tools (2023) · checked 6 August 2026
- Smith et al., Ten simple rules for using large language models in science, version 1.0, PLoS Computational Biology (2024) · checked 6 August 2026
- Du et al., Machine learning models for abstract screening task - A systematic literature review application for health economics and outcome research, BMC Medical Research Methodology (2024) · checked 6 August 2026
- Pautasso, Ten Simple Rules for Writing a Literature Review, PLoS Computational Biology (2013) · checked 6 August 2026