How do I screen hundreds of abstracts for a systematic review?
Screen in two passes. First pass on title and abstract only, deciding include or exclude with no middle category, erring towards include when unsure. Second pass on full text, where exclusions must be recorded with a reason. Trying to make the final decision from an abstract is what makes screening slow.
Updated
Screening is boring in a way that produces errors. Attention drops after about an hour, and the failure mode is not random: tired screeners start including things to avoid the decision, which pushes the work into the full-text stage where it costs far more. Work in blocks of forty-five minutes to an hour and stop even when it feels like you could continue.
The error rate is measurable, so plan for it. Wang et al. checked 329,332 abstract-screening decisions made by 85 reviewers across 25 systematic reviews covering nine clinical areas, and found that 10.76 percent of decisions were false inclusions or false exclusions, ranging from 5.76 percent to 21.11 percent depending on the area and the type of question (Wang et al., 2020, p. 4). Assume roughly a tenth of your own calls are wrong and build a process that catches them, rather than one that assumes a careful screener does not make them.
The two-pass structure works because the first pass is supposed to over-include. In the same sample, 18.07 percent of retrieved records passed abstract screening but only 5.48 percent survived full text and entered the review (Wang et al., 2020, p. 4). That gap is the cost of keeping uncertain records alive, and it is the cheaper error when missing a relevant study would damage the review. The authors do not test any particular threshold for deciding what counts as uncertain, so that judgement stays yours.
Deduplicate before you start. Exports from three databases will overlap heavily, and screening the same record three times wastes effort and produces inconsistent decisions on the same paper. Any reference manager will do this, and it is worth checking the results by hand, since near-duplicates with different metadata often survive automatic matching.
Order matters more than people expect. Screening in a random order rather than by database keeps you from calibrating differently on different sources. Screening the same records in the same order as your co-screener makes disagreements easier to discuss.
Keep a decision log from the first record, not from the point where you realize you need one. For every exclusion at full text you will need a reason, and the reasons have to be from a fixed list rather than freely written, or the counts in your report will not add up.
The log has to cover the process, not only the records. PRISMA 2020 asks you to report how many reviewers screened each record and each report, whether they worked independently, and what automation tools were used in the process (Page et al., 2021, p. 6). Note those three things at the start, because reconstructing them from memory at writing-up time is how single-screener projects end up describing themselves inaccurately.
If you are considering a screening tool, treat it as something to evaluate on your own records. Du et al. framed abstract screening as a binary classification problem and compared machine-learning and deep-learning models by precision, recall, accuracy, and F1-score on two corpora, one of 1,697 citations with 538 included and one of 2,865 citations with 711 included (Du et al., 2024, pp. 1, 3). Both corpora were clinical, covering HPV prevalence in head and neck squamous-cell carcinomas and pneumococcal disease in children, so the reported performance says little about a humanities or social-science corpus. Wang et al. make the opposite-facing point, that since human decisions already carry substantial error, a tool with a comparable error rate may be acceptable and save time (Wang et al., 2020, p. 6). They also warn that some of what looks like reviewer error may be incomplete abstract reporting instead (Wang et al., 2020, pp. 5-6).
What makes a good inclusion and exclusion criterion?
One that two people applying it to the same abstract reach the same answer on. Test yours on twenty records before starting properly. If agreement is poor, the criterion is vague rather than the screeners careless, and rewriting it now saves reconciling hundreds of disagreements later.
The words that cause trouble are the ones that sound precise and are not. "Adult", "recent", "large sample", "relevant setting" all need numbers or lists behind them. Write the criterion as something you could hand to a stranger.
Write exclusions as an ordered list too, since a record can fail several criteria and you need to report a single reason. Applying them in a fixed order makes that decision automatic instead of arbitrary.
What do I do when two screeners disagree?
Discuss the record against the written criterion, not against intuition, and either agree or refine the criterion and reapply it to everything already screened. A third person breaks genuine ties. Record how many disagreements arose, since that number tells a reader how reliable the screening was.
Refining mid-screen is allowed and often necessary. What is not allowed is refining silently. Note the change, the date, and the fact that earlier records were rescreened against the new wording, because a criterion that shifted halfway through without correction makes the whole set inconsistent.
If you are screening alone, which many doctoral students are, do not pretend otherwise. Screen a random ten percent twice, weeks apart, report the agreement between your own two passes, and state the single-screener limitation plainly.
Sources
- Wang et al., Error rates of human reviewers during abstract screening in systematic reviews, PLoS One (2020) · checked 6 August 2026
- Du et al., Machine learning models for abstract screening task - A systematic literature review application for health economics and outcome research, BMC Medical Research Methodology (2024) · checked 6 August 2026
- Page et al., The PRISMA 2020 statement: An updated guideline for reporting systematic reviews, PLoS Medicine (2021) · checked 6 August 2026