How do I build a search strategy for a systematic review?

Break your question into concepts, list every term each concept is described by including author variants and controlled vocabulary, combine terms within a concept with OR, then combine concepts with AND. Test the result against papers you already know should appear, and widen it until they all do.

Updated

The known-item test is the part most people skip and the part that catches the most errors. Before you trust a strategy, assemble five to ten papers you are certain belong in the review, from your own reading or from a supervisor. Run your search. If any of the ten are missing, the strategy has a hole, and finding out now costs an afternoon instead of finding out at screening and costing a month.

When a known paper is missing, look at how it describes itself. Usually it uses a term you did not list, because fields rename things and authors are inconsistent. Add the term, rerun, and check the others are still there. Two or three rounds of this is normal.

There is a second test that pairs with the known-item check. Broaden the strategy, then run the wider version against the original as a set difference, "Strategy 2 NOT Strategy 1", and screen only the unique records to see whether the added terms recover anything relevant (Bramer et al., 2018, p. 536). After screening, check the reverse direction. If a paper you found by reference-list checking was in the database all along, read its title, abstract, and indexing terms to work out which word your strategy missed (Bramer et al., 2018, p. 538).

The second common error is over-specification. A strategy that combines four concepts with AND returns almost nothing, and the temptation is to conclude the literature is thin. Drop the weakest concept. In most questions, two well-built concept blocks find more of the relevant work than four narrow ones, and the extra irrelevant hits are cheaper to screen out than missing papers are to recover.

Controlled vocabulary is worth the effort where it exists. Subject headings catch papers whose titles and abstracts never use your keyword, and combining headings with free-text terms is standard practice rather than belt and braces.

Bramer et al. report that conventional systematic review search development can take 100 hours or more, and that the whole task is a trade-off between sensitivity and specificity (Bramer et al., 2018, p. 531). Their method builds the strategy in order, starting with the most important and specific concept and adding broader ones until the review team agrees the screening volume is manageable, and it compares the records retrieved by thesaurus terms against those retrieved by free-text synonyms to identify candidate terms you have not listed yet (Bramer et al., 2018, p. 533). They also treat the strategy as something two other people check, an information specialist working through the PRESS checklist for Boolean operators, field codes, and syntax, and a researcher who knows the topic checking the terms themselves (Bramer et al., 2018, p. 537).

Pick two or three subject databases that cover your field, add one multidisciplinary index, and add a grey literature source if unpublished work matters to your question. Ask a subject librarian, because database coverage varies by discipline in ways that are not obvious from the outside.

Google Scholar is useful for finding known items and for citation chaining, and weak as a primary systematic search, since its result set is not stable, not fully exportable, and not reproducible from the same query.

Haddaway et al. compared Google Scholar title searches against Web of Science topic searches across seven reviews, and the gap runs in both directions. A peatland greenhouse gas review returned 1,120 records in Google Scholar against 4,151 in Web of Science, while an oil palm biodiversity review returned 968 against 290 (Haddaway et al., 2015, p. 5). In their grey literature case study, Google Scholar found 61 of 84 known articles and missed 23, and a normal systematic review string did not retrieve even those 61 (Haddaway et al., 2015, p. 14). That is the case for running it alongside subject databases rather than in place of them. The same study found the most grey literature around page 35 of the title results, which counts against the habit of screening only the first 50 to 100 hits (Haddaway et al., 2015, p. 12). Their evidence is one case study and the ranking algorithm is undisclosed, so treat page 35 as a reason to state and justify your stopping point, not as a fixed depth to copy.

Whatever you choose, name the choice and justify it in one sentence. A reviewer's objection is almost never "you searched too few databases", it is "you did not say why these ones".

How do I document a search so someone else can reproduce it?

Record, per database, the exact string as entered, any limits or filters applied, the interface used, the date run, and the number of results. Save the strings in a file rather than reconstructing them later. Search syntax differs between platforms, so a string is only reproducible with its platform.

Reconstruction after the fact is where reproducibility quietly dies. Six months on, nobody remembers whether the truncation was two characters or three, or whether the date limit was applied at search or at export, and the strategy in the appendix stops matching the one that produced the results.

PRISMA 2020 sets the minimum you have to report. It asks for the full search strategy for every database, register, and website, including any filters and limits used, and for the name of each information source with the date it was last searched (Page et al., 2021, p. 7). Its flow diagram asks separately for records identified through websites, organisations, and citation searching, for reports you sought but could not retrieve, and for the reasons reports were excluded (Page et al., 2021, p. 10). Write the log against that list while you are searching, because every one of those fields is easier to record than to reconstruct.

Two moves complete the record. Note the citation-chaining you did by hand, forwards and backwards, since that is part of how the set was assembled. And keep the raw export files, so the numbers in your flow diagram can be checked against something rather than trusted.