Why can I never find my PDF highlights again?

Because a highlight lives inside one file and your question spans the library. You want every passage about a claim, and highlights are organized by document. The colored text is also a record that something mattered, without any record of why, which is the part you needed.

Updated

Highlighting works well for one paper and fails for a hundred, and the reason is structural rather than a matter of discipline. Annotations are stored per document. Your writing question is never per document. When you sit down to write about measurement error, you want every passage about measurement error across forty papers, and there is no view that shows you that.

The second failure is informational. A highlight records that a sentence seemed important at the time. It does not record why, what you were thinking, or which of your claims it supports. Those three things are the entire value, and they are the ones that evaporate.

This is a known failure of reading habits, not a personal defect. Méndez states it plainly, "The more you read, the more you forget," and the recommended fix is a personal review that stores the basic information or main message of each primary paper in a spreadsheet, text document, or reference manager, so that both confirmatory and negative evidence stays retrievable when you write an introduction or a discussion (Méndez, 2018, p. 3). That record sits outside the PDF on purpose, because the PDF cannot answer a question that spans your library.

There is also a volume problem. Enthusiastic highlighting produces papers where a third of the text is colored, which conveys nothing. When everything is marked, the marking has stopped carrying information, and rereading the highlights is only slightly faster than rereading the paper.

None of this means stop highlighting. It means treat the highlight as a bookmark with a short shelf life, and convert it into something durable before the context is gone.

How should I annotate a paper instead?

Highlight while reading, then spend five minutes afterwards converting the highlights into two or three written claims with page numbers. The highlight becomes a pointer rather than the record. Converting immediately matters, because the reason a passage mattered fades within days.

Colour coding helps if the scheme is small and fixed. Three colours, meaning something like evidence for, evidence against, and method worth borrowing, is usable. Seven colours whose meanings you invented over time are not, and you will not remember the scheme by year two.

Carey, Steiner, and Petri describe one workflow that separates the passes: read the article a first time for general exposure, a second time for understanding, and a third time for taking notes (Carey et al., 2020, p. 5). Their annotation advice for that third pass is to mark your questions, look up terms you do not know, highlight the important statements, and draw the link between each figure and the interpretation given in the discussion (Carey et al., 2020, p. 5). They also say the format can be paper or an equivalent digital method, so they do not treat PDF annotation as the weaker option. What matters is whether the mark records a question or a connection rather than only a location.

The conversion step is where the actual reading happens. Restating a passage as a claim forces you to decide what it establishes, and that decision is the thing you were reading for.

How do I make highlights searchable across a whole library?

Extract them into one place. Most readers can export annotations, and a single file or table of every extracted passage with its source and page is searchable in a way that per-document highlights are not. The extraction is worth automating, since doing it by hand stops within a month.

Once extracted, tag the passages by claim rather than by paper. That is the move that turns an archive into a working set, because it lets you pull every passage bearing on one argument in a single filter.

Whatever tool holds the extracted passages, make sure it keeps the page number and the source with each one. A searchable pile of quotations with no locators is a citation problem waiting to happen at exactly the point where you can least afford one, which is the week before submission.

Tag with your own words. Méndez recommends running a reference manager, updating it daily with what you have just read, keeping unread papers in a separate "to read" folder, and assigning personalized keywords instead of importing the publisher's keywords (Méndez, 2018, p. 3). Publisher keywords describe the paper as the field sees it, and your search a year later will use the vocabulary of your own chapters.

The storage side has its own requirements. Vandendorpe and colleagues write about electronic lab notebooks rather than PDF annotation, so treat this as an infrastructure parallel and not as direct evidence about highlighting, but their list is the right checklist for an extraction file: capture the record early, attach metadata, use persistent identifiers, link out to reference-management databases and repositories, then back up, publish, and archive (Vandendorpe et al., 2024, p. 6). Applied to your reading, that means one canonical copy of each PDF, one bibliographic record it links to, and one note per highlighted claim that says what the claim is and why you kept it.

Keep the original PDFs alongside the extraction. Exported annotations lose their surrounding context, and there will be a passage whose meaning depends on the sentence before it.