Most AI research tools are not competing with each other
We spent a week comparing ten of them. The useful split is not which one is smartest. It is what corpus each tool reads, what its citations point at, and what you are left holding when you close the tab.

A researcher we spoke to last month had four tabs open: NotebookLM, Connected Papers, Elicit, and ChatGPT. All four describe themselves as AI research assistants. She was trying to work out which one to keep paying for.
That framing is the problem. Three of those four are not alternatives to each other in any meaningful sense. One builds a similarity graph and never opens a PDF. One answers questions about documents you upload and cannot search for a paper you do not already have. One screens thousands of papers into a table. Asking which is best is like asking whether a reference manager beats a highlighter.
We recently sat down and compared ten of these tools properly, because we were writing comparison pages and did not want to publish the usual grid where every competitor row is a red cross. What came out was more useful than a scorecard, so here it is.
The market is four jobs, not one category
Sorted by what they actually do rather than what they call themselves:
Discovery and citation mapping. Connected Papers, ResearchRabbit, Litmaps. These work on citation graphs and metadata. Connected Papers computes similarity from co-citation and bibliographic coupling, which is why it surfaces conceptually close papers that share none of your keywords. None of the three reads full text. They answer "what should I read".
Answer engines over an open index. Elicit, Consensus, scite.ai, SciSpace. These query a large public corpus and return structured findings. Elicit screens 5,000 papers into extraction tables on its Pro tier. Consensus reports how a field splits across yes, no, and possibly. scite classifies 1.2 billion citation statements as supporting, contrasting, or mentioning, which answers a question none of the others touch: was this finding later contradicted. They answer "what does the literature say".
Own-corpus workspaces. NotebookLM, Atlas, and Agent Bayes. These read only what you give them. They answer "what do my sources say".
General research agents. ChatGPT Deep Research. Reads the open web for tens of minutes and writes a long cited report. It answers almost anything, at whatever depth the public web supports.
Six of the ten are not substitutes for the workspace category at all. If you are choosing between Connected Papers and Agent Bayes, you have mis-framed the decision, and we say so on that page.
Four questions that actually separate them
Model quality is not one of them. Most of these products sit on the same frontier models, and the differences you feel come from architecture and scope.
What corpus does it read
This is the first fork and it drives almost everything downstream.
An open index gives you reach. Consensus searches more than 220 million papers, SciSpace around 280 million, Elicit more than 138 million. You will find work you did not know existed.
It also means a ranker chose your evidence. You did not set inclusion criteria before retrieval ran, and you did not check the journals. For orientation that is the right trade. For a chapter you will defend, the first question anyone asks is why these papers, and "the model picked them" is not an answer.
Own-corpus tools invert it. You decide what goes in, in advance, and the tool never leaves that boundary. The cost is real: you have to build the library first, and nothing will tell you about the paper you never downloaded.
Deep Research complicates this usefully. Since February 2026 you can connect it to MCP servers and restrict searches to trusted sites, so it is no longer accurate to call it unconstrained. What it still cannot do is authenticate into publisher platforms and open the paywalled articles your institution subscribes to. In most fields that is where the literature lives.
What does a citation point at
Every tool here says it cites its sources. They mean four different things.
Deep Research cites a URL. NotebookLM highlights the passage inside its own interface, which is good for checking an answer in the moment and does not give you the printed page number of the journal. Elicit and Consensus cite the paper. scite cites the sentence in the citing paper, which is a different kind of evidence again: reception rather than grounding.
The gap that matters is between a citation you can check in the tool and a citation you can put in a bibliography. Between "this paper supports the claim" and "page 114 of this paper supports the claim" there is a manual step where you open the article and go looking, repeated for every claim you keep.
What are you left holding
Ask a good question in most of these tools and you get an answer. Ask forty good questions over a term and you have forty answers.
A Connected Papers graph is regenerated from its seed. An Elicit table belongs to one query. A Deep Research report is the end of a run, and the next run starts from nothing, with no memory of what you established or rejected. NotebookLM notebooks persist, which is more than most, but there is no version history over an argument you are developing and no record of what changed when.
None of that is a defect. It is what those products are for. It becomes a problem only when the work runs for months and the thing you need at the end is a structure, not a stack of outputs.
What happens when your sources disagree
This is the one we care most about, and it is the least discussed.
Most tools resolve disagreement, because coherent prose is what a language model produces by default. Two of your sources conflict, and the answer has to pick a line, hedge, or average. Consensus is the interesting exception: its meter counts the split across the index, which tells you the field is contested without telling you why.
For a lot of questions, resolution is fine. For a literature review it is a loss, because the disagreement is frequently the finding. Methodological disputes do not survive being summarised into one sentence, and they do not fit in a table cell either.
Where we put Agent Bayes, and where we do not
Agent Bayes is an own-corpus workspace. Knowledge bases are libraries you assemble, and every retrieval stays inside them.
Each PDF runs through OCR that preserves multicolumn text, tables, figure captions, and footnotes, then semantic chunking, then distillation into structured bullet points. Those bullet points are always written in English regardless of the source language, so a German methods paper is searchable next to your English ones. A limitation buried in an appendix is as retrievable as an abstract.
Every claim the agent writes links to author, year, printed page range, and the exact chunk of source text, and clicking it opens the PDF at that page. When your sources conflict, the competing positions become sibling nodes with separate evidence chains rather than one smoothed statement, because the editing stage is built not to average them. And the map persists, with a rolling version stack, named snapshots, provenance on every change, and a replay of exactly what the agent altered on the last pass before you accept it.
There is also a step we have not seen elsewhere. Write a claim in your own words, attach the citations you believe support it, then select the node. Agent Bayes reads the full text behind each citation, scores how well the evidence supports what you actually wrote, and tightens phrasing that overstates. That check is the one everyone skips under deadline pressure, and the one a reviewer finds.
Here is what we do not do, and it is a longer list than a marketing post usually admits. No open-literature search, so if you have no library you have to start somewhere else. No alerts, so Litmaps and ResearchRabbit keep that job. No drafting, deliberately, because the work you defend in a viva is the thinking rather than the sentences. No reception analysis across the citation graph, which is scite's territory. No screening at Elicit's scale. No free tier. And indexing a corpus costs time and credits before you get your first answer, where an open-index tool answers in seconds from a cold start.
A workflow that uses several of them
The honest recommendation is not one subscription.
Map an unfamiliar area in Connected Papers or ResearchRabbit. Keep the library in Zotero, where it belongs. Index the papers that matter in Agent Bayes, which the Zotero plugin does in place with no re-uploading, and build the argument there with every claim carrying its page. Let Litmaps watch for new work. Before you submit, run the finished reference list through scite to catch anything retracted or heavily contested.
That sequence uses each tool for the thing it is genuinely good at, and none of them is doing another's job badly.
We wrote up all ten comparisons in more detail, including pricing checked on 3 August 2026 and a note on each page about what the other tool does better than we do. They are here. If you find something out of date, tell us and we will fix it.
New posts, straight to your inbox
No newsletter fluff, just an email when we publish something new.
Email me when a new post is published on the Agent Bayes blog. You can unsubscribe anytime. We'll first send a confirmation email, and we only use your details for this. See our Privacy Policy.