Launch bonus Yearly plans 20% 40% off
Blog
Engineering

GPT-6.1 Sol, GPT-6 Astra, or Claude Opus 5.5 for cited research?

We gave three models the same five research questions three times each and judged 540 cited claims blind. Opus cites the most sources but drops the sources' hedges more often. Sol stays as close to its sources as Astra at a fifth of the cost.

GZ

Guy Zana · Founder

October 7, 2026 · 9 min read

Share
GPT-6.1 Sol, GPT-6 Astra, or Claude Opus 5.5 for cited research?

Agent Bayes offers GPT-6.1 Sol, GPT-6 Astra, and Claude Opus 5.5, and the price gap between them is large. So we asked a plain question: on real research work, what does the extra money buy?

We gave each model the same five research questions over a library of 155 documents, three times each, and checked every sampled citation blind against the passage it points to. Then we put a hypothesis to each model in two longer conversations, and in one of them pushed back on its answer.

None of the three made up a claim. Across 540 judged claims, zero were unsupported by their cited passage.

The models differ in what they do around that shared baseline. Opus cites the most sources, almost twice as many per question, but it drops the sources' hedges far more often: about a dozen times in 180 claims, against once each for Sol and Astra. Sol and Astra are equally careful, and Astra costs five times as much.

Which model for which task

"Fully supported" in this table means a blind judge found every part of the sentence in the passage it cites. A sentence falls short when it adds a detail the passage does not state, even a true one, or turns the source's "probably" into a plain fact.

Costs are in credits, the unit Agent Bayes charges in. The $19 plan's 600 monthly credits cover about 27 research tasks like these on Sol, 7 on Opus, or 5 on Astra.

You are doingPickWhy
A first overview of a field: who argues what, from which papersClaude Opus 5.5It cites 11.6 different sources per question against 6.5 for the GPT models.
Sentences you will quote, submit, or build an argument onGPT-6.1 Sol92% of its sampled claims are fully supported, against 83% for Opus, at about a quarter of Opus's cost.
Testing your own hypothesis against the literatureSol or OpusIn two blind-judged conversations, Opus scored 34.5 of 35 and Sol 34. Astra scored 30.
Keeping cost downGPT-6.1 Sol22 credits per task. Opus used 83 and Astra 116.
Paying more for accuracyNot GPT-6 AstraIts 94% fully supported cannot be told apart from Sol's 92%, at 5.2 times the credits.

The rest of this post is the benchmark behind that table.

What we measured

The library holds 155 documents on Iron Age Levantine architecture: journal papers, excavation reports, and reference volumes, plus a set of methods papers. It is the library we use to tune the agent.

Five research questions stay fixed. Four are archaeological debates, such as the date of Megiddo's six-chambered gate and the origin of the proto-Aeolic capital. The fifth asks for a survey of quantitative methods for tracing art motifs between cultures.

For each question, the agent searches the library and adds 3 to 4 paragraphs of cited claims to a mindmap. A mindmap is the working structure of a research project in Agent Bayes, where every claim is linked to the passages that support it.

Every model answered each question in three rounds, on the same day, at effort Medium, the product default. That makes 15 runs per model and 45 in all. Three measurements follow.

  • Breadth: how many different documents a run cites, how many sources each paragraph combines, and how many of a fixed list of key papers it finds. The list holds 4 to 8 papers per question and was set before the runs.
  • Claim support, judged blind: from each run, up to 12 cited claims are sampled with the exact passages they cite. Model names are removed and the packets are shuffled. A judge marks each claim supported, partial, unsupported, or misattributed. Partial means the core is in the passage but one element is not.
  • Verifier score: the AI verifier that ships in the product reads the same claims with their passages and returns a score from 0 to 1, with notes on what does not match. The verifier is a GPT model and the judges are Claude models, so each family checks the other.

Opus cites more of the library

Mean of 15 runsSolAstraOpus
Different sources cited6.56.511.6
Sources combined in one paragraph3.13.14.2
Key papers found, of the list57%58%69%

Opus cited the most sources on every one of the five questions. Its least-cited run used 7 sources, where Sol's and Astra's used 3 and 4.

On the Megiddo gate debate, Opus cited 13.3 sources per run and found 5.7 of the 8 key papers. Sol cited 6.7 and found 4.7. Astra cited 7.0 and found 5.0.

More sources did not always mean more of the key papers. Opus found more of them in two of the three rounds. In the third round it found 57% against 60% for the GPT models, and on the motif survey it found fewer than both.

A paragraph that combines four sources compares them the way a literature review does. A paragraph built from one source at a time reads more like a list.

Sol and Astra add the least

Judged claims, 180 per modelSolAstraOpus
Fully supported92%94%83%
Partial151129
Unsupported000
Misattributed001

The gap between Opus and each GPT model is larger than run-to-run noise. If we shuffle whole runs between models at random, a gap this large appears 2.6% of the time against Sol and 0.3% against Astra. Sol and Astra cannot be told apart.

Most of Opus's partial claims fall into two groups.

  • It drops hedges. A passage says the rear rooms of a palace are "presumed to be domestic quarters". Opus writes that the rear wing "often held domestic rooms". Another passage says cult objects were redated, "or at least many of them". Opus redates all of them.
  • It adds details the passage does not state. It names "the Chicago expedition" where the passage says "the expedition", and it calls a pair of stone volutes "two-ton" where the passage mentions their "weight and dimension".

The second habit is not special to Opus. Sol adds outside details about as often, such as a year or a site name the passage leaves out.

The dropped hedge is what separates Opus: about 12 times in 180 claims, against once for Sol and once for Astra.

Added details may well be true, but a reader who opens the citation will not find them. A dropped hedge is worse: the sentence still points to the right page, and the page says something weaker.

The product verifier gives all three nearly the same score: 0.92 for Sol, 0.93 for Astra, and 0.92 for Opus. It does name the dropped hedge or the added detail in its notes on most of these claims, and it writes the most notes on Opus, but it lowers the score only a little. Read the verifier's notes, not only its score.

The one misattribution was Opus's. A review paper described another team's network of Cretan sites, and Opus credited the network to the review's author.

Sol and Opus argue best

The second test is two conversations. One covers the date and function of Building 338 at Megiddo, a structure read by some as a shrine and by others as a residence. The other covers what the proto-Aeolic capital's volutes depict.

Each conversation lays out the debate, follows up on it, and asks for a plain explanation. Then the user offers a hypothesis of their own and asks the model to assume it, list what must be true, weigh the sources for and against, and give a credence. The capital conversation adds a final turn that pushes back on the answer.

A blind judge scores the three transcripts side by side, with points for naming each position, weighing the evidence, stating where the model assumes against a cited scholar, and giving a calibrated credence.

Two conversations, 9 turnsSolAstraOpus
Rubric score, of 35343034.5
Credits for both conversations1481,045582

Sol and Opus are close. They tied on the capital conversation, and Opus led by half a point on Building 338. One judge scored each conversation.

On Building 338, Opus named the one physical fact that most limits the user's hypothesis, a shrine built on the building's own foundation walls, and said how unreasonable it would be to assume against it.

Astra lost most of its points on one turn, weighing the capital hypothesis condition by condition.

When pushed back on, all three moved or held their credence with a reason and said plainly where they were assuming against a cited scholar.

What the credits buy

Per research task, mean of 15SolAstraOpus
Credits2211683
Relative to Sol1x5.2x3.7x
Time to finish, median192 s217 s157 s

Opus is the fastest of the three and the second most expensive. Astra is the slowest and the most expensive. All three ran at the same time on the same server, so the times compare fairly with each other but are not a promise for a single run.

How to work with each

  • Use Opus to open a question. It cites more of the library and combines more sources in each paragraph. Before you keep a sentence, open its citation and check for a hedge the sentence lost.
  • Use Sol for anything you will quote, and as the default when cost matters. 92% of its claims are fully supported, at 22 credits per task.
  • Use Sol or Opus to test an idea of your own against the literature. Both weighed the sources, stated their assumptions, and moved their credence under pushback.
  • Skip Astra for this kind of work for now. We could not find a task here where it earns five times Sol's price.
  • Whatever model wrote the sentence, open the citation before you quote a number, a name, or a date. That is what the link is for.

Limits

One library, five questions, three runs per question, and 540 judged claims. The judging is single-pass, and the judges are Claude models. They ranked Claude's own Opus last on support, which argues against favoritism, and a GPT verifier checked the same claims.

Every model ran at effort Medium. Higher effort may change the picture, especially for Astra. Two results held in every round: Opus cited the most sources, and Opus had the lowest share of fully supported claims. The exact rates are descriptive.

If you compare models for literature work, where do you draw the line between a supported claim and one that quietly drops the source's "probably"?

New posts, straight to your inbox

No newsletter fluff, just an email when we publish something new.

Email me when a new post is published on the Agent Bayes blog. You can unsubscribe anytime. We'll first send a confirmation email, and we only use your details for this. See our Privacy Policy.

Enjoyed this? Share it.

GZ

Written by

Guy Zana · Founder

We are researchers and engineers building tools that help people reason over large bodies of literature without losing the thread back to the source.