Launch Bonus Yearly plans offer a 20% 40% discount bonus for a limited time. See plans

Blog
Engineering

Gemini 3.8 Flash, GPT-5.4, or GPT-5.6 Terra for cited research synthesis?

We benchmarked Gemini 3.8 Flash, GPT-5.4, and GPT-5.6 Terra on citation-grounded synthesis over 150 indexed archaeology papers. Gemini weaves twice as many sources into each paragraph as Terra, and Terra's sentences stay closest to what the cited passage states. GPT-5.4 sits between them at two to three times the cost.

GZ

Guy Zana · Founder

September 8, 2026 · 12 min read

Share
Gemini 3.8 Flash, GPT-5.4, or GPT-5.6 Terra for cited research synthesis?

Agent Bayes now offers Gemini 3.8 Flash next to the GPT models, so the question is which one to pick. We ran Gemini 3.8 Flash, GPT-5.6 Terra, and GPT-5.4 through the synthesis benchmark we use to tune the agent, over a corpus of 150 indexed papers, on the same day under the same judging protocol.

The benchmark asks each model to do the two things the agent does: answer a question in chat with cited prose, and build a citation-backed mindmap. The answer is a tradeoff rather than a winner.

Gemini reads wider: twice the sources per paragraph, four times the alternation between them, and mindmaps with about half again as many claim nodes.

Terra reads closer. The fidelity numbers in this post come from the AI verifier that ships in the product, which scores each cited sentence against its cited passages from 0 to 1. Terra averages 0.96 in chat answers and 0.98 in mindmap nodes, against 0.89 and 0.87 for Gemini, because Gemini tends to add a label or a qualifier that the cited passage does not state.

GPT-5.4 does both reasonably well and costs two to three times as much per run. It synthesizes more than Terra and less than Gemini. On the verifier it matches Terra's fidelity in mindmap nodes and trails it in chat answers.

In the sentences we traced, some of Gemini's additions came from passages it retrieved but did not cite, and others were in nothing it read during the run: it appears to be hard for Gemini to keep its own knowledge of the field out of a cited sentence.

Across 200 claims judged blind against their cited passages, one was misattributed to the wrong author and none were unsupported. The difference between the models is what each adds on top of a grounded core, and that decides the task.

Which model for which task

Fidelity figures in this table are the share of cited claims the product's AI verifier scores fully supported, 1.0 on its scale. A claim loses points when the sentence adds a qualifier the cited passage does not contain, such as calling a "shrine" the "primary" shrine or a "palace" the "governor's residence".

Most such additions are ordinary synthesis, a detail the field would accept and that is often true. The point is that a reader who opens the citation will not find it there.

You are doingPickWhy
A first pass over a new corpus, mapping the debates and who holds which positionGemini 3.8 FlashIt pulls 5 to 11 sources into a mindmap where Terra pulls 4 to 7, builds 19 to 22 claim nodes against 13 to 15, and switches between sources four times as often inside a paragraph. Its claims are mostly right with one added specific each, and none are unsupported.
Nodes or paragraphs you will quote, submit, or build an argument onGPT-5.6 Terra69% of its cited chat sentences and 85% of its nodes contain no such addition. For Gemini the figures are 50% and 33%, for GPT-5.4 44% and 85%.
A mindmap you will both build on and quote from, when cost is not the constraintGPT-5.4Between the two on synthesis, 2.8 sources per paragraph. On fidelity it matches Terra in mindmap nodes, 85%, and sits below it in chat, 44% against 69%. About 37 credits per chat run and 64 per mindmap.
Answering a question in chat to orient yourselfGemini or TerraGemini gives about 570 words with 25 citations, Terra about 420 words with 11, for about the same credits.
Checking a Gemini mindmap before you rely on itRun AI verification on the nodesThe verifier lists each added specific it finds: about one per Gemini claim, against 0.2 to 0.4 for Terra and 0.2 to 0.7 for GPT-5.4.
Keeping cost downGemini or TerraGemini's list price per token is about a third of Terra's, but it writes more and calls more tools, so measured credits per run are close: Gemini 15 and Terra 13 in chat, 22 and 24 in research mode. GPT-5.4 spends 37 and 64.

The rest of this post is the benchmark behind that table.

What we measured

The benchmark runs the mindmap agent against a fixed corpus of 150 papers on Iron Age Levantine architecture. Three prompts are held constant: the dating debate over Megiddo's six-chambered gate, the shrine-or-residence debate over Building 338, and the origin of the volute capital.

Each prompt runs in two modes. In ask mode the agent writes three to four paragraphs of cited prose in chat. In research mode it adds cited nodes to a mindmap.

The mindmap is the working structure of a research project in Agent Bayes: each node is one claim linked to the passages that support it, consecutive claim nodes under a parent form a paragraph, and parents above them act as section titles. Research mode therefore measures how much of that structure a model builds from the corpus, and how faithful each node is to the passage it cites.

Settings are fixed. A cell is one prompt in one mode, so there are six cells per model, and each cell is one completed run.

Three measurements follow.

  • Structure, computed by a lint pass over the archived output:
    • Mixing: how many distinct sources a cited paragraph draws on. A synthesis measure.
    • Interleave: how often a paragraph switches between its sources. A synthesis measure.
    • Adjacency: whether each citation sits inside the sentence it supports. A traceability measure.
    • Pooled paragraphs: citations gathered at paragraph end instead of in their sentences. A traceability measure.
    • Band rate: whether each mindmap node stays inside the 80 to 220 character band the editor is instructed to keep.
  • Claim support, judged blind. The scorer samples up to twelve claims per cell, fewer when the output has fewer cited claims, attaches the exact chunks the agent retrieved for each citation, and strips the model and mode labels. A judge reads only that packet and assigns one of four verdicts: supported, partial (the core is grounded but one material element is not), unsupported, or misattributed. A wrong or added specific is at best partial.
  • Verifier score, from the AI verifier that ships in the product. It reads a sentence and its cited passages and returns a graded 0 to 1 support score with a list of discrepancies. It uses a different model and prompt than the judge and never sees the judge's verdicts.

Gemini synthesizes more sources

Ask mode, mean of 3 promptsGemini 3.8 FlashGPT-5.4GPT-5.6 Terra
Distinct sources per cited paragraph (mixing)3.82.82.0
Source alternations per multi-cited paragraph (interleave)4.12.61.1
Citation inside its own sentence (adjacency)0.880.830.71
Paragraphs with citations pooled at the end000
Words per citation243438
Answer length in words568608424
Research mode, per prompt (gate, Building 338, capitals) unless marked meanGemini 3.8 FlashGPT-5.4GPT-5.6 Terra
Nodes inside the 80 to 220 character band, mean1.001.001.00
Claim nodes per mindmap19, 21, 2220, 17, 1715, 13, 14
Distinct sources cited11, 10, 58, 5, 57, 5, 4
Sources per cited paragraph of nodes (mixing), mean3.52.41.8

All three models build the same shape of mindmap: a title node, three to four paragraph parents, and claim nodes beneath them, every node inside the size band. Gemini fills that shape with more claims from more sources, GPT-5.4 is close behind on node count, and Terra is the most compact.

Gemini weaves twice as many sources into each paragraph as Terra, switches between them four times as often, attaches twice as many citations, and builds mindmaps with about half again as many claim nodes. GPT-5.4 lands between the two on every synthesis measure.

Synthesis is the harder part of this work. Keeping a citation next to its sentence is a formatting habit. Drawing four sources into one paragraph and alternating between them is what a literature review exists to do.

A paragraph of single-cited claims in a row scores 1.0 on adjacency and is a list, not a synthesis. Mixing and interleave are what separate the two, and Gemini leads there.

Terra stays inside the passage

Model, modeClaims judgedSupportedPartialUnsupportedMisattributed
Gemini 3.8 Flash, ask36231201
GPT-5.4, ask3634200
GPT-5.6 Terra, ask3228400
Gemini 3.8 Flash, research36251100
GPT-5.4, research3330300
GPT-5.6 Terra, research2726100
Model, modeMean verifier scoreShare scored 1.0Discrepancies per claim
Gemini 3.8 Flash, ask0.8950%0.9
GPT-5.4, ask0.9044%0.7
GPT-5.6 Terra, ask0.9669%0.4
Gemini 3.8 Flash, research0.8733%1.0
GPT-5.4, research0.9885%0.2
GPT-5.6 Terra, research0.9885%0.2

The two measurements agree on the ranking and on what the gaps are, with one exception: in chat, the judge puts GPT-5.4 first and the verifier puts Terra first. Both place Gemini last. Gemini's typical cited sentence is mostly right with one gap. Terra's and GPT-5.4's typical cited sentences have no gap.

Gemini's gaps are almost all of one kind.

  • A passage says "other finds". The sentence says "ritual artifacts".
  • A passage says an author "suggests" a date. The sentence calls the evidence "decisive".
  • A passage says the excavators read a structure as a "palace/fortress". The sentence says "secular administrative" and drops the fortress.
  • A passage says the site overlooks "its surroundings". The sentence says "the Jezreel Valley".

Most of these additions are the kind of detail a well-read archaeologist would supply from memory, and several may be true. That is what makes the pattern matter for anyone who will quote the sentence. The citation is real and the passage is on topic, and a reader who opens it finds that the specific they were about to quote is not there.

The one misattribution is Gemini's: views on where the capitals stood, held by other scholars, pinned to Shiloh.

Where the additions come from

The agent retrieves several passages per page, and the judge sees only the passages a sentence cites. So a flagged specific could come from a passage the model read but did not cite, or from outside the run entirely.

In a separate tracing exercise on earlier runs of the same prompts, we searched every tool result Gemini and Terra received during a run (retrieval, page reads, citation reads) for the specifics the judges had flagged in that run. GPT-5.4 was not part of that check.

In the sentences we traced, every Terra specific was in a passage it had read. It cites the wrong neighbour occasionally, and nothing it had not seen.

Gemini's split in half. Half were in retrieved passages it had not cited. The other half, "fortified residence", "primary shrine", the Low Chronology and its radiocarbon and seriation methods, were in nothing the model read that turn.

That half is the kind of gap that remains in the tables above. It is why the agent's prompt tells every role to cite only what the passage states and to cite each passage it relies on, and why we recommend Terra when claim fidelity matters.

The judge is part of the measurement

Under this rubric, two judges given the same outputs can differ by about twenty points on the absolute supported rate. The boundary between "supported" and "partial" is set by the judge, not the rubric.

Absolute supported rates from an LLM judge, the share of claims it marks fully supported, are not comparable across judging passes. The gap between models under one judge, and a graded verifier score, are. Every comparison in this post comes from one day, one judging protocol, and one verifier pass over all three models.

If you compare models with an LLM judge, calibrate on the same outputs with the same judge before believing any absolute number.

How to work with each

  • Use Gemini to open a corpus: broad mindmaps, many sources, claims already split to node size. Then run AI verification on the nodes you will build on and read the discrepancy list. Expect the verifier to flag about two nodes in three, usually one added label or qualifier each on a claim whose core stays grounded, and expect the fix to be a deletion rather than a rewrite.
  • Use Terra for anything you will quote. 69% of its cited chat sentences and 85% of its nodes contain nothing beyond the passage, and most of the remaining gaps are single qualifiers.
  • Use GPT-5.4 when you want one model for both the mindmap and the prose and the credits are not the constraint. It reads more sources than Terra, its nodes are as faithful on the verifier and its chat claims are ahead of Terra's on the blind judge, and it costs two to three times as much per run.
  • Use GPT-5.6 Luna when cost matters more than either. It was not in this comparison, but it remains the cheapest model in the picker and the one we recommend trying first for ordinary questions.
  • Whatever model wrote the sentence, open the citation before you quote a number, a date, or a stratum. That is what the link is for.

Limits

Three prompts on one corpus, one completed run per cell, 72 Gemini claims, 69 GPT-5.4 claims, and 59 Terra claims judged. Judging is single-pass. The direction of every result above survives that design. The exact rates do not, and we report them as descriptive.

If you run citation-grounded synthesis with a scoring pipeline of your own, we would like to know where your rubric draws the line between a supported claim and a claim that says ten percent more than its source.

New posts, straight to your inbox

No newsletter fluff, just an email when we publish something new.

Email me when a new post is published on the Agent Bayes blog. You can unsubscribe anytime. We'll first send a confirmation email, and we only use your details for this. See our Privacy Policy.

Enjoyed this? Share it.

GZ

Written by

Guy Zana · Founder

We are researchers and engineers building tools that help people reason over large bodies of literature without losing the thread back to the source.