Multilingual corpora
Your French, German and Japanese sources, searchable with the rest
Half your literature is not in English, and the tools behave as though that is an edge case. Indexing distills every passage into structured English bullet points whatever the source language, so one semantic query reaches the whole corpus and every citation still points at the original.
Free to start, no card. 200 credits when you verify your email.
The reason you are here
The most important paper in your field was published in 1974, in German
So it sits outside whatever tool you are using. You search in English, get English results, and quietly build an argument on the subset of your literature that happens to be in the language your software prefers.
Meanwhile the scanned monograph with two columns and dense footnotes defeats the extraction entirely, and the passage you actually needed is in one of those footnotes.
A corpus that only works in one language is not your corpus. It is the part of your corpus your tooling could read.
How the language barrier goes away
Indexed in English, cited in the original
The retrieval layer speaks one language. The scholarship stays in its own.
Index the paper in its own language
Nothing is translated up front and nothing is discarded. OCR reads the page as it was typeset, including the parts that are not prose.
It becomes searchable in English
Passages are segmented and distilled into structured English bullet points that carry enough context to stand on their own, so a single query crosses every language in your library.
Cite it in the original
The claim in your mindmap cites the source as published, with author, year and printed page range. Ask for an inline translation of the passage when you need to quote it.
In the product
One query, every language in your library
Semantic search returns the extracted bullet points, the source passage, the page range and the attribution, whatever language the paper was written in.
What you can rely on
What survives the pipeline
Summaries always in English
Each indexed passage is distilled into structured bullet points written in English regardless of the source language, so one semantic query reaches your whole corpus at once.
Layout survives the OCR
Multicolumn text, tables, figure captions, footnotes and headers are preserved rather than flattened, which matters more in older and non-English typesetting.
Inline translation while reading
Request a translation of any passage in the reader. It preserves academic terminology and the author's intent rather than smoothing both away.
Real page numbers, any language
Page references use the actual printed pages of the journal or book, so a citation to a German monograph drops into your bibliography as cleanly as an English one.
What Agent Bayes does not do
- It does not produce a publication-quality translation of a whole paper. Inline translation is for reading and quoting a passage, not for replacing a translator.
- It does not write your mindmap in another language. The working language of the map and the extracted summaries is English.
- It cannot rescue a scan that is genuinely illegible. Good OCR is not magic, and a bad photocopy stays a bad photocopy.
The second one is worth knowing before you start. If you need the map itself in another language, this is not the tool for that today.
Before you ask
Questions researchers ask first
Which languages does this cover?
Indexing writes its structured summaries in English whatever the source language. French, German and Japanese are the examples we use because they came up first, and the mechanism is not a per-language feature list.
Does it translate the whole paper?
No, and it does not need to. Indexing writes English summaries of each passage for retrieval, and the full text stays as published. Translation happens on request, per passage, while you read.
Will the terminology survive translation?
That is what the inline translation is built for. It preserves academic terminology and the author's intent instead of producing fluent prose that has lost the distinction you needed.
Can I mix languages in one knowledge base?
Yes, and that is the point. One knowledge base can hold papers in several languages, and retrieval treats them as one corpus.
What does it cost to try?
The free plan holds up to five indexed files and gives you 200 one-time credits once you verify your email. Indexing an average 20 to 30 page PDF costs about 20 credits.
Related
If you are also wrestling with
Create your account
Index the paper your other tools could not read
Start with five papers, and make at least one of them the awkward multilingual scan you have been working around.
- No credit card needed
- One corpus, several languages
- Citations point at the original