TranslationBench 1.0 · Evidence for editorial teams
Preserve the meaning. Across the whole manuscript.
ProsaBridge combines greater fidelity to the original with fluent, readable prose. In a blind comparison of a 67,127-word scholarly book, ProsaBridge reduced weighted error points by 59% compared with the same AI model prompted directly.
- Book comparison
- Same AI model, with ProsaBridge and prompted directly
- Manuscript length
- 67,127 words · 61 sections
- Evaluation
- Independent blind scoring, identical criteria across both workflows
- Published evidence
- Book: DE → EN · short texts: EN → DE
The ProsaBridge advantage
- Fewer error points across the whole book
- −59%
- Across the complete book, the ProsaBridge translation received fewer weighted error points than the same AI model used directly. Independent evaluators assessed both versions without knowing how they were produced.
- Fewer points for meaning errors
- −64%
- ProsaBridge cut weighted error points for changes in meaning, incorrect terms, and omissions by nearly two thirds, while keeping the translation fluent and readable.
Book study · September 16, 2026
A whole book, translated more faithfully with ProsaBridge
−59%
fewer weighted error points with ProsaBridge than when the same model was prompted directly. The accuracy error score fell by 64%, combining greater fidelity to the original with fluent, readable prose.
ProsaBridge brought greater fidelity to a 67,127-word scholarly book translated from German into English. We compared it with the same AI model prompted directly with the complete manuscript. Two independent evaluators checked every section against the original, without knowing which workflow produced either translation.
The book comparison shows the value ProsaBridge brings to a complete manuscript. You can explore the full results and evaluation method below.
What you get with ProsaBridge
Set the terminology and style before translation begins, carry the manuscript’s context into every section, and review the finished translation in Word.
- Your terminology. Your style.
- Review terminology and style instructions, and resolve questions about the source text before translation begins.
- Context for every section
- Every section is translated with the manuscript summary, glossary, and style instructions, for consistent terminology and tone throughout.
- A document ready for review
- Work in Word and receive a translated .docx that retains the document’s structure and formatting.
Keep your team in control, from the first terminology decisions to the final review in Word.
See plans and pricingMethod and complete results
The book study and short-text benchmark are separate experiments with different texts, language directions and model settings. This appendix contains the scoring method, all published results and downloadable records.
Book study: scores from both evaluators
Book study · September 16, 2026 · One book · 61 sections · 67,127 words · German → English · blind evaluation · text confidential
Book scores are weighted error points per 100 source words, pooled across all sections and both judges. Minor errors count 1 point, major errors 5, and critical errors 10. Lower scores mean fewer weighted error points. Accuracy covers changed meaning, incorrect terms, and omissions; readability covers how the translation reads. These scores do not count incorrect words or predict how often an editor will need to make a correction.
| Entry | Error score | Score | Accuracy | Readability | Fable 5.1 medium | GPT-6 Astra medium |
|---|---|---|---|---|---|---|
| GPT-6 Astra inside ProsaBridge | 0.08 | 0.06 | 0.02 | 0.12 | 0.04 | |
| GPT-6 Astra prompted directly | 0.20 | 0.18 | 0.02 | 0.28 | 0.12 | |
AccuracyReadability
How to read the short-text results
An entry is one AI model used in one of two ways, or one translation service. There are 23 entries.
- With ProsaBridge
- The AI model worked inside ProsaBridge. It read the whole text first, prepared a translation plan, and translated using that plan.
- Model on its own
- The same AI model received the text once and translated it in a single step.
- Translation service
- DeepL, Google Translate, or Amazon Translate, each translating the whole document through its own service.
- Error score
- Each judge assigns the errors they find to Multidimensional Quality Metrics (MQM) categories. In this benchmark, a minor error counts 1 point, a major error 5, and a critical error 10. The error score is the sum of penalty points per 100 words, averaged over both judges and all twelve texts. Lower is better.
- The score reflects both the number and severity of errors: one major error in meaning carries the same weight as five minor errors. It is not a count of errors or a percentage of incorrect words, and it does not predict how often an editor will need to make a correction.
- Accuracy
- The translation departs from the original meaning: an incorrect fact, an inaccurate term, or an omitted sentence.
- Readability
- The translation conveys the meaning but reads poorly: awkward phrasing, grammatical mistakes, or inappropriate tone.
Complete short-text results
All 23 published entries from twelve English-to-German texts, with model settings, separate judge scores and relative compute costs. The results describe the configurations tested in this study. Compute costs are model usage charges expressed as multiples of the lowest measured cost; they exclude document handling and editorial work and are not ProsaBridge prices.
Click a column to sort · sorted by Score
| Rank | Entry | ||||||
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra†With ProsaBridge · high/xhigh/xhigh | 0.14 | 0.06 | 0.08 | 0.27 | 0.02 | 501.9× |
| 2 | GPT-5.6 SolWith ProsaBridge · high/xhigh/xhigh | 0.25 | 0.07 | 0.18 | 0.34 | 0.16 | 158.9× |
| 3 | GPT-6 Astra†Model on its own · xhigh | 0.25 | 0.14 | 0.11 | 0.31 | 0.20 | 184.2× |
| 4 | GPT-6 Astra†With ProsaBridge · medium/medium/medium | 0.25 | 0.09 | 0.16 | 0.42 | 0.09 | 109.5× |
| 5 | GPT-5.6 SolModel on its own · xhigh | 0.40 | 0.20 | 0.21 | 0.47 | 0.34 | 84.9× |
| 6 | GPT-6 Astra†Model on its own · medium | 0.44 | 0.12 | 0.32 | 0.61 | 0.27 | 21.4× |
| 7 | Fable 5.1†Model on its own · xhigh | 0.45 | 0.18 | 0.27 | 0.33 | 0.58 | 96.6× |
| 8 | Fable 5.1†With ProsaBridge · high/xhigh/xhigh | 0.48 | 0.23 | 0.25 | 0.29 | 0.67 | 354.4× |
| 9 | Gemini 3.8 FlashWith ProsaBridge · high/xhigh/xhigh | 0.51 | 0.19 | 0.32 | 0.54 | 0.47 | 47.5× |
| 10 | GPT-5.6 LunaWith ProsaBridge · medium/xhigh/xhigh | 0.66 | 0.31 | 0.35 | 0.77 | 0.56 | 10.9× |
| 11 | Opus 5With ProsaBridge · high/xhigh/xhigh | 0.70 | 0.29 | 0.41 | 0.48 | 0.92 | 174.6× |
| 12 | Fable 5.1†Model on its own · medium | 0.81 | 0.28 | 0.53 | 0.77 | 0.85 | 45.1× |
| 13 | Gemini 3.8 FlashModel on its own · xhigh | 0.82 | 0.41 | 0.41 | 0.78 | 0.85 | 17.6× |
| 14 | Fable 5.1†With ProsaBridge · medium/medium/medium | 0.83 | 0.41 | 0.42 | 0.64 | 1.02 | 169.1× |
| 15 | GPT-5.6 LunaModel on its own · xhigh | 1.00 | 0.53 | 0.47 | 0.92 | 1.08 | 4.2× |
| 16 | Grok 4.6With ProsaBridge · high/xhigh/xhigh | 1.33 | 0.42 | 0.91 | 0.99 | 1.67 | 101× |
| 17 | GLM-5.3 FlashWith ProsaBridge · xhigh/xhigh/xhigh | 1.46 | 0.60 | 0.85 | 1.59 | 1.32 | 9.5× |
| 18 | GLM-5.3 FlashModel on its own · xhigh | 1.65 | 0.66 | 0.99 | 1.57 | 1.73 | 1× |
| 19 | Opus 5Model on its own · xhigh | 1.83 | 1.07 | 0.76 | 1.68 | 1.99 | 17.6× |
| 20 | Grok 4.6Model on its own · xhigh | 2.57 | 1.01 | 1.57 | 2.51 | 2.64 | 15.2× |
| 21 | DeepLTranslation service | 2.88 | 1.90 | 0.98 | 2.60 | 3.16 | – |
| 22 | GoogleTranslation service | 7.13 | 5.98 | 1.15 | 6.30 | 7.96 | – |
| 23 | AmazonTranslation service | 15.93 | 12.60 | 3.33 | 14.42 | 17.44 | – |
† This entry used the same AI model as one of the judges. The published score combines both judges’ assessments.
How the short-text score works
Two AI judges, Claude Fable 5.1 and GPT-6 Astra, read every translation next to its original. They were not told which entry produced it and did not see each other’s work. Each judge recorded errors with a category, a severity weight, and an explanation. The categories follow MQM (Multidimensional Quality Metrics), a framework used to assess translation quality.
A minor error counts 1 point, a major error 5, a critical error 10. The penalty points for a translation are summed, divided by its word count, and normalized to 100 words. An entry’s score for a text is the average of the two judges; its published score is the average over the twelve texts.
In the short-text study, each ProsaBridge configuration read the text, prepared a translation plan, and translated using that plan. The direct configuration received the text and required output format. The published data identifies the model and settings used for each entry. The results describe the configurations tested at the time of each study.
›How we checked the ranking
The follow-up comparisons support much of the ranking: the three translation services stay at the bottom, the order from the last AI entry down to Amazon Translate is confirmed, and 10 of the 22 neighbouring pairs keep their published order. The top entries are close together, so read the first few places as a group.
After the ranking was fixed, both judges compared each entry with the one ranked directly below it, text by text, without knowing which was which. A text only counts toward one side when both judges agree.
- 1GPT-6 Astra, With ProsaBridgevs.2GPT-5.6 Sol, With ProsaBridge1 texts where both judges preferred the higher-ranked entry, 4 texts where both judges saw no difference, 7 texts where the judges disagreedToo close to call
- 2GPT-5.6 Sol, With ProsaBridgevs.3GPT-6 Astra, Model on its own2 texts where both judges saw no difference, 3 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreedToo close to call
- 3GPT-6 Astra, Model on its ownvs.4GPT-6 Astra, With ProsaBridge1 texts where both judges preferred the higher-ranked entry, 3 texts where both judges saw no difference, 8 texts where the judges disagreedToo close to call
- 4GPT-6 Astra, With ProsaBridgevs.5GPT-5.6 Sol, Model on its own5 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 5 texts where the judges disagreedOrder confirmed
- 5GPT-5.6 Sol, Model on its ownvs.6GPT-6 Astra, Model on its own2 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 2 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreedToo close to call
- 6GPT-6 Astra, Model on its ownvs.7Fable 5.1, Model on its own4 texts where both judges preferred the higher-ranked entry, 8 texts where the judges disagreedToo close to call
- 7Fable 5.1, Model on its ownvs.8Fable 5.1, With ProsaBridge3 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 4 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreedJudges preferred the lower-ranked entry
- 8Fable 5.1, With ProsaBridgevs.9Gemini 3.8 Flash, With ProsaBridge2 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 1 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreedToo close to call
- 9Gemini 3.8 Flash, With ProsaBridgevs.10GPT-5.6 Luna, With ProsaBridge2 texts where both judges preferred the higher-ranked entry, 1 texts where both judges preferred the lower-ranked entry, 9 texts where the judges disagreedToo close to call
- 10GPT-5.6 Luna, With ProsaBridgevs.11Opus 5, With ProsaBridge1 texts where both judges preferred the higher-ranked entry, 1 texts where both judges preferred the lower-ranked entry, 10 texts where the judges disagreedToo close to call
- 11Opus 5, With ProsaBridgevs.12Fable 5.1, Model on its own6 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 1 texts where both judges preferred the lower-ranked entry, 4 texts where the judges disagreedOrder confirmed
- 12Fable 5.1, Model on its ownvs.13Gemini 3.8 Flash, Model on its own1 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 7 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreedJudges preferred the lower-ranked entry
- 13Gemini 3.8 Flash, Model on its ownvs.14Fable 5.1, With ProsaBridge5 texts where both judges preferred the higher-ranked entry, 2 texts where both judges preferred the lower-ranked entry, 5 texts where the judges disagreedOrder confirmed
- 14Fable 5.1, With ProsaBridgevs.15GPT-5.6 Luna, Model on its own4 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 7 texts where the judges disagreedToo close to call
- 15GPT-5.6 Luna, Model on its ownvs.16Grok 4.6, With ProsaBridge9 texts where both judges preferred the higher-ranked entry, 3 texts where the judges disagreedOrder confirmed
- 16Grok 4.6, With ProsaBridgevs.17GLM-5.3 Flash, With ProsaBridge4 texts where both judges preferred the higher-ranked entry, 3 texts where both judges preferred the lower-ranked entry, 5 texts where the judges disagreedOrder confirmed
- 17GLM-5.3 Flash, With ProsaBridgevs.18GLM-5.3 Flash, Model on its own6 texts where both judges preferred the higher-ranked entry, 3 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreedOrder confirmed
- 18GLM-5.3 Flash, Model on its ownvs.19Opus 5, Model on its own5 texts where both judges preferred the higher-ranked entry, 5 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreedEven
- 19Opus 5, Model on its ownvs.20Grok 4.6, Model on its own8 texts where both judges preferred the higher-ranked entry, 2 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreedOrder confirmed
- 20Grok 4.6, Model on its ownvs.21DeepL, Translation service5 texts where both judges preferred the higher-ranked entry, 4 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreedOrder confirmed
- 21DeepL, Translation servicevs.22Google, Translation service11 texts where both judges preferred the higher-ranked entry, 1 texts where the judges disagreedOrder confirmed
- 22Google, Translation servicevs.23Amazon, Translation service11 texts where both judges preferred the higher-ranked entry, 1 texts where the judges disagreedOrder confirmed
- texts where both judges preferred the higher-ranked entry
- texts where both judges saw no difference
- texts where both judges preferred the lower-ranked entry
- texts where the judges disagreed
Grok 4.6 on its own and DeepL are too close to separate. Scores that lie a few hundredths apart may swap places on another day; the download has every pair.
About the studies
Both book translations used GPT-6 Astra with the translation reasoning level set to very high. The model prompted directly received the complete manuscript and was instructed to make the translation read as though it had originally been written in English. Evaluators had no access to ProsaBridge’s glossary or notes.
- The book study assessed translation quality; it did not measure editing time or financial savings. Editorial review is still needed before publication.
- The book study compared one translation per workflow on one book under blind evaluation. The model prompted directly did not receive ProsaBridge’s glossary. The comparison assessed the complete workflows without isolating individual steps. The results apply to the settings tested at the time and do not guarantee the same improvement for other manuscripts. Models and settings can change as ProsaBridge develops. The customer’s book remains confidential; scores and section counts are available to download.
- The short-text benchmark covers English to German; the book study covers German to English. These results do not establish performance for other language pairs.
- The short-text benchmark and book study use different texts, translation directions, and ways of combining individual scores. Read each comparison within its own study; the results do not establish how document length affects the advantage.
- One test round. The numbers are not averages over repeated rounds. Differences of a few hundredths of a point may reflect chance variation.
- The judges are AI models, and two of the models that took part are also the judges. Their separate scores are shown so you can check the effect yourself. The winner is first when scored by either judge on its own.
What we keep private
- The twelve texts. They were written for this benchmark and have never been published. Publishing them could allow future models to encounter them during training, weakening later comparisons. Their titles, text types, lengths, and translation challenges are listed below.
- The instructions ProsaBridge gives the models, and the instructions the judges follow.
- What the test rounds cost. The results table shows only multiples of the cheapest entry.
- The judges’ notes and the translations themselves, because they would reveal the texts.
The twelve texts
Seven short stories and five nonfiction pieces, each written to be hard to translate in a particular way.
| Text | Kind | Words | What makes it hard | |
|---|---|---|---|---|
| 1 | The Lease | Fiction, dialogue | 431 | A shift from formal to informal address, dialogue, dry humor |
| 2 | Gravel | Fiction, inner voice | 485 | Thoughts in free indirect speech, colloquial asides, unmarked questions |
| 3 | The Orchard Ledger | Fiction, lyrical | 463 | One long metaphor that must hold up from start to finish |
| 4 | Southpaw | Fiction, scene | 449 | Short, hard sentences next to one long one |
| 5 | The Committee | Fiction, satire | 446 | Long nested sentences, asides, irony that depends on word order |
| 6 | Moving the Needle | Fiction, workplace | 462 | Many idioms and one pun the whole text depends on |
| 7 | The Visitor | Fiction, inner voice | 458 | A deliberately unclear “she” who must remain unclear |
| 8 | Remarks at the Dedication of the Cedar Run Flood Memorial | Nonfiction, speech | 441 | American civic tone, repetition for effect, names of institutions |
| 9 | Allision of the Ferry Marguerite Bay at Harbor Point Terminal | Nonfiction, accident report | 478 | Nautical terms, passive voice, exact times and measurements |
| 10 | The Sleep of Swifts | Nonfiction, popular science | 493 | Lyrical science writing that moves between data and imagery |
| 11 | Expanding District Heating in Mid-Sized Cities | Nonfiction, policy brief | 436 | “Should”, “must”, and “may” that have to keep their exact strength; false friends; consistent terms |
| 12 | Midterm Evaluation of the Aldercreek Adult Literacy Program | Nonfiction, evaluation report | 437 | Paired terms that must stay distinct, careful attribution, measured recommendations |
Models we removed before publishing
These models took part during development and are missing from the published results. Each removal is recorded with its evidence.
- Claude Sonnet 5
- Removed after seven texts because of weak translation quality: 3.18 and 3.22 error points per 100 words, behind every other AI model at the time.
- Kimi K2 Thinking
- Removed after it hit the provider’s output limit on three of the seven texts without delivering an answer.
- Mistral Large 3
- Removed after it added document formatting of its own and returned translated text where the process requires fixed reference codes.
- Gemini 3.1 Pro and GPT-5.6 Terra
- Withdrawn before the published test round: for what they cost, their scores were too weak. Cheaper models scored better, and better models cost about the same. Their earlier results are archived.
Download the data
The benchmark round: each entry on each text as scored by each judge, error counts by severity and category, relative compute costs, and comparisons between neighbouring entries in the ranking.
Download the 1.0 datasetThe book study: overall scores for both translations, scores by error category and judge, and the book’s word and section counts. The book text is not included.
Download the book study data- Test round
- translationbench-run-fa35fa02-3190-4f82-b379-8d05b7dde5d7
- Version
- 1.0 (7.2)
- Completed
- September 12, 2026