Back to home

TranslationBench 1.0 · Evidence for editorial teams

Preserve the meaning. Across the whole manuscript.

ProsaBridge combines greater fidelity to the original with fluent, readable prose. In a blind comparison of a 67,127-word scholarly book, ProsaBridge reduced weighted error points by 59% compared with the same AI model prompted directly.

Book comparison
Same AI model, with ProsaBridge and prompted directly
Manuscript length
67,127 words · 61 sections
Evaluation
Independent blind scoring, identical criteria across both workflows
Published evidence
Book: DE → EN · short texts: EN → DE

The ProsaBridge advantage

Fewer error points across the whole book
−59%
Across the complete book, the ProsaBridge translation received fewer weighted error points than the same AI model used directly. Independent evaluators assessed both versions without knowing how they were produced.
Fewer points for meaning errors
−64%
ProsaBridge cut weighted error points for changes in meaning, incorrect terms, and omissions by nearly two thirds, while keeping the translation fluent and readable.

Book study · September 16, 2026

A whole book, translated more faithfully with ProsaBridge

59%

fewer weighted error points with ProsaBridge than when the same model was prompted directly. The accuracy error score fell by 64%, combining greater fidelity to the original with fluent, readable prose.

ProsaBridge brought greater fidelity to a 67,127-word scholarly book translated from German into English. We compared it with the same AI model prompted directly with the complete manuscript. Two independent evaluators checked every section against the original, without knowing which workflow produced either translation.

The book comparison shows the value ProsaBridge brings to a complete manuscript. You can explore the full results and evaluation method below.

What you get with ProsaBridge

Set the terminology and style before translation begins, carry the manuscript’s context into every section, and review the finished translation in Word.

Your terminology. Your style.
Review terminology and style instructions, and resolve questions about the source text before translation begins.
Context for every section
Every section is translated with the manuscript summary, glossary, and style instructions, for consistent terminology and tone throughout.
A document ready for review
Work in Word and receive a translated .docx that retains the document’s structure and formatting.

Keep your team in control, from the first terminology decisions to the final review in Word.

See plans and pricing
Method and complete results

The book study and short-text benchmark are separate experiments with different texts, language directions and model settings. This appendix contains the scoring method, all published results and downloadable records.

Book study: scores from both evaluators

Book study · September 16, 2026 · One book · 61 sections · 67,127 words · German → English · blind evaluation · text confidential

Book scores are weighted error points per 100 source words, pooled across all sections and both judges. Minor errors count 1 point, major errors 5, and critical errors 10. Lower scores mean fewer weighted error points. Accuracy covers changed meaning, incorrect terms, and omissions; readability covers how the translation reads. These scores do not count incorrect words or predict how often an editor will need to make a correction.

EntryError scoreScoreAccuracyReadabilityFable 5.1 mediumGPT-6 Astra medium
GPT-6 Astra inside ProsaBridge0.080.060.020.120.04
GPT-6 Astra prompted directly0.200.180.020.280.12

AccuracyReadability

How to read the short-text results

An entry is one AI model used in one of two ways, or one translation service. There are 23 entries.

With ProsaBridge
The AI model worked inside ProsaBridge. It read the whole text first, prepared a translation plan, and translated using that plan.
Model on its own
The same AI model received the text once and translated it in a single step.
Translation service
DeepL, Google Translate, or Amazon Translate, each translating the whole document through its own service.
Error score
Each judge assigns the errors they find to Multidimensional Quality Metrics (MQM) categories. In this benchmark, a minor error counts 1 point, a major error 5, and a critical error 10. The error score is the sum of penalty points per 100 words, averaged over both judges and all twelve texts. Lower is better.
The score reflects both the number and severity of errors: one major error in meaning carries the same weight as five minor errors. It is not a count of errors or a percentage of incorrect words, and it does not predict how often an editor will need to make a correction.
Accuracy
The translation departs from the original meaning: an incorrect fact, an inaccurate term, or an omitted sentence.
Readability
The translation conveys the meaning but reads poorly: awkward phrasing, grammatical mistakes, or inappropriate tone.

Complete short-text results

All 23 published entries from twelve English-to-German texts, with model settings, separate judge scores and relative compute costs. The results describe the configurations tested in this study. Compute costs are model usage charges expressed as multiples of the lowest measured cost; they exclude document handling and editorial work and are not ProsaBridge prices.

Click a column to sort · sorted by Score

RankEntry
1GPT-6 AstraWith ProsaBridge · high/xhigh/xhigh0.140.060.080.270.02501.9×
2GPT-5.6 SolWith ProsaBridge · high/xhigh/xhigh0.250.070.180.340.16158.9×
3GPT-6 AstraModel on its own · xhigh0.250.140.110.310.20184.2×
4GPT-6 AstraWith ProsaBridge · medium/medium/medium0.250.090.160.420.09109.5×
5GPT-5.6 SolModel on its own · xhigh0.400.200.210.470.3484.9×
6GPT-6 AstraModel on its own · medium0.440.120.320.610.2721.4×
7Fable 5.1Model on its own · xhigh0.450.180.270.330.5896.6×
8Fable 5.1With ProsaBridge · high/xhigh/xhigh0.480.230.250.290.67354.4×
9Gemini 3.8 FlashWith ProsaBridge · high/xhigh/xhigh0.510.190.320.540.4747.5×
10GPT-5.6 LunaWith ProsaBridge · medium/xhigh/xhigh0.660.310.350.770.5610.9×
11Opus 5With ProsaBridge · high/xhigh/xhigh0.700.290.410.480.92174.6×
12Fable 5.1Model on its own · medium0.810.280.530.770.8545.1×
13Gemini 3.8 FlashModel on its own · xhigh0.820.410.410.780.8517.6×
14Fable 5.1With ProsaBridge · medium/medium/medium0.830.410.420.641.02169.1×
15GPT-5.6 LunaModel on its own · xhigh1.000.530.470.921.084.2×
16Grok 4.6With ProsaBridge · high/xhigh/xhigh1.330.420.910.991.67101×
17GLM-5.3 FlashWith ProsaBridge · xhigh/xhigh/xhigh1.460.600.851.591.329.5×
18GLM-5.3 FlashModel on its own · xhigh1.650.660.991.571.73
19Opus 5Model on its own · xhigh1.831.070.761.681.9917.6×
20Grok 4.6Model on its own · xhigh2.571.011.572.512.6415.2×
21DeepLTranslation service2.881.900.982.603.16
22GoogleTranslation service7.135.981.156.307.96
23AmazonTranslation service15.9312.603.3314.4217.44

This entry used the same AI model as one of the judges. The published score combines both judges’ assessments.

How the short-text score works

Two AI judges, Claude Fable 5.1 and GPT-6 Astra, read every translation next to its original. They were not told which entry produced it and did not see each other’s work. Each judge recorded errors with a category, a severity weight, and an explanation. The categories follow MQM (Multidimensional Quality Metrics), a framework used to assess translation quality.

A minor error counts 1 point, a major error 5, a critical error 10. The penalty points for a translation are summed, divided by its word count, and normalized to 100 words. An entry’s score for a text is the average of the two judges; its published score is the average over the twelve texts.

In the short-text study, each ProsaBridge configuration read the text, prepared a translation plan, and translated using that plan. The direct configuration received the text and required output format. The published data identifies the model and settings used for each entry. The results describe the configurations tested at the time of each study.

How we checked the ranking

The follow-up comparisons support much of the ranking: the three translation services stay at the bottom, the order from the last AI entry down to Amazon Translate is confirmed, and 10 of the 22 neighbouring pairs keep their published order. The top entries are close together, so read the first few places as a group.

After the ranking was fixed, both judges compared each entry with the one ranked directly below it, text by text, without knowing which was which. A text only counts toward one side when both judges agree.

  1. 1GPT-6 Astra, With ProsaBridgevs.2GPT-5.6 Sol, With ProsaBridge
    1 texts where both judges preferred the higher-ranked entry, 4 texts where both judges saw no difference, 7 texts where the judges disagreed
    Too close to call
  2. 2GPT-5.6 Sol, With ProsaBridgevs.3GPT-6 Astra, Model on its own
    2 texts where both judges saw no difference, 3 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreed
    Too close to call
  3. 3GPT-6 Astra, Model on its ownvs.4GPT-6 Astra, With ProsaBridge
    1 texts where both judges preferred the higher-ranked entry, 3 texts where both judges saw no difference, 8 texts where the judges disagreed
    Too close to call
  4. 4GPT-6 Astra, With ProsaBridgevs.5GPT-5.6 Sol, Model on its own
    5 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 5 texts where the judges disagreed
    Order confirmed
  5. 5GPT-5.6 Sol, Model on its ownvs.6GPT-6 Astra, Model on its own
    2 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 2 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreed
    Too close to call
  6. 6GPT-6 Astra, Model on its ownvs.7Fable 5.1, Model on its own
    4 texts where both judges preferred the higher-ranked entry, 8 texts where the judges disagreed
    Too close to call
  7. 7Fable 5.1, Model on its ownvs.8Fable 5.1, With ProsaBridge
    3 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 4 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreed
    Judges preferred the lower-ranked entry
  8. 8Fable 5.1, With ProsaBridgevs.9Gemini 3.8 Flash, With ProsaBridge
    2 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 1 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreed
    Too close to call
  9. 9Gemini 3.8 Flash, With ProsaBridgevs.10GPT-5.6 Luna, With ProsaBridge
    2 texts where both judges preferred the higher-ranked entry, 1 texts where both judges preferred the lower-ranked entry, 9 texts where the judges disagreed
    Too close to call
  10. 10GPT-5.6 Luna, With ProsaBridgevs.11Opus 5, With ProsaBridge
    1 texts where both judges preferred the higher-ranked entry, 1 texts where both judges preferred the lower-ranked entry, 10 texts where the judges disagreed
    Too close to call
  11. 11Opus 5, With ProsaBridgevs.12Fable 5.1, Model on its own
    6 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 1 texts where both judges preferred the lower-ranked entry, 4 texts where the judges disagreed
    Order confirmed
  12. 12Fable 5.1, Model on its ownvs.13Gemini 3.8 Flash, Model on its own
    1 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 7 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreed
    Judges preferred the lower-ranked entry
  13. 13Gemini 3.8 Flash, Model on its ownvs.14Fable 5.1, With ProsaBridge
    5 texts where both judges preferred the higher-ranked entry, 2 texts where both judges preferred the lower-ranked entry, 5 texts where the judges disagreed
    Order confirmed
  14. 14Fable 5.1, With ProsaBridgevs.15GPT-5.6 Luna, Model on its own
    4 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 7 texts where the judges disagreed
    Too close to call
  15. 15GPT-5.6 Luna, Model on its ownvs.16Grok 4.6, With ProsaBridge
    9 texts where both judges preferred the higher-ranked entry, 3 texts where the judges disagreed
    Order confirmed
  16. 16Grok 4.6, With ProsaBridgevs.17GLM-5.3 Flash, With ProsaBridge
    4 texts where both judges preferred the higher-ranked entry, 3 texts where both judges preferred the lower-ranked entry, 5 texts where the judges disagreed
    Order confirmed
  17. 17GLM-5.3 Flash, With ProsaBridgevs.18GLM-5.3 Flash, Model on its own
    6 texts where both judges preferred the higher-ranked entry, 3 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreed
    Order confirmed
  18. 18GLM-5.3 Flash, Model on its ownvs.19Opus 5, Model on its own
    5 texts where both judges preferred the higher-ranked entry, 5 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreed
    Even
  19. 19Opus 5, Model on its ownvs.20Grok 4.6, Model on its own
    8 texts where both judges preferred the higher-ranked entry, 2 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreed
    Order confirmed
  20. 20Grok 4.6, Model on its ownvs.21DeepL, Translation service
    5 texts where both judges preferred the higher-ranked entry, 4 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreed
    Order confirmed
  21. 21DeepL, Translation servicevs.22Google, Translation service
    11 texts where both judges preferred the higher-ranked entry, 1 texts where the judges disagreed
    Order confirmed
  22. 22Google, Translation servicevs.23Amazon, Translation service
    11 texts where both judges preferred the higher-ranked entry, 1 texts where the judges disagreed
    Order confirmed
  • texts where both judges preferred the higher-ranked entry
  • texts where both judges saw no difference
  • texts where both judges preferred the lower-ranked entry
  • texts where the judges disagreed

Grok 4.6 on its own and DeepL are too close to separate. Scores that lie a few hundredths apart may swap places on another day; the download has every pair.

About the studies

Both book translations used GPT-6 Astra with the translation reasoning level set to very high. The model prompted directly received the complete manuscript and was instructed to make the translation read as though it had originally been written in English. Evaluators had no access to ProsaBridge’s glossary or notes.

  • The book study assessed translation quality; it did not measure editing time or financial savings. Editorial review is still needed before publication.
  • The book study compared one translation per workflow on one book under blind evaluation. The model prompted directly did not receive ProsaBridge’s glossary. The comparison assessed the complete workflows without isolating individual steps. The results apply to the settings tested at the time and do not guarantee the same improvement for other manuscripts. Models and settings can change as ProsaBridge develops. The customer’s book remains confidential; scores and section counts are available to download.
  • The short-text benchmark covers English to German; the book study covers German to English. These results do not establish performance for other language pairs.
  • The short-text benchmark and book study use different texts, translation directions, and ways of combining individual scores. Read each comparison within its own study; the results do not establish how document length affects the advantage.
  • One test round. The numbers are not averages over repeated rounds. Differences of a few hundredths of a point may reflect chance variation.
  • The judges are AI models, and two of the models that took part are also the judges. Their separate scores are shown so you can check the effect yourself. The winner is first when scored by either judge on its own.

What we keep private

  • The twelve texts. They were written for this benchmark and have never been published. Publishing them could allow future models to encounter them during training, weakening later comparisons. Their titles, text types, lengths, and translation challenges are listed below.
  • The instructions ProsaBridge gives the models, and the instructions the judges follow.
  • What the test rounds cost. The results table shows only multiples of the cheapest entry.
  • The judges’ notes and the translations themselves, because they would reveal the texts.

The twelve texts

Seven short stories and five nonfiction pieces, each written to be hard to translate in a particular way.

TextKindWordsWhat makes it hard
1The LeaseFiction, dialogue431A shift from formal to informal address, dialogue, dry humor
2GravelFiction, inner voice485Thoughts in free indirect speech, colloquial asides, unmarked questions
3The Orchard LedgerFiction, lyrical463One long metaphor that must hold up from start to finish
4SouthpawFiction, scene449Short, hard sentences next to one long one
5The CommitteeFiction, satire446Long nested sentences, asides, irony that depends on word order
6Moving the NeedleFiction, workplace462Many idioms and one pun the whole text depends on
7The VisitorFiction, inner voice458A deliberately unclear “she” who must remain unclear
8Remarks at the Dedication of the Cedar Run Flood MemorialNonfiction, speech441American civic tone, repetition for effect, names of institutions
9Allision of the Ferry Marguerite Bay at Harbor Point TerminalNonfiction, accident report478Nautical terms, passive voice, exact times and measurements
10The Sleep of SwiftsNonfiction, popular science493Lyrical science writing that moves between data and imagery
11Expanding District Heating in Mid-Sized CitiesNonfiction, policy brief436“Should”, “must”, and “may” that have to keep their exact strength; false friends; consistent terms
12Midterm Evaluation of the Aldercreek Adult Literacy ProgramNonfiction, evaluation report437Paired terms that must stay distinct, careful attribution, measured recommendations

Models we removed before publishing

These models took part during development and are missing from the published results. Each removal is recorded with its evidence.

Claude Sonnet 5
Removed after seven texts because of weak translation quality: 3.18 and 3.22 error points per 100 words, behind every other AI model at the time.
Kimi K2 Thinking
Removed after it hit the provider’s output limit on three of the seven texts without delivering an answer.
Mistral Large 3
Removed after it added document formatting of its own and returned translated text where the process requires fixed reference codes.
Gemini 3.1 Pro and GPT-5.6 Terra
Withdrawn before the published test round: for what they cost, their scores were too weak. Cheaper models scored better, and better models cost about the same. Their earlier results are archived.

Download the data

The benchmark round: each entry on each text as scored by each judge, error counts by severity and category, relative compute costs, and comparisons between neighbouring entries in the ranking.

Download the 1.0 dataset

The book study: overall scores for both translations, scores by error category and judge, and the book’s word and section counts. The book text is not included.

Download the book study data
Test round
translationbench-run-fa35fa02-3190-4f82-b379-8d05b7dde5d7
Version
1.0 (7.2)
Completed
September 12, 2026