‹ Culture 3.14 · What we ask ourselves
Question · Language
How do you evaluate a shift in meaning that reads perfectly natural?
Automatic metrics reward fluency, and the most expensive error in a translation is precisely a fluent one.
Why it matters
The errors that catch themselves are the clumsy ones. The one that matters is the one that produces an impeccable sentence with a different obligation, a missing condition or an inverted nuance.
No similarity-based metric flags it, because the sentence looks a great deal like the original.
What we know so far
Human review aimed at specific categories — conditions, negations, quantifiers, modality, scope — finds considerably more than free review does. Comparing the pair of texts sentence by sentence, with the differences marked, speeds that work up a lot.
We measure terminological compliance separately, which is automatable, from the shift in meaning, which for now is not.
What remains open
We are evaluating whether a model can detect these categories reliably in specific language pairs. The results vary too much between domains to generalise.
Nor do we know how much human review is proportionate when the volume is high and the risk per piece is low.