AI Translations Are Fine But Humans Win
Based on research by Yves Ferstler, Adam Podoxin, Ty Brassington, Roman Grundkiewicz, Maite Taboada
You are reading a translated novel. The prose flows, the plot makes sense, and the characters feel real. So why does a nagging sense of something missing linger in the back of your mind? New research suggests that while AI translation has crossed the threshold of adequacy, it still fails to deliver the soul of literature.
Researchers conducted a deep dive into how readers experience machine-generated translations compared to human ones. They asked fifteen avid readers to compare excerpts from fifteen recent novels translated from French, Polish, and Japanese into English. The participants engaged in two types of reading: immersive reading of full excerpts and close reading of aligned chunks. The goal was to measure immersiveness and literary effect, aspects that standard automated metrics often miss.
The results reveal a subtle but significant divide. Readers described the AI translations as "fine," yet they consistently preferred human translations. At the excerpt level, the preference was slight, but during close reading, the human advantage became clear. Readers cited human translations as easier to read, clearer, and more immersive. Interestingly, the quality of AI translations varied more within a single book than human translations did. Perhaps most surprisingly, readers could not reliably tell the difference between the two. When they guessed, they tended to prefer the version they believed was human, highlighting a powerful bias toward the human touch.
This disconnect exposes a major flaw in how we currently evaluate translation technology. Automatic metrics, including those that use large language models as judges, failed to reflect reader preferences and instead favored the machine translations. To address this, the researchers released LAIT, a reader-centered evaluation dataset containing thousands of judgments and annotations. The takeaway is clear: fluency is not enough. For literature, the human element remains indispensable, and our tools need to evolve to measure what truly matters to the reader.