Randomized trial compares translation accuracy in literary texts between human and DeepL, suggesting AI can improve quality.
The advent of artificial intelligence (AI) and the proliferation of automatic translation tools have generated questions about the quality of AI-generated literary translation. This study set out to apply Larson’s (1984) model to compare the human and DeepL translations of Nicole Brossard’s Desert mauve, to judge the accuracy, clarity, and naturalness of each version and to determine which is precise. After analysing twenty purposively selected excerpts, the study revealed that on one hand, the human translator was constrained by sociological and the communicative realities of her recipients, which made just 50% of her excerpts accurate, for she sometimes over translated or under translated. Since DeepL, on the other hand, did not function under such constraints, it produced 75% accurate renderings. It thus concluded that the human translator was not accurate and failed to precisely convey the source text author’s intention to target readers because she lacked a framework for literary analysis other than metatexts, which made her assume the author’s intention. The study resolved that DeepL is accurate for rendering literary texts, although the translations it produces must be post-edited for them to be completely exact, clear and natural.
No takes yet. Share an insight, caveat, or question.
Tanyitiku Enaka Agbor Bayee (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: