← all papers · overview

GPT-4 Vs. Human Translators: A Comprehensive Evaluation Of Translation Quality Across Languages, Domains, And Expertise Levels

Abstract

This study comprehensively evaluates the translation quality of Large Language Models (LLMs), specifically GPT-4, against human translators of varying expertise levels across multiple language pairs and domains. Through carefully designed annotation rounds, we find that GPT-4 performs comparably to junior translators in terms of total errors made but lags behind medium and senior translators. We a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).