← all papers · overview

MQM-APE: Toward High-quality Error Annotation Predictors With Automatic Post-editing In LLM Translation Evaluators

Abstract

Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment, providing both scores and fine-grained feedback. Although approaches such as GEMBA-MQM have shown state-of-the-art performance on reference-free evaluation, the predicted errors do not align well with those annotated by human, limiting their interpretability as feedback signals.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).