← all papers · overview

Is My Meeting Summary Good? Estimating Quality With A Multi-llm Evaluator

Abstract

The quality of meeting summaries generated by natural language generation (NLG) systems is hard to measure automatically. Established metrics such as ROUGE and BERTScore have a relatively low correlation with human judgments and fail to capture nuanced errors. Recent studies suggest using large language models (LLMs), which have the benefit of better context understanding and adaption of error def

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).