← all papers · overview

Evalsense: A Framework For Domain-specific LLM (meta-)evaluation

Abstract

Robust and comprehensive evaluation of large language models (LLMs) is essential for identifying effective LLM system configurations and mitigating risks associated with deploying LLMs in sensitive domains. However, traditional statistical metrics are poorly suited to open-ended generation tasks, leading to growing reliance on LLM-based evaluation methods. These methods, while often more flexible,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).