← all papers · overview

Compound-qa: A Benchmark For Evaluating Llms On Compound Questions

Abstract

Large language models (LLMs) demonstrate remarkable performance across various tasks, prompting researchers to develop diverse evaluation benchmarks. However, most benchmarks typically measure the ability of LLMs to respond to individual questions, neglecting the complex interactions in real-world applications. We introduce Compound Question Synthesis (CQ-Syn) to build Compound-QA, a benchmark tar

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).