QA benchmarks
Emerging4papers using it
2026first seen
The 'QA benchmarks' are datasets used to evaluate the performance of agents in knowledge-intensive question answering tasks, specifically measuring their accuracy and efficiency in retrieving and reasoning with information.
Papers using QA benchmarks (4)
- Grad Detect: Gradient-Based Hallucination Detection in LLMsWhen Can Conformal Risk Control Certify LLM Outputs? Bounds, Impossibility, and Adaptation for Structured GenerationModelLens: Finding the Best for Your Task from Myriads of ModelsCalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learning in Knowledge-Intensive Tasks