← all datasets

multi-hop QA benchmarks

Emerging
6papers using it
2025first seen

Multi-hop QA benchmarks are datasets used to evaluate the ability of models to answer questions that require reasoning over multiple pieces of evidence or information.

Papers using multi-hop QA benchmarks (6)

multi-hop QA benchmarks dataset β€” papers, benchmarks & downloads Β· AI Agents