← all datasets

Who&When

Emerging
7papers using it
2025first seen

The 'Who&When' dataset/benchmark contains data used to evaluate the step-level accuracy of causal attribution in LLM agents, specifically focusing on identifying which step in an agent's decision-making process caused a failure.

Papers using Who&When (6)

Who&When dataset β€” papers, benchmarks & downloads Β· AI Agents