← all papers · overview

Tapilot-crossing: Benchmarking And Evolving Llms Towards Interactive Data Analysis Agents

Abstract

Interactive Data Analysis, the collaboration between humans and LLM agents, enables real-time data exploration for informed decision-making. The challenges and costs of collecting realistic interactive logs for data analysis hinder the quantitative evaluation of Large Language Model (LLM) agents in this task. To mitigate this issue, we introduce Tapilot-Crossing, a new benchmark to evaluate LLM ag

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).