← all datasets

DevEval

Emerging
2papers using it
2026first seen

The 'DevEval' dataset/benchmark contains a collection of real-world coding tasks used to evaluate the performance of multi-agent Large Language Model systems in software engineering.

Papers using DevEval (1)

DevEval dataset β€” papers, benchmarks & downloads Β· AI Agents