DevEval
Emerging2papers using it
2026first seen
The 'DevEval' dataset/benchmark contains a collection of real-world coding tasks used to evaluate the performance of multi-agent Large Language Model systems in software engineering.
The 'DevEval' dataset/benchmark contains a collection of real-world coding tasks used to evaluate the performance of multi-agent Large Language Model systems in software engineering.