SWE-bench
Emerging9papers using it
66,082HF downloads
144HF likes
2025first seen
Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 2,294 Issue-Pull Request pairs from 12 popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released a
Papers using SWE-bench (9)
- Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement LearningTerminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward ModelsOrchard: An Open-Source Agentic Modeling FrameworkA Rubric-Supervised Critic from Sparse Real-World OutcomesReinforcement Learning for Chain of Thought Compression with One-Domain-to-All GeneralizationSkyRL-Agent: Efficient RL Training for Multi-turn LLM AgentSWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource ConstraintsSatori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering