SWE-Bench Pro
Emerging9papers using it
2025first seen
Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks. Paper: https://static.scale.com/uploads/654197dc94d34f66c0f5184e/SWEAP_Eval_Scale%20(9).pdf See the related evaluation Github: https://github.com/scaleapi/SWE-bench_Pro-os Datas
Papers using SWE-Bench Pro (8)
- Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent SkillsEvolving Agents in the Dark: Retrospective Harness Optimization via Self-PreferenceExploration Structure in LLM Agents for Multi-File Change LocalizationOpen-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering AgentsSkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to EvolutionEvaluating Plan Compliance In Autonomous Programming AgentsSWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue ResolutionToward Training Superintelligent Software Agents through Self-Play SWE-RL