SWE-bench Verified
Canonical45papers using it
2025first seen
A human-validated subset of SWE-bench with confirmed-solvable, well-specified issues.
Papers using SWE-bench Verified (45)
- Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation ModelsHarnessX: A Composable, Adaptive, and Evolvable Agent Harness FoundrySocratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent SkillsHarnessBridge: Learnable Bidirectional Controller for LLM Agent HarnessSingle-Rollout Asynchronous Optimization for Agentic Reinforcement LearningLLM-as-a-Verifier: A General-Purpose Verification FrameworkDecentralized Multi-Agent Systems with Shared ContextSWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning ContextFrom Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness FlawsDon't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent QualityCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentsThe Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous AgentsOpen-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering AgentsLong Live the Librarian! A Persistent Search Sub-Agent for Energy-Efficient Multi-Agent Software Engineering SystemsAgentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent HarnessesHybrid-gym: Training Coding Agents To Generalize Across TasksWhat Context Does a Coding Agent Actually Need to Act?Lean4Agent: Formal Modeling and Verification for Agent Workflow and TrajectoryFrontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming LanguagesSHERLOC: Structured Diagnostic Localization for Code Repair AgentsProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding AgentsAutomated Benchmark Auditing for AI Agents and Large Language ModelsCoMem: Context Management with A Decoupled Long-Context ModelSwe-bench-cl: Continual Learning For Coding AgentsSwe-prot\'eg\'e: Learning To Selectively Collaborate With An Expert Unlocks Small Language Models As Software Engineering AgentsCRANE: Constrained Reasoning Injection for Code Agents via Nullspace EditingRatchet: A Minimal Hygiene Recipe for Self-Evolving LLM AgentsGuardrails Beat Guidance: A Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding AgentsSWE-Edit: Rethinking Code Editing for Efficient SWE-AgentEvaluating Plan Compliance In Autonomous Programming AgentsCoherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right CodeAsk or Assume? Uncertainty-Aware Clarification-Seeking in Coding AgentsSWE-Universe: Scale Real-World Verifiable Environments to MillionsSWE-Master: Unleashing the Potential of Software Engineering Agents via Post-TrainingEvoMAS: Evolutionary Generation of Multi-Agent SystemsGroup-evolving Agents: Open-ended Self-improvement Via Experience SharingLearning Adaptive Parallel Execution for Efficient Code LocalizationToward Training Superintelligent Software Agents through Self-Play SWE-RLSWE-EVO: Benchmarking Coding Agents In Long-horizon Software Evolution ScenariosSelf-Abstraction from Grounded Experience for Plan-Guided Policy RefinementR2e-gym: Procedural Environments And Hybrid Verifiers For Scaling Open-weights SWE AgentsPutting It All Into Context: Simplifying Agents With LclmsA Self-improving Coding AgentEstablishing Best Practices For Building Rigorous Agentic BenchmarksGuided Search Strategies In Non-serializable Environments With Applications To Software Engineering Agents