SWE-bench
Canonical49papers using it
2024first seen
Real GitHub issues paired with their fix PRs across Python repositories; models must produce patches that pass the repo's tests.
Papers using SWE-bench (49)
- Swe-agent: Agent-computer Interfaces Enable Automated Software EngineeringQwen3-Coder-Next Technical ReportPACE: A Proxy for Agentic Capability EvaluationOrchard: An Open-Source Agentic Modeling FrameworkAgentcgroup: Understanding And Controlling OS Resources Of AI AgentsAgentic Software: How AI Agents Are Restructuring the Software ParadigmSelf-Evolving Agents with Anytime-Valid CertificatesCoACT: Action-Preserving Observation Compression for Coding AgentsWhen Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design AblationThe Hidden Footprint: Making Storage a First-Class Metric for LLM Agent EvaluationUMoE:Unlocking Every Expert in Domain-Specific TrainingThe Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes NecessaryADK Arena: Evaluating Agent Development Kits via LLM-as-a-DeveloperYour Agent Has a Genome: Sequence-Level Behavioral Analysis and Runtime Governance of LLM-Powered Autonomous AgentsProbe-and-Refine Tuning of Repository Guidance for Coding AgentsTwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM RoutingSpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM AgentsAdaptOrch: Task-Adaptive Multi-Agent Orchestration in the Era of LLM Performance ConvergenceAdarubric: Task-adaptive Rubrics For LLM Agent EvaluationHybrid-gym: Training Coding Agents To Generalize Across TasksToken Reduction Is Not Cost ReductionProcess Reward Informed Tree Rollout for Effective Multi-Turn RLInfantagent-next: A Multimodal Generalist Agent For Automated Computer InteractionCalibrating Conservatism for Scalable OversightEvaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?When Agents Go Astray: Course-correcting SWE Agents With PrmsPatchpilot: A Cost-efficient Software Engineering Agent With Early Attempts On Formal VerificationTaming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel FaultsOmniCode: A Benchmark for Evaluating Software Engineering AgentsBeyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic JudgmentHow Do AI Agents Spend Your Money? Analyzing And Predicting Token Consumption In Agentic Coding TasksSwe-prot\'eg\'e: Learning To Selectively Collaborate With An Expert Unlocks Small Language Models As Software Engineering AgentsAn Executable Benchmarking Suite for Tool-Using AgentsSame Signal, Different Semantics: A Cross-Framework Behavioral Analysis of Software Engineering AgentsFrom Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to PythonSWE-Edit: Rethinking Code Editing for Efficient SWE-AgentAOrchestra: Automating Sub-Agent Creation for Agentic OrchestrationEvoMAS: Evolutionary Generation of Multi-Agent SystemsAgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code GenerationAgent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path ForwardSGAgent: Suggestion-Guided LLM-Based Multi-Agent Framework for Repository-Level Software RepairSkyRL-Agent: Efficient RL Training for Multi-turn LLM AgentKimi-dev: Agentless Training As Skill Prior For Swe-agentsUnderstanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical StudyBreakpoint: Scalable Evaluation Of System-level Reasoning In LLM Code AgentsPutting It All Into Context: Simplifying Agents With LclmsSwe-rebench: An Automated Pipeline For Task Collection And Decontaminated Evaluation Of Software Engineering AgentsHyperagent: Generalist Software Engineering Agents To Solve Coding Tasks At ScaleREDO: Execution-free Runtime Error Detection For Coding Agents