SWE-bench Lite
Emerging16papers using it
2025first seen
Dataset Summary SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The datase
Papers using SWE-bench Lite (16)
- Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation ModelsSocratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent SkillsContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program RepairNo Time Like the Present: Agentic Test-Time Training for LLM AgentsGoal-Autopilot: A Verifiable Anti-Fabrication Firewall for Unattended Long-Horizon AgentsPhoenix: Safe GitHub Issue Resolution via Multi-Agent LLMsSIADAFIX: Issue Description Response For Adaptive Program RepairSHERLOC: Structured Diagnostic Localization for Code Repair AgentsCausal state binding predicts action control in language agentsBLAgent: Agentic RAG for File-Level Bug LocalizationAgentForge: Execution-Grounded Multi-Agent LLM Framework for Autonomous Software EngineeringA Self-Evolving Framework for Efficient Terminal Agents via Observational Context CompressionMarketbench: Evaluating AI Agents As Market ParticipantsSWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue ResolutionVulnResolver: A Hybrid Agent Framework for LLM-Based Automated Vulnerability Issue ResolutionOrcaLoca: An LLM Agent Framework for Software Issue Localization