LiveCodeBench
Emerging17papers using it
2025first seen
The 'LiveCodeBench' dataset/benchmark contains a collection of coding tasks and is used to evaluate the performance of large language models in generating correct and robust code solutions.
Papers using LiveCodeBench (17)
- Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation ModelsSakana Fugu Technical ReportDon't Let Gains FADE: Breaking Down Policy Gradient Weights in RLDecompRL: Solving Harder Problems by Learning Modular Code GenerationLearning to Orchestrate Agents in Natural Language with the ConductorATLAS: Agentic Test-time Learning-to-Allocate ScalingEvoSyn: Generalizable Evolutionary Data Synthesis for Verifiable LearningCodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming SolutionsReVeal: Self-Evolving Code Agents via Reliable Self-VerificationACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference OptimizationCollabCoder: Plan-Code Co-Evolution via Collaborative Decision-Making for Efficient Code GenerationThink Anywhere in Code GenerationTeam of Thoughts: Efficient Test-time Scaling of Agentic Systems through Orchestrated Tool CallingTRINITY: An Evolved LLM CoordinatorA Self-improving Coding AgentSEW: Self-Evolving Agentic Workflows for Automated Code GenerationPangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition