LiveCodeBench
Canonical79papers using it
2024first seen
A contamination-resistant coding benchmark that continuously collects new competitive-programming problems over time.
Papers using LiveCodeBench (79)
- KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for
CodingKodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for CodingFLARE: Fine-Grained Diagnostic Feedback for LLM Code RefinementInferring Code Correctness from SpecificationFocused-DPO: Enhancing Code Generation Through Focused Preference Optimization on Error-Prone PointsCodeElo: Benchmarking Competition-level Code Generation of LLMs with
Human-comparable Elo RatingsCodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming SolutionsOpenThoughts: Data Recipes for Reasoning ModelsAutomated Repair of Ambiguous Problem Descriptions for LLM-Based Code GenerationOpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMsEnhancing LLM Code Generation with Ensembles: A Similarity-Based Selection ApproachFunction-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation ModelsApriel-Reasoner: RL Post-Training for General-Purpose and Efficient ReasoningBenchEvolver: Frontier Task Synthesis via Solution-Centric EvolutionMulti-LCB: Extending LiveCodeBench to Multiple Programming LanguagesARIADNE: Agentic Reward-Informed Adaptive Decision Exploration via Blackboard-Driven MCTS for Competitive Program GenerationUsing Semantic Distance to Estimate Uncertainty in LLM-Based Code GenerationAn Execution-Verified Multi-Language Benchmark for Code Semantic ReasoningPrimal Generation, Dual Judgment: Self-Training from Test-Time ScalingStepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement LearningACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference OptimizationBeyond Retrieval: A Multitask Benchmark and Model for Code SearchDuET: Dual Execution for Test Output Prediction with Generated Code and PseudocodeCollabCoder: Plan-Code Co-Evolution via Collaborative Decision-Making for Efficient Code GenerationDefective Task Descriptions in LLM-Based Code Generation: Detection and AnalysisScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models$V_1$: Unifying Generation and Self-Verification for Parallel ReasonersReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement LearningScaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging ProblemsEnsemble-Based Uncertainty Estimation for Code Correctness EstimationThink Anywhere in Code GenerationIntentCoding: Amplifying User Intent in Code GenerationBridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code GenerationUnderstanding Specification-Driven Code Generation with LLMs: An Empirical Study DesignCodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case GenerationDAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code GenerationFunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code GenerationDreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM CodingDemystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution SimulationReducing Hallucinations in LLM-Generated Code via Semantic TriangulationCWM: An Open-Weights LLM for Research on Code Generation with World ModelsQueST: Incentivizing LLMs to Generate Difficult ProblemsInspectCoder: Dynamic Analysis-Enabled Self Repair through interactive LLM-Debugger CollaborationLarge Language Model enabled Mathematical ModelingChain of Execution Supervision Promotes General Reasoning in Large Language ModelsCWM: An Open-Weights LLM for Research on Code Generation with World
ModelsAgnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning EnvironmentCodeGrad: Integrating Multi-Step Verification with Gradient-Based LLM RefinementAlignment with Fill-In-the-Middle for Enhancing Code GenerationRethinking Verification for LLM Code Generation: From Generation to TestingOpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-CritiqueTurning the Tide: Repository-based Code ReflectionMemoCoder: Automated Function Synthesis using LLM-Supported AgentsRethinking Verification for LLM Code Generation: From Generation to
TestingOpenCodeReasoning-II: A Simple Test Time Scaling Approach via
Self-CritiqueICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming ContestsEliminating Hallucination-Induced Errors in LLM Code Generation with Functional ClusteringReVeal: Self-Evolving Code Agents via Reliable Self-VerificationWhich Data Attributes Stimulate Math and Code Reasoning? An
Investigation via Influence FunctionsAre Large Language Models Robust in Understanding Code Against Semantics-Preserving Mutations?CRPE: Expanding The Reasoning Capability of Large Language Model for Code GenerationAceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement LearningWhich Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence FunctionsrStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified DatasetAM-Thinking-v1: Advancing the Frontier of Reasoning at 32B ScaleAceReason-Nemotron: Advancing Math and Code Reasoning through
Reinforcement LearningrStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale
Verified DatasetSEW: Self-Evolving Agentic Workflows for Automated Code GenerationOpenCodeReasoning: Advancing Data Distillation for Competitive CodingA Self-Improving Coding AgentProgramming Language Confusion: When Code LLMs Can't Keep their Languages StraightThink Like Human Developers: Harnessing Community Knowledge for
Structured Code ReasoningACECODER: Acing Coder RL via Automated Test-Case SynthesisAssessing Correctness in LLM-Based Code Generation via Uncertainty EstimationS*: Test Time Scaling for Code GenerationLearning to Solve and Verify: A Self-Play Framework for Code and Test GenerationPlanning In Natural Language Improves LLM Search For Code GenerationHow Do Your Code LLMs Perform? Empowering Code Instruction Tuning with
High-Quality Data$\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases