QuixBugs
Emerging21papers using it
2018first seen
QuixBugs is a benchmark dataset used to evaluate the bug-fixing accuracy of small language models in automated program repair tasks.
Papers using QuixBugs (21)
- LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83
Software Engineering BenchmarksMokav: Execution-driven Differential Testing with LLMsThe Impact of Fine-tuning Large Language Models on Automated Program RepairBoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward ModelsHow Small is Enough? Empirical Evidence of Quantized Small Language Models for Automated Program RepairExploring and Lifting the Robustness of LLM-powered Automated Program
Repair with Metamorphic TestingContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case PairsAn Analysis of the Automatic Bug Fixing Performance of ChatGPTCURE: Code-Aware Neural Machine Translation for Automatic Program RepairKNOD: Domain Knowledge Distilled Tree Decoder for Automated Program RepairNuances are the Key: Unlocking ChatGPT to Find Failure-Inducing Tests
with Differential PromptingGAMMA: Revisiting Template-based Automated Program Repair via Mask
PredictionThinkRepair: Self-Directed Automated Program RepairDeepDebug: Fixing Python Bugs Using Stack Traces, Backtranslation, and
Code SkeletonsENCORE: Ensemble Learning using Convolution Neural Machine Translation for Automatic Program RepairPrompting Code Interpreter to Write Better Unit Tests on Quixbugs
FunctionsAutomatic Program Repair with OpenAI's Codex: Evaluating QuixBugsCan GPT-O1 Kill All Bugs? An Evaluation of GPT-Family LLMs on QuixBugsA Comprehensive Study of Automatic Program Repair on the QuixBugs
BenchmarkA Quick Repair Facility for DebuggingLecPrompt: A Prompt-based Approach for Logical Error Correction with
CodeBERT