Awesome AI for Code
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
benchmark
loadingβ¦
π€
Ask AI
Awesome benchmark β curated papers, datasets & benchmarks Β· Awesome AI for Code
β all topics
overview
benchmark
19 papers tagged benchmark β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
19 papers Β· trending (default)
numbers = π₯ heat
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
(2026)
Michael Solodko et al.
3.51
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
(2026)
Bytedance Seed
2.00
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
(2026)
Shijie Li et al.
2.00
From Foundation to Application: Improving VLA Models in Practice
(2026)
Wei Wu et al.
2.00
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
(2026)
Kaifeng Zhao et al.
2.00
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
(2026)
Tongkun Guan et al.
1.94
EnvX: Agentize Everything with Agentic AI
(2025)
Linyao Chen et al.
1.44
RExBench: Can coding agents autonomously implement AI research extensions?
(2025)
Nicholas Edwards et al.
1.28
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
(2025)
Hao Liang et al.
1.06
DeepSeek-V3 Technical Report
(2024)
DeepSeek-AI et al.
β
WebArena: A Realistic Web Environment for Building Autonomous Agents
(2023)
Shuyan Zhou et al.
β
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
(2024)
Zhihong Shao et al.
β
Pre-training Small Base LMs with Fewer Tokens
(2024)
Sunny Sanyal et al.
β
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
(2024)
Huajian Xin et al.
β
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
(2024)
Wenhao Shi et al.
β
SciCode: A Research Coding Benchmark Curated by Scientists
(2024)
Minyang Tian et al.
β
MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
(2024)
Pei Wang et al.
β
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
(2024)
Jiale Cheng et al.
β
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
(2024)
Zihan Liu et al.
β