LiveCodeBench livecodebench Leaderboard
Contamination-free competitive-programming eval (LeetCode/AtCoder/ Codeforces) with a rolling problem window. Live from the Artificial Analysis API. Β· Metric: Accuracy (higher is better) Β· π’ Updated 10h ago
| # | Model | Accuracy | Paper |
|---|---|---|---|
| 1 | Gemini 3 Pro Preview (high) | 91.70 | β |
| 2 | Gemini 3 Flash Preview (Reasoning) | 90.80 | β |
| 3 | DeepSeek V3.2 Speciale | 89.60 | β |
| 4 | GLM-4.7 (Reasoning) | 89.40 | β |
| 5 | GPT-5.2 (medium) | 89.40 | β |
| 6 | GPT-5.2 (xhigh) | 88.90 | β |
| 7 | gpt-oss-120b (high) | 87.80 | β |
| 8 | Claude Opus 4.5 (Reasoning) | 87.10 | β |
| 9 | GPT-5.1 (high) | 86.80 | β |
| 10 | MiMo-V2-Flash (Reasoning) | 86.80 | β |
| 11 | DeepSeek V3.2 (Reasoning) | 86.20 | β |
| 12 | o4-mini (high) | 85.90 | β |
| 13 | Gemini 3 Pro Preview (low) | 85.70 | β |
| 14 | Kimi K2 Thinking | 85.30 | β |
| 15 | GPT-5.1 Codex (high) | 84.90 | β |
| 16 | GPT-5 (high) | 84.60 | β |
| 17 | GPT-5 Codex (high) | 84.00 | β |
| 18 | GPT-5 mini (high) | 83.80 | β |
| 19 | GPT-5.1 Codex mini (high) | 83.60 | β |
| 20 | Grok 4 Fast (Reasoning) | 83.20 | β |
| 21 | MiniMax-M2 | 82.60 | β |
| 22 | Grok 4.1 Fast (Reasoning) | 82.20 | β |
| 23 | Grok 4 | 81.90 | β |
| 24 | ERNIE 5.0 Thinking Preview | 81.20 | β |
| 25 | MiniMax-M2.1 | 81.00 | β |
| 26 | o3 | 80.80 | β |
| 27 | Apriel-v1.6-15B-Thinker | 80.70 | β |
| 28 | Gemini 2.5 Pro | 80.10 | β |
| 29 | DeepSeek V3.1 Terminus (Reasoning) | 79.80 | β |
| 30 | Gemini 3 Flash Preview (Non-reasoning) | 79.70 | β |