Output Speed (Artificial Analysis) aa-output-speed Leaderboard
Median output throughput (tokens/sec) across providers, from the Artificial Analysis API. Higher is faster. Β· Metric: Output Speed (higher is better) Β· π’ Updated 17h ago
| # | Model | Output Speed | Paper |
|---|---|---|---|
| 1 | Mercury 2 | 1118.92 | β |
| 2 | HyperNova 60B 2605 | 420.96 | β |
| 3 | Granite 4.0 H Small | 406.40 | β |
| 4 | Step 3.7 Flash | 397.59 | β |
| 5 | LFM2.5-VL-1.6B | 364.75 | β |
| 6 | Gemini 3.5 Flash-Lite | 362.19 | β |
| 7 | LFM2.5-8B-A1B | 331.57 | β |
| 8 | gpt-oss-120b (low) | 323.73 | β |
| 9 | Nemotron 3 Nano Omni 30B A3B Reasoning | 322.69 | β |
| 10 | Gemini 3.1 Flash-Lite | 298.64 | β |
| 11 | Step 3.5 Flash 2603 | 297.02 | β |
| 12 | Nova Micro | 289.82 | β |
| 13 | gpt-oss-120b (high) | 275.72 | β |
| 14 | gpt-oss-20b (low) | 273.04 | β |
| 15 | Gemini 3.5 Flash (medium) | 265.21 | β |
| 16 | Gemini 3.5 Flash (high) | 250.26 | β |
| 17 | Ministral 3 3B | 248.58 | β |
| 18 | gpt-oss-20b (high) | 242.74 | β |
| 19 | Gemini 3.5 Flash (minimal) | 228.30 | β |
| 20 | Qwen3.5 Omni Flash | 226.36 | β |
| 21 | Nova 2.0 Lite (Non-reasoning) | 225.16 | β |
| 22 | Gemini 3.6 Flash (high) | 219.36 | β |
| 23 | Qwen3.7 Max | 199.57 | β |
| 24 | Command A+ | 197.45 | β |
| 25 | NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) | 195.86 | β |
| 26 | Qwen3 Next 80B A3B Instruct | 190.67 | β |
| 27 | Qwen3 Next 80B A3B (Reasoning) | 182.11 | β |
| 28 | Nemotron 3 Ultra 550B A55B (Reasoning) | 180.38 | β |
| 29 | Nova 2.0 Lite (low) | 179.28 | β |
| 30 | Nova 2.0 Lite (high) | 175.55 | β |