Updated weekly · Oct 3, 2026
Which LLM is #1 right now?
We rank every major model on the things you actually care about — code that works, writing that sounds human, answers that are right, and what it costs you per million tokens.
Muse Spark 1.3
Muse Spark 1.3 scores 48.1 on the Artificial Analysis Intelligence Index (highest reasoning-effort setting). List price $1.25 input / $4.25 output per 1M tokens.
- Overall score
- 83.2
- Input $/M
- $1.25
- Output $/M
- $4.25
- Context
- 1M
Overall leaderboard
19 of 19 models · sort any column
| Coding | Writing | Reasoning | Speed | Cost-efficiency | Long | Multimodal | Open-source | Weights | Updated | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Muse Spark 1.3 Meta | 83.2 | $1.25 | $4.25 | 1M | 6 | — | 7 | 4 | 6 | — | — | — | Closed | Oct 3, 2026 |
| 2 | Claude Sonnet 5.5 Anthropic | 81.9 | $2.00 | $10.00 | 1M | — | — | 2 | 5 | 12 | — | — | — | Closed | Oct 3, 2026 |
| 3 | Gemini 4 Argon Google DeepMind | 81.9 | $2.00 | $10.00 | 1M | — | — | 5 | — | 13 | — | — | — | Closed | Oct 3, 2026 |
| 4 | Claude Fable 5.1 Anthropic | 81.7 | $10.00 | $50.00 | 1M | 1 | — | 3 | 13 | 18 | — | — | — | Closed | Oct 3, 2026 |
| 5 | Gemini 3.8 Flash Google DeepMind | 81.7 | $0.75 | $3.75 | 1M | 3 | — | 14 | 1 | 7 | — | — | — | Closed | Oct 3, 2026 |
| 6 | Claude Opus 5.5 Anthropic | 78.7 | $4.00 | $20.00 | 1M | — | — | 1 | 8 | 16 | — | — | — | Closed | Oct 3, 2026 |
| 7 | GPT-6 Astra OpenAI | 78.4 | $10.00 | $50.00 | 1M | 2 | — | 4 | 15 | 19 | — | — | — | Closed | Oct 3, 2026 |
| 8 | GLM-5.3 z.ai | 76.4 | $1.40 | $4.40 | 1M | 7 | — | 11 | 12 | 8 | — | — | 1 | Open | Oct 3, 2026 |
| 9 | Qwen3.8 Max Alibaba Qwen | 75.5 | $2.00 | $6.00 | 984K | 4 | — | 10 | 17 | 11 | — | — | — | Closed | Oct 3, 2026 |
| 10 | DeepSeek V4.1 Flash DeepSeek | 74.1 | $0.30 | $1.20 | 1M | — | — | 15 | 2 | 3 | — | — | 3 | Open | Oct 3, 2026 |
| 11 | Kimi K3 Moonshot AI | 72.4 | $3.00 | $15.00 | 1.1M | 5 | — | 13 | 18 | 15 | — | — | 2 | Open | Oct 3, 2026 |
| 12 | MiMo-V2.6-Pro Xiaomi | 71.7 | $0.43 | $0.87 | 1M | — | — | 9 | 16 | 2 | — | — | — | Closed | Oct 3, 2026 |
| 13 | GPT-6.1 Sol OpenAI | 71.4 | $2.00 | $10.00 | 1M | — | — | 6 | 14 | 14 | — | — | — | Closed | Oct 3, 2026 |
| 14 | GPT-6 Luna OpenAI | 70.2 | $0.10 | $0.50 | 1M | — | — | 16 | 6 | 1 | — | — | — | Closed | Oct 3, 2026 |
| 15 | DeepSeek V4 Pro DeepSeek | 68.5 | $1.32 | $3.96 | 1M | 8 | — | 17 | 7 | 9 | — | — | 4 | Open | Oct 3, 2026 |
| 16 | Grok 4.7 xAI | 67.8 | $2.00 | $6.00 | 500K | — | — | 8 | 10 | 10 | — | — | — | Closed | Oct 3, 2026 |
| 17 | Step 5 Preview StepFun | 67.2 | $1.00 | $2.70 | 1M | — | — | 12 | 9 | 5 | — | — | — | Closed | Oct 3, 2026 |
| 18 | MiniMax-M3 MiniMax | 59.6 | $0.30 | $1.20 | 1M | 9 | — | 18 | 11 | 4 | — | — | 5 | Open | Oct 3, 2026 |
| 19 | Mistral Medium 3.5 Mistral AI | 42.3 | $1.50 | $7.50 | — | 10 | — | 19 | 3 | 17 | — | — | — | Closed | Oct 3, 2026 |
- In $/M
- $1.25
- Out $/M
- $4.25
- Context
- 1M
- In $/M
- $2.00
- Out $/M
- $10.00
- Context
- 1M
- In $/M
- $2.00
- Out $/M
- $10.00
- Context
- 1M
- In $/M
- $10.00
- Out $/M
- $50.00
- Context
- 1M
- In $/M
- $0.75
- Out $/M
- $3.75
- Context
- 1M
- In $/M
- $4.00
- Out $/M
- $20.00
- Context
- 1M
- In $/M
- $10.00
- Out $/M
- $50.00
- Context
- 1M
- In $/M
- $1.40
- Out $/M
- $4.40
- Context
- 1M
- In $/M
- $2.00
- Out $/M
- $6.00
- Context
- 984K
- In $/M
- $0.30
- Out $/M
- $1.20
- Context
- 1M
- In $/M
- $3.00
- Out $/M
- $15.00
- Context
- 1.1M
- In $/M
- $0.43
- Out $/M
- $0.87
- Context
- 1M
- In $/M
- $2.00
- Out $/M
- $10.00
- Context
- 1M
- In $/M
- $0.10
- Out $/M
- $0.50
- Context
- 1M
- In $/M
- $1.32
- Out $/M
- $3.96
- Context
- 1M
- In $/M
- $2.00
- Out $/M
- $6.00
- Context
- 500K
- In $/M
- $1.00
- Out $/M
- $2.70
- Context
- 1M
- In $/M
- $0.30
- Out $/M
- $1.20
- Context
- 1M
- In $/M
- $1.50
- Out $/M
- $7.50
- Context
- —
Claude Fable 5.1 has the highest Artificial Analysis Coding Index of the models we track — an independent test of writing and fixing real code.
Claude Opus 5.5 has the highest Artificial Analysis Intelligence Index of the models we track, which combines ten hard tests of maths, science, coding and reasoning.
Gemini 3.8 Flash produces answers faster than any other model we track, based on Artificial Analysis's measured median output speed.
Get the weekly LLM rankings
One email a week: what moved, what's new, and which model to actually use for your work. No hype, no benchmark jargon.