The model cheatsheet
Which AI is best for which job?
There's no single "best" model — only the best one for the task in front of you. Here's what each is actually good at, ranked from real benchmark, human-preference, and price data. Not opinions.
The cheatsheet
Best models for each job
Pick the job, get the shortlist. Each list is ranked by the metric that actually matters for that task — and refreshes as new models land.
Reasoning & hard problems
Deep multi-step thinking, math, analysis
- 1 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) 60.1
- 2 GPT-5.5 (xhigh) 54.8
- 3 Gemini 3.5 Flash (high) 50.2
- 4 MO Kimi K3 (low) 46.6
- 5 MI MiniMax-M3 44.4
Ranked by Intelligence Index
Writing & shipping code
Generating, refactoring, and fixing code
- 1 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) 77.0
- 2 GPT-5.5 (xhigh) 74.9
- 3 MO Kimi K3 (low) 72.0
- 4 Gemini 3.5 Flash (high) 70.1
- 5 XI MiMo-V2.5-Pro 60.2
Ranked by Coding Index
Agents & tool use
Autonomous workflows that call tools
- 1 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) 54.5
- 2 GPT-5.6 Luna (max) 45.6
- 3 Gemini 3.5 Flash (high) 37.4
- 4 MO Kimi K3 (low) 37.0
- 5 DeepSeek V4 Pro (Reasoning, Max Effort) 36.4
Ranked by Agentic Index
General chat & writing
Everyday assistant, drafting, Q&A
- 1 claude-opus-4-6-thinking 1,503
- 2 gpt-5.4 1,499
- 3 gemini-3.6-flash 1,491
- 4 muse-spark-1.1 1,483
- 5 qwen3.7-max-preview 1,480
Ranked by LMArena (human votes)
Real-time & low latency
Voice, autocomplete, anything live
- 1 ST Step 3.7 Flash 413 tok/s
- 2 LA LFM2.5-VL-1.6B 401 tok/s
- 3 Gemini 3.5 Flash-Lite 368 tok/s
- 4 IB Granite 3.3 8B (Non-reasoning) 368 tok/s
- 5 gpt-oss-120b (low) 346 tok/s
Ranked by Output tokens/sec
Huge documents & long context
Whole codebases, books, long transcripts
- 1 Llama 4 Scout 17b 128e Instruct Maas 10M
- 2 Gemini Exp 1206 2.1M
- 3 Grok 4 Fast Reasoning 2M
- 4 GPT 5.5 1.1M
- 5 DA Databricks Gemini 2 5 Flash 1M
Ranked by Context window
Best bang for the buck
The most intelligence per dollar
- 1 Qwen3.5 4B (Non-reasoning) 266.7 pts/$
- 2 Gemma 4 E4B (Non-reasoning) 222.5 pts/$
- 3 DeepSeek V4 Flash (Reasoning, High Effort) 213.7 pts/$
- 4 XI MiMo-V2.5 212.6 pts/$
- 5 ST Step 3.5 Flash 2603 173.3 pts/$
Ranked by Intelligence per $/Mtok
High-volume on a budget
Cheap, good-enough, at scale
- 1 Llama 3.1 8b $0.035/Mtok
- 2 Meta Llama 3.2 1B Instruct $0.05/Mtok
- 3 CO Command R7b 12 2024 $0.066/Mtok
- 4 Mistral Small Latest $0.09/Mtok
- 5 ZA GLM 4 32b 0414 128k $0.1/Mtok
Ranked by Lowest blended $/Mtok
Ranked from live data · updated 8 hours ago. Model & provider names are trademarks of their owners, shown here only to report public benchmark and price data.
No favorites
How the picks are made
Benchmarks, not vibes
Reasoning, coding, agentic, speed and value come from independent Artificial Analysis indices. Chat is the LMArena leaderboard — millions of blind human votes. Context and budget come from the live price catalog.
Self-updating
Nothing is hand-picked. When a new model tops a benchmark or a price changes, the cheatsheet re-ranks itself on the next daily sync. No stale "best of 2024" lists.
One axis at a time
A model can win one job and lose another. We rank each category by the single metric that matters for it, so the shortlist is honest about trade-offs.
Quality & speed from Artificial Analysis; human preference from LMArena; prices from the MyTokenTracker catalog. See the full methodology.
Citation
Use this in your work
Open data, free to cite. Pair it with the price-vs-cost breakdown and the full State of AI.
Copy a citation
Free to use and cite under CC BY 4.0. See how this is measured.
Champlin Enterprises. (2026). Which AI for which job — the model cheatsheet (MyTokenTracker) [Data set]. MyTokenTracker. Retrieved July 31, 2026, from https://mytokentracker.io/which-ai
@misc{mytokentracker-which-ai,
title = {Which AI for which job — the model cheatsheet (MyTokenTracker)},
author = {{Champlin Enterprises}},
year = {2026},
howpublished = {MyTokenTracker, \url{https://mytokentracker.io/which-ai}},
note = {Accessed July 31, 2026. Licensed CC BY 4.0.},
url = {https://mytokentracker.io/which-ai}
}
Need a fixed point in time? Every day’s data is permanently archived in the open-data repository, so you can cite a specific date by linking that day’s committed file.
Free weekly digest
The best model keeps changing
New models top these lists every few weeks. Get the weekly digest — what moved, what's now best for what, and what it costs. Free, no account.