Value for money

What you actually get for your dollar

Price is only half the story. Below: the lowest cost per million tokens across 2,308 models, and, as the community grows, the real cost per successful task. The cheapest model is not always the best value.

Lowest cost per million tokens

Blended price at a fixed 3:1 input-to-output ratio, text models, from the live catalog.

# Model Provider Input Output Blended
1 nscale/Qwen/Qwen2.5-Coder-3B-Instruct nscale $0.01 $0.03 $0.015
2 nscale/Qwen/Qwen2.5-Coder-7B-Instruct nscale $0.01 $0.03 $0.015
3 nebius/Qwen/Qwen2.5-Coder-7B nebius $0.01 $0.03 $0.015
4 lambda_ai/llama3.2-11b-vision-instruct lambda_ai $0.02 $0.03 $0.018
5 lambda_ai/llama3.2-3b-instruct lambda_ai $0.02 $0.03 $0.018
6 deepinfra/meta-llama/Llama-3.2-3B-Instruct deepinfra $0.02 $0.02 $0.020
7 novita/paddlepaddle/paddleocr-vl novita $0.02 $0.02 $0.020
8 deepinfra/meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo deepinfra $0.02 $0.03 $0.023
9 deepinfra/mistralai/Mistral-Nemo-Instruct-2407 deepinfra $0.02 $0.04 $0.025
10 nscale/deepseek-ai/DeepSeek-R1-Distill-Llama-8B nscale $0.03 $0.03 $0.025
11 novita/meta-llama/llama-3.1-8b-instruct novita $0.02 $0.05 $0.028
12 darkbloom/gpt-oss-20b darkbloom $0.01 $0.07 $0.028
13 lambda_ai/hermes3-8b lambda_ai $0.03 $0.04 $0.029
14 lambda_ai/lfm-7b lambda_ai $0.03 $0.04 $0.029
15 lambda_ai/llama3.1-8b-instruct lambda_ai $0.03 $0.04 $0.029
16 nscale/meta-llama/Llama-3.1-8B-Instruct nscale $0.03 $0.03 $0.030
17 nebius/meta-llama/Meta-Llama-3.1-8B-Instruct nebius $0.02 $0.06 $0.030
18 nebius/Qwen/Qwen2-VL-7B-Instruct nebius $0.02 $0.06 $0.030
19 novita/deepseek/deepseek-ocr novita $0.03 $0.03 $0.030
20 novita/qwen/qwen3-4b-fp8 novita $0.03 $0.03 $0.030

Lowest price is not the same as best value. A cheaper model that fails more often can cost more per finished task. That is what the next section measures.

From the community

Real-world value for money

Cost per successful task, success rate, average quality, and latency, aggregated across everyone who opts in. Anonymized, and only shown when enough people contribute. Last 30 days.

This fills in as the community grows

Send the optional success, quality_score, and latency_ms fields with your events and we can show which models are actually worth it, not just which are cheapest. The more people who do, the sharper this gets.

Cost per success = total spend divided by tasks reported successful. We never show a model unless enough separate contributors reported it. See the community dashboard for usage totals.