How we rank

Models are ranked by blended cost using a 1:3 input-to-output ratio (typical for chat/assistant workloads). A model costing $0.10/M input + $0.30/M output scores at $1.00/M blended. Context window and provider are noted but don't affect ranking. All prices sourced directly from provider pricing pages.

Price trends: โ†“ Dropped = price decreased recently ยท โ†’ Stable = no change in past 30 days ยท โ†‘ Increased = price went up. Trends reflect changes since May 2026.

#1

Qwen 3.7 Flash

Alibaba's cheapest multimodal model. 1M context with vision support at the lowest price in the industry.

Alibaba/Qwen 1M context Cheapest model overall ($0.03/M input) โ†“ Cheapest
Input
$0.03 /M tokens
Output
$0.13 /M tokens
#2

GPT-5 nano

OpenAI's ultra-budget model. Cheapest OpenAI option at $0.05/M input.

OpenAI 128K context Cheapest OpenAI model โ†“ New (Jul)
Input
$0.05 /M tokens
Output
$0.40 /M tokens
#3

GPT-oss 20B

OpenAI's open-source 20B model. Self-host or use via Hugging Face at $0.08/M input.

OpenAI 128K context Open-source, self-hostable โ†’ Stable
Input
$0.08 /M tokens
Output
$0.35 /M tokens
#4

GPT-4.1 nano

OpenAI's nano-tier model. 1M context at budget pricing โ€” ideal for bulk processing.

OpenAI 1M context Best value 1M context from OpenAI โ†’ Stable
Input
$0.10 /M tokens
Output
$0.40 /M tokens
#5

Gemini 2.5 Flash-Lite

Google's budget model. 1M context with vision support at $0.10/M input.

Google 1M context Cheapest Google model โ†’ Stable
Input
$0.10 /M tokens
Output
$0.40 /M tokens
#6

Ministral 3 3B

Mistral's cheapest edge model. Equal input/output at $0.10/M โ€” great for high-volume symmetric workloads.

Mistral 128K context Cheapest output price ($0.10/M) โ†“ New
Input
$0.10 /M tokens
Output
$0.10 /M tokens
#7

DeepSeek V4 Flash

DeepSeek's speed-optimized model. Excellent blended cost at $0.22/M input.

DeepSeek 128K context Best blended cost ratio โ†’ Stable
Input
$0.22 /M tokens
Output
$0.28 /M tokens
#8

GPT-oss 120B

OpenAI's larger open-source model. More capable than the 20B at still-reasonable pricing.

OpenAI 128K context Open-source, self-hostable โ†’ Stable
Input
$0.15 /M tokens
Output
$0.60 /M tokens
#9

Mistral Small 4

Mistral's efficient small model. Competitive pricing with 128K context.

Mistral 128K context Good quality-to-price ratio โ†’ Stable
Input
$0.15 /M tokens
Output
$0.60 /M tokens
#10

GPT-5.6 Luna

OpenAI's budget long-context model. 1.05M context at $0.20/M input โ€” 80% price cut in July 2026.

OpenAI 1.05M context Cheapest OpenAI model with 1M+ context โ†“ 80% price cut
Input
$0.20 /M tokens
Output
$1.20 /M tokens

How much could you save?

Enter your monthly token usage to see the cost difference between the cheapest and most expensive models.

Qwen 3.7 Flash (cheapest)
$42.00
GPT-5.5 Pro (most expensive)
$57,000

You'd save $56,958/mo by choosing the cheapest model.

Cheapest Model Per Provider

Alibaba/Qwen

Cheapest: $0.03/M input
Model: Qwen 3.7 Flash
1 model

OpenAI

Cheapest: $0.05/M input
Model: GPT-5 nano
22 models total

Google

Cheapest: $0.10/M input
Model: Gemini 2.5 Flash-Lite
11 models total

Mistral

Cheapest: $0.10/M input
Model: Ministral 3 3B
5 models total

DeepSeek

Cheapest: $0.22/M input
Model: DeepSeek V4 Flash
4 models total

Meta (Together.ai)

Cheapest: $0.18/M input
Model: Llama 4 Scout
3 models total

AI21

Cheapest: $0.20/M input
Model: Jamba Mini
3 models total

xAI

Cheapest: $0.30/M input
Model: Grok Build 0.1
3 models total

Cohere

Cheapest: $0.50/M input
Model: Command A
3 models total

Anthropic

Cheapest: $1.00/M input
Model: Claude Haiku 4.5
12 models total

Moonshot

Cheapest: $0.95/M input
Model: Kimi K2.7 Code
3 models total

Calculate your exact costs

Use our free calculator to compare all 95 models and find the cheapest option for your specific usage pattern.

Check My Costs โ€” Free Compare Any Two Models