Lowest Input Price
$0.075
Gemini 2.5 Flash-Lite ($0.075 in / $0.30 out)

Calculate Your Exact Costs

Enter your token usage and see exactly how much each of these 15 models costs for your workload.

Open Cost Calculator →

Price Tier Breakdown

These 15 models split into three clear tiers based on input pricing. Here's what you get at each level.

Ultra-Cheap

Under $0.15/M input

  • GPT-oss 20B at $0.08/$0.35 is the cheapest model
  • GPT-oss 20B — $0.08/$0.35 · 128K context · OpenAI's open-source entry
  • Llama 3.1 8B — $0.10/$0.10 · 128K context · Open source, lowest total cost
  • Gemini 2.5 Flash-Lite — $0.10/$0.40 · 1M context · Google's balanced budget option
  • DeepSeek V4 Flash — $0.14/$0.28 · 1M context · Best budget coding model

Best for: high-volume classification, simple Q&A, data extraction, embedding pipelines

Cheap

$0.15 — $0.30/M input

  • GPT-oss 120B — $0.15/$0.60 · 128K context · Strong general-purpose
  • GPT-4o mini — $0.15/$0.60 · 128K context · OpenAI's budget workhorse
  • Mistral Small 4 ($0.15/$0.60) · 128K context · EU data sovereignty
  • Llama 4 Scout — $0.18/$0.59 · 1M context · Open source, MIT license
  • DeepSeek V3.2 — $0.23/$0.34 · 128K context · Proven production model
  • Llama 4 Maverick — $0.27/$0.85 · 1M context · Open source flagship

Best for: chatbots, content generation, RAG pipelines, code assistance

Budget-Friendly

$0.40 — $1.00/M input

  • DeepSeek V4 Pro — $0.435/$0.87 · 1M context · Best value for complex tasks
  • Mistral Large 3 — $0.50/$1.50 · 262K context · Strong at RAG and retrieval
  • Command R — $0.50/$1.50 · 128K context · Cohere's enterprise RAG model
  • Kimi K2.6 — $0.95/$4.00 · 256K context · Excellent reasoning capabilities

Best for: code generation, complex analysis, RAG, nuanced writing

Use Case Recommendations

Different tasks need different models. Here's the best sub-$1 model for each major use case.

💬

Chatbot

DeepSeek V4 Flash

$0.14/$0.28 — cheapest model that handles multi-turn conversations naturally. Used in production by thousands of apps.

💻

Code Generation

DeepSeek V4 Pro

$0.435/$0.87 — outperforms GPT-4o on coding benchmarks at 80% less cost. Best value coding model under $1.

📚

RAG Pipeline

Command R

$0.50/$1.50 — purpose-built for retrieval-augmented generation with strong context following and factual accuracy.

✍️

Content Writing

Kimi K2.6

$0.95/$4.00 — excellent reasoning and long-form generation. The most capable writing model under $1/M input.

📊

Data Extraction

Llama 3.1 8B

$0.10/$0.10 — lowest total cost at $0.20/M. Perfect for high-volume structured extraction tasks.

🌐

Multilingual

Mistral Small 4

$0.15/$0.60 — strong multilingual support with EU data sovereignty. Handles 30+ languages well.

🔒

No Vendor Lock-in

Llama 4 Scout

$0.18/$0.59 — open source MIT license, self-hostable, 1M context window. Full control over your stack.

Highest Volume

Gemini 2.5 Flash-Lite

$0.075/$0.30 — cheapest input price of any model. When you need to process millions of tokens daily.

Track Every Dollar with APIpulse

Set cost alerts, compare models in real-time, and optimize your API spend across all 15 budget models. Free forever.

Try APIpulse Free →

Provider Breakdown

Eight providers offer models under $1/M input in July 2026. Here's how they compare.

  • Google — 2 models. Cheapest input price with Gemini Flash Lite ($0.075). Both models have 1M context windows.
  • OpenAI — 3 models. GPT-oss 20B and 120B are open-source. GPT-4o mini is the industry standard budget model.
  • Meta — 3 models. Llama 3.1 8B is the lowest total cost. Llama 4 Scout and Maverick both have 1M context, MIT license.
  • DeepSeek — 3 models. V4 Flash ($0.14) is best for coding. V3.2 is proven in production. V4 Pro is best value for complex tasks.
  • Mistral — 2 models. EU-based with data sovereignty. Mistral Small 4 competes directly with GPT-4o mini.
  • Cohere — 1 model. Command R at $0.50 is purpose-built for RAG and enterprise search.
  • Moonshot — 1 model. Kimi K2.6 at $0.95 is the most capable reasoning model under $1/M input.

Compare All 53 Models Side by Side

Our comparison tool lets you filter by price, context window, provider, and capabilities across every tracked model.

Open Comparison Tool →

The Bottom Line

The era of expensive AI is over. With 15 models under $1/M input tokens, every startup and indie developer can afford production-quality AI. The cheapest option, Gemini 2.5 Flash-Lite at $0.075/M, lets you process 13 million tokens for a dollar.

Here's the quick decision tree for choosing among these 15 models:

  • Lowest possible cost → Llama 3.1 8B ($0.10/$0.10, total $0.20)
  • Cheapest input for high volume → Gemini 2.5 Flash-Lite ($0.075 input)
  • Best production chatbot → DeepSeek V4 Flash ($0.14/$0.28)
  • Best budget coding → DeepSeek V4 Pro ($0.435/$0.87)
  • No vendor lock-in → Llama 4 Scout ($0.18/$0.59, MIT license)
  • EU data sovereignty → Mistral Small 4 ($0.15/$0.60)
  • Best RAG on a budget → Command R ($0.50/$1.50)
  • Best reasoning under $1 → Kimi K2.6 ($0.95/$4.00)

Use the APIpulse cost calculator to model your exact usage and find the cheapest model that meets your quality bar.

💸 Looking for DeepSeek V4 Flash Alternatives?
5 models ranked by cost — some offer better quality at similar prices.
See 5 DeepSeek V4 Flash Alternatives →
💸 Looking for Llama 4 Maverick Alternatives?
5 models ranked by cost — some are 95% cheaper.
See 5 Llama 4 Maverick Alternatives →
💸 Looking for Mistral Small 4 Alternatives?
5 models ranked by cost — some are 90% cheaper.
See 5 Mistral Small 4 Alternatives →
💸 Looking for Llama 4 Scout Alternatives?
5 models ranked by cost — some are 95% cheaper.
See 5 Llama 4 Scout Alternatives →
🔧 Free Embeddable Pricing Widget
Add live AI API pricing to your docs, blog, or README with one script tag. 87 models, auto-updating.
Get the Free Widget → Free MCP Server →

Stop guessing — get exact API cost comparisons

No signup required to 67-model comparison, migration code snippets, PDF reports, price alerts, and cost monitoring. ✅ All tools free.

Free Tools →