Ministral 3 3B API Pricing: Mistral's Cheapest Model at $0.10/M
Ministral 3 3B offers symmetric pricing at $0.10/$0.10 — the same cost for input and output. Ideal for output-heavy tasks where other models' output premiums add up.
TL;DR
- Price: $0.10/M input, $0.10/M output — symmetric pricing, same cost for both
- Context: 128K tokens — standard for small models
- Size: 3B parameters — fast, efficient, edge-deployable
- Provider: Mistral API (api.mistral.ai) or self-host via Hugging Face
- Best for: Classification, extraction, output-heavy tasks, edge deployment
- Trade-off: Smaller model means less capable reasoning — not for complex tasks
Ministral 3 3B Pricing Breakdown
At $0.10 per million tokens for both input and output, Ministral 3 3B has the simplest pricing in the budget tier. Here's how the costs scale:
| Monthly Volume | Input Cost | Output Cost | Total Cost |
|---|---|---|---|
| 1M tokens | $0.10 | $0.10 | $0.20 |
| 10M tokens | $1.00 | $1.00 | $2.00 |
| 100M tokens | $10.00 | $10.00 | $20.00 |
| 1B tokens | $100.00 | $100.00 | $200.00 |
Compare this to GPT-5 nano ($0.05/$0.40): at 100M tokens with a 50/50 input/output split, Ministral costs $20 while GPT-5 nano costs $22.50. But if your workload is 30% input / 70% output, Ministral drops to $16 vs nano's $29 — a 45% savings.
Why Symmetric Pricing Matters
Most AI providers charge significantly more for output tokens than input. This penalizes tasks that generate lots of text. Here's how Ministral's symmetric pricing compares for output-heavy workloads:
| Model | Input $/M | Output $/M | Output:Input Ratio | Cost at 70% Output |
|---|---|---|---|---|
| Ministral 3 3B | $0.10 | $0.10 | 1:1 | $10.00 |
| GPT-5 nano | $0.05 | $0.40 | 8:1 | $16.50 |
| DeepSeek V4 Flash | $0.14 | $0.28 | 2:1 | $12.60 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 4:1 | $16.00 |
| Qwen 3.7 Flash | $0.03 | $0.13 | 4.3:1 | $5.20 |
At 100M tokens/month with 70% output, Ministral ($10) beats GPT-5 nano ($16.50) and Gemini 2.5 Flash-Lite ($16). Only Qwen 3.7 Flash ($5.20) and DeepSeek V4 Flash ($12.60) are competitive, but Qwen is from Alibaba and DeepSeek is from China — Ministral is the cheapest European/US option for output-heavy workloads.
Ministral 3 3B vs Other Budget Models
| Model | Input $/M | Output $/M | Context | Size | Provider |
|---|---|---|---|---|---|
| Qwen 3.7 Flash | $0.03 | $0.13 | 1M | — | Alibaba |
| GPT-5 nano | $0.05 | $0.40 | 128K | — | OpenAI |
| Ministral 3 3B | $0.10 | $0.10 | 128K | 3B | Mistral |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | — | |
| DeepSeek V4 Flash | $0.14 | $0.28 | 1M | — | DeepSeek |
| Ministral 3 8B | $0.15 | $0.15 | 128K | 8B | Mistral |
Among Mistral models, the Ministral family offers 3 tiers: 3B ($0.10/$0.10), 8B ($0.15/$0.15), and 14B ($0.20/$0.20). All have symmetric pricing. The 3B is cheapest; upgrade to 8B or 14B when you need better reasoning.
The Ministral Family: 3B vs 8B vs 14B
| Model | Input $/M | Output $/M | Context | Best For |
|---|---|---|---|---|
| Ministral 3 3B | $0.10 | $0.10 | 128K | Simple classification, extraction, edge |
| Ministral 3 8B | $0.15 | $0.15 | 128K | Moderate reasoning, code, analysis |
| Ministral 3 14B | $0.20 | $0.20 | 128K | Complex tasks, agentic workflows |
All three Ministral models share symmetric pricing and 128K context. The difference is capability: 3B handles simple tasks, 8B adds moderate reasoning, and 14B approaches larger model quality. The price premium is small ($0.10 → $0.20), so consider starting with 14B and scaling down if quality is sufficient.
When to Use Ministral 3 3B
📊 Classification
Sentiment analysis, intent detection, spam filtering. The 3B size handles straightforward classification fast and cheap.
📝 Output-Heavy Tasks
Summarization, translation, content generation. Symmetric pricing means you don't pay a premium for generating text.
🔍 Data Extraction
Entity extraction, structured data from unstructured text. Fast inference at the lowest per-token cost.
🤖 Edge Deployment
3B parameters fit on consumer GPUs. Deploy locally for zero-latency, zero-cost inference after setup.
💬 Simple Q&A
FAQ bots, help desk automation, knowledge base queries. Handles straightforward questions efficiently.
🔄 High-Volume Pipelines
Process millions of items per day. Low per-token cost and fast inference make high-volume processing affordable.
Real-World Cost: Summarizing 50,000 Articles/Month
Suppose you're building a news summarization service that processes 50,000 articles per month, averaging 2,000 input tokens and generating 500 output tokens per article:
| Model | Input Cost | Output Cost | Monthly Total |
|---|---|---|---|
| Ministral 3 3B | $10.00 | $2.50 | $12.50 |
| GPT-5 nano | $5.00 | $10.00 | $15.00 |
| Gemini 2.5 Flash-Lite | $10.00 | $10.00 | $20.00 |
| DeepSeek V4 Flash | $14.00 | $7.00 | $21.00 |
For this output-heavy workload (25% output ratio), Ministral 3 3B at $12.50/month is the cheapest option. Its symmetric pricing means you don't pay extra for generating summaries.
Compare 93 AI Models Side by Side
Ministral 3 3B is one of 93 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.
Frequently Asked Questions
Related Pages
- Mistral Provider Page — All Mistral models compared
- Qwen 3.7 Flash Pricing — The cheapest AI model at $0.03/$0.13
- GPT-5 nano Pricing — OpenAI's cheapest model at $0.05/$0.40
- Gemini Flash-Lite Pricing — Google's 3 budget tiers
- Top 10 Cheapest LLM APIs — Ranked by input price