Ministral 3 8B API Pricing: Symmetric $0.15/M for Better Reasoning
Ministral 3 8B offers symmetric pricing at $0.15/$0.15 — the same cost for input and output. More capable than 3B at only 50% higher price, making it the sweet spot in the Ministral family.
TL;DR
- Price: $0.15/M input, $0.15/M output — symmetric pricing, same cost for both
- Context: 128K tokens — standard for small models
- Size: 8B parameters — significantly better reasoning than 3B
- Provider: Mistral API (api.mistral.ai) or self-host via Hugging Face
- Best for: Code generation, analysis, moderate reasoning, output-heavy tasks
- Trade-off: 128K context only — not for ultra-long documents
Ministral 3 8B Pricing Breakdown
At $0.15 per million tokens for both input and output, here's how the costs scale:
| Monthly Volume | Input Cost | Output Cost | Total Cost |
|---|---|---|---|
| 1M tokens | $0.15 | $0.15 | $0.30 |
| 10M tokens | $1.50 | $1.50 | $3.00 |
| 100M tokens | $15.00 | $15.00 | $30.00 |
| 1B tokens | $150.00 | $150.00 | $300.00 |
Compare to Ministral 3 3B at $0.10/$0.10: the 8B costs 50% more but delivers significantly better reasoning. For most production workloads, the quality improvement easily justifies the small price increase.
Why Symmetric Pricing Matters
Most AI providers charge significantly more for output tokens than input. This penalizes tasks that generate lots of text. Here's how Ministral 3 8B's symmetric pricing compares for output-heavy workloads:
| Model | Input $/M | Output $/M | Output:Input Ratio | Cost at 70% Output |
|---|---|---|---|---|
| Ministral 3 8B | $0.15 | $0.15 | 1:1 | $15.00 |
| Ministral 3 3B | $0.10 | $0.10 | 1:1 | $10.00 |
| GPT-5 nano | $0.05 | $0.40 | 8:1 | $16.50 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 4:1 | $16.00 |
| DeepSeek V4 Flash | $0.14 | $0.28 | 2:1 | $12.60 |
At 100M tokens/month with 70% output, Ministral 3 8B ($15) is competitive with GPT-5 nano ($16.50) and Gemini 2.5 Flash-Lite ($16). But Ministral offers symmetric pricing — so the more output you generate, the more you save compared to asymmetric models.
Ministral 3 8B vs Other Budget Models
| Model | Input $/M | Output $/M | Context | Size | Provider |
|---|---|---|---|---|---|
| Qwen 3.7 Flash | $0.03 | $0.13 | 1M | — | Alibaba |
| GPT-5 nano | $0.05 | $0.40 | 128K | — | OpenAI |
| Ministral 3 3B | $0.10 | $0.10 | 128K | 3B | Mistral |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | — | |
| DeepSeek V4 Flash | $0.14 | $0.28 | 1M | — | DeepSeek |
| Ministral 3 8B | $0.15 | $0.15 | 128K | 8B | Mistral |
| Ministral 3 14B | $0.20 | $0.20 | 128K | 14B | Mistral |
Among Mistral models, the Ministral family offers 3 tiers: 3B ($0.10/$0.10), 8B ($0.15/$0.15), and 14B ($0.20/$0.20). All have symmetric pricing. The 8B is the sweet spot — better reasoning than 3B at only 50% more cost.
The Ministral Family: 3B vs 8B vs 14B
| Model | Input $/M | Output $/M | Context | Best For |
|---|---|---|---|---|
| Ministral 3 3B | $0.10 | $0.10 | 128K | Simple classification, extraction, edge |
| Ministral 3 8B | $0.15 | $0.15 | 128K | Moderate reasoning, code, analysis |
| Ministral 3 14B | $0.20 | $0.20 | 128K | Complex tasks, agentic workflows |
All three Ministral models share symmetric pricing and 128K context. The difference is capability: 3B handles simple tasks, 8B adds moderate reasoning and code generation, and 14B approaches larger model quality. The 8B is the best balance of cost and capability for most workloads.
When to Use Ministral 3 8B
💻 Code Generation
Generate code snippets, functions, or boilerplate. The 8B size handles programming tasks significantly better than 3B at minimal extra cost.
📝 Output-Heavy Tasks
Summarization, translation, content generation. Symmetric pricing means you don't pay a premium for generating text.
🔍 Data Analysis
Analyze documents, extract insights, answer questions. 8B parameters provide better reasoning for analytical tasks.
🤖 Chatbots
Build conversational AI with better reasoning than 3B. The 128K context handles multi-turn conversations easily.
📊 Classification
Sentiment analysis, intent detection, categorization. More accurate than 3B for nuanced classification tasks.
🔄 High-Volume Pipelines
Process millions of items per day. Low per-token cost and good quality make high-volume processing affordable.
Real-World Cost: Generating 50,000 Code Snippets/Month
Suppose you're building a code generation service that produces 50,000 code snippets per month, averaging 500 input tokens and 1,000 output tokens per snippet:
| Model | Input Cost | Output Cost | Monthly Total |
|---|---|---|---|
| Ministral 3 8B | $3.75 | $7.50 | $11.25 |
| Ministral 3 3B | $2.50 | $5.00 | $7.50 |
| GPT-5 nano | $1.25 | $20.00 | $21.25 |
| Gemini 2.5 Flash-Lite | $2.50 | $20.00 | $22.50 |
| DeepSeek V4 Flash | $3.50 | $14.00 | $17.50 |
For this output-heavy workload (67% output), Ministral 3 8B at $11.25/month is significantly cheaper than GPT-5 nano ($21.25) or Gemini Flash-Lite ($22.50). The symmetric pricing saves you money on every generated token.
Compare 93 AI Models Side by Side
Ministral 3 8B is one of 93 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.
Frequently Asked Questions
Related Pages
- Ministral 3 3B Pricing — Cheaper at $0.10/$0.10 but less capable
- GPT-5 nano Pricing — OpenAI's cheapest at $0.05/$0.40
- Gemini 2.5 Flash-Lite Pricing — Google's cheapest at $0.10/$0.40
- Full Model Rankings — All 93 models ranked by price