Ministral 3 3B API Pricing: Mistral's Cheapest Model at $0.10/M

Ministral 3 3B offers symmetric pricing at $0.10/$0.10 — the same cost for input and output. Ideal for output-heavy tasks where other models' output premiums add up.

Updated Aug 7, 2026 · 93 models tracked across 11 providers

TL;DR

Symmetric Pricing: Most AI models charge 2–10x more for output than input. Ministral 3 3B charges the same rate ($0.10/M) for both. This makes it especially cost-effective for tasks that generate lots of output — like summarization, translation, or content generation.

Ministral 3 3B Pricing Breakdown

At $0.10 per million tokens for both input and output, Ministral 3 3B has the simplest pricing in the budget tier. Here's how the costs scale:

Monthly Volume Input Cost Output Cost Total Cost
1M tokens$0.10$0.10$0.20
10M tokens$1.00$1.00$2.00
100M tokens$10.00$10.00$20.00
1B tokens$100.00$100.00$200.00

Compare this to GPT-5 nano ($0.05/$0.40): at 100M tokens with a 50/50 input/output split, Ministral costs $20 while GPT-5 nano costs $22.50. But if your workload is 30% input / 70% output, Ministral drops to $16 vs nano's $29 — a 45% savings.

Why Symmetric Pricing Matters

Most AI providers charge significantly more for output tokens than input. This penalizes tasks that generate lots of text. Here's how Ministral's symmetric pricing compares for output-heavy workloads:

Model Input $/M Output $/M Output:Input Ratio Cost at 70% Output
Ministral 3 3B $0.10 $0.10 1:1 $10.00
GPT-5 nano $0.05 $0.40 8:1 $16.50
DeepSeek V4 Flash $0.14 $0.28 2:1 $12.60
Gemini 2.5 Flash-Lite $0.10 $0.40 4:1 $16.00
Qwen 3.7 Flash $0.03 $0.13 4.3:1 $5.20

At 100M tokens/month with 70% output, Ministral ($10) beats GPT-5 nano ($16.50) and Gemini 2.5 Flash-Lite ($16). Only Qwen 3.7 Flash ($5.20) and DeepSeek V4 Flash ($12.60) are competitive, but Qwen is from Alibaba and DeepSeek is from China — Ministral is the cheapest European/US option for output-heavy workloads.

Ministral 3 3B vs Other Budget Models

Model Input $/M Output $/M Context Size Provider
Qwen 3.7 Flash $0.03 $0.13 1M Alibaba
GPT-5 nano $0.05 $0.40 128K OpenAI
Ministral 3 3B $0.10 $0.10 128K 3B Mistral
Gemini 2.5 Flash-Lite $0.10 $0.40 1M Google
DeepSeek V4 Flash $0.14 $0.28 1M DeepSeek
Ministral 3 8B $0.15 $0.15 128K 8B Mistral

Among Mistral models, the Ministral family offers 3 tiers: 3B ($0.10/$0.10), 8B ($0.15/$0.15), and 14B ($0.20/$0.20). All have symmetric pricing. The 3B is cheapest; upgrade to 8B or 14B when you need better reasoning.

The Ministral Family: 3B vs 8B vs 14B

Model Input $/M Output $/M Context Best For
Ministral 3 3B $0.10 $0.10 128K Simple classification, extraction, edge
Ministral 3 8B $0.15 $0.15 128K Moderate reasoning, code, analysis
Ministral 3 14B $0.20 $0.20 128K Complex tasks, agentic workflows

All three Ministral models share symmetric pricing and 128K context. The difference is capability: 3B handles simple tasks, 8B adds moderate reasoning, and 14B approaches larger model quality. The price premium is small ($0.10 → $0.20), so consider starting with 14B and scaling down if quality is sufficient.

When to Use Ministral 3 3B

📊 Classification

Sentiment analysis, intent detection, spam filtering. The 3B size handles straightforward classification fast and cheap.

📝 Output-Heavy Tasks

Summarization, translation, content generation. Symmetric pricing means you don't pay a premium for generating text.

🔍 Data Extraction

Entity extraction, structured data from unstructured text. Fast inference at the lowest per-token cost.

🤖 Edge Deployment

3B parameters fit on consumer GPUs. Deploy locally for zero-latency, zero-cost inference after setup.

💬 Simple Q&A

FAQ bots, help desk automation, knowledge base queries. Handles straightforward questions efficiently.

🔄 High-Volume Pipelines

Process millions of items per day. Low per-token cost and fast inference make high-volume processing affordable.

Real-World Cost: Summarizing 50,000 Articles/Month

Suppose you're building a news summarization service that processes 50,000 articles per month, averaging 2,000 input tokens and generating 500 output tokens per article:

ModelInput CostOutput CostMonthly Total
Ministral 3 3B$10.00$2.50$12.50
GPT-5 nano$5.00$10.00$15.00
Gemini 2.5 Flash-Lite$10.00$10.00$20.00
DeepSeek V4 Flash$14.00$7.00$21.00

For this output-heavy workload (25% output ratio), Ministral 3 3B at $12.50/month is the cheapest option. Its symmetric pricing means you don't pay extra for generating summaries.

Compare 93 AI Models Side by Side

Ministral 3 3B is one of 93 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.

Frequently Asked Questions

How much does Ministral 3 3B cost?
Ministral 3 3B costs $0.10 per million input tokens and $0.10 per million output tokens. This symmetric pricing means input and output cost the same — unusual in the AI industry where output is typically 2-10x more expensive.
What is the context window of Ministral 3 3B?
Ministral 3 3B has a 128K token context window. This is standard for small models — sufficient for most chat, classification, and extraction tasks, though smaller than the 1M windows offered by larger models.
Is Ministral 3 3B cheaper than GPT-5 nano?
It depends on the task. Ministral 3 3B ($0.10/$0.10) is 2x more expensive on input than GPT-5 nano ($0.05/$0.40) but 4x cheaper on output. For output-heavy tasks (generation, summarization), Ministral is significantly cheaper. For input-heavy tasks (classification, extraction), GPT-5 nano wins.
What is Ministral 3 3B good for?
Ministral 3 3B excels at classification, extraction, simple Q&A, and edge deployment. Its 3B parameter size makes it fast and efficient for high-volume, low-complexity tasks. The symmetric pricing is ideal for output-heavy workloads where other models' output premiums add up.
Can I self-host Ministral 3 3B?
Yes. Ministral 3 3B is available through Mistral's API (api.mistral.ai) and can also be self-hosted via Hugging Face. Self-hosting eliminates per-token costs but requires GPU infrastructure.

Related Pages