AI API Rate Limits Compared 2026: RPM & TPM for Every Provider
A comprehensive breakdown of requests per minute (RPM), tokens per minute (TPM), and tier-based limits across all 11 major AI API providers โ with strategies for handling 429 errors.
At 2 requests per user per minute, DeepSeek can only handle 30 concurrent users. That's fine for internal tools and small chatbots, but not for production apps with real traffic. Google Gemini Flash Lite handles 100x more users at 1/10th the cost.
If you're building anything beyond a hobby project, rate limits matter. A 429 error at peak traffic means lost revenue and frustrated users. This guide breaks down exactly what limits you'll hit โ and how to work around them.
The Bottom Line
For most developers, rate limits are not the bottleneck โ cost is. But if you're building a high-traffic chatbot or real-time API, rate limits become critical. Google Gemini Flash Lite offers the best combination of high RPM (6,000), high TPM (8M), and low cost ($0.075/$0.30). OpenAI scales well with spend. DeepSeek and Mistral are limited to 60 RPM โ fine for low-traffic apps, but you'll hit the ceiling fast.
Rate Limits by Provider: Quick Comparison
This table shows the maximum RPM and TPM available at each provider's highest paid tier. For tiered providers (OpenAI, Anthropic), limits increase as you spend more.
| Provider | Max RPM | Max TPM | Tiered? | Cheapest Model |
|---|---|---|---|---|
| OpenAI | 10,000โ20,000 | 2Mโ4M | 5 tiers | GPT-5 Mini ($0.30/$1.25) |
| Anthropic | 4,000โ15,000 | 1Mโ2M | 4 tiers | Haiku 4.5 ($1/$5) |
| 2,000โ6,000 | 4Mโ8M | Free + paid | Flash-Lite ($0.075/$0.30) | |
| DeepSeek | 60 | 100Kโ200K | Flat | V4 Flash ($0.07/$0.27) |
| Mistral | 60 | 200K | Flat | Ministral 3 3B ($0.10/$0.10) |
| Cohere | 100 | 200K | Flat | Command R ($0.50/$1.50) |
| xAI | 60 | 100K | Flat | Grok Build 0.1 ($0.30/$0.50) |
| Together.ai | 60 | 200K | Flat | Llama 4 Scout ($0.08/$0.30) |
| Moonshot | 60 | 100K | Flat | Kimi K2.6 ($0.30/$1.00) |
| AI21 | 60 | 200K | Flat | Jamba 1.5 Large ($0.50/$1.50) |
๐ก Key Takeaway
Google Gemini models have the highest TPM limits (up to 8M tokens/min on Flash Lite) โ ideal for long-context or high-throughput workloads. OpenAI has the highest RPM (20K on budget models at Tier 5). Most smaller providers cap at 60 RPM regardless of spend.
OpenAI Rate Limits
OpenAI uses a 5-tier system based on cumulative API spend. You start at Tier 1 ($5+ spent) and progress automatically. Limits are per-model โ budget models (GPT-5 Mini, GPT-4o mini) get 2x the RPM of flagship models.
| Tier | Spend Required | Flagship RPM | Budget RPM | Flagship TPM | Budget TPM |
|---|---|---|---|---|---|
| Tier 1 | $5 | 500 | 1,000 | 40K | 80K |
| Tier 2 | $50 | 2,000 | 4,000 | 150K | 300K |
| Tier 3 | $100 | 5,000 | 10,000 | 400K | 800K |
| Tier 4 | $250 | 8,000 | 15,000 | 800K | 1.5M |
| Tier 5 | $1,000 | 10,000 | 20,000 | 2M | 4M |
Flagship models: GPT-5.5, GPT-5, GPT-5.4, GPT-5.6 Sol/Terra/Luna, GPT-oss 120B/20B. Budget models: GPT-5 Mini, GPT-4o mini. The Batch API has separate, higher limits and costs 50% less.
Anthropic (Claude) Rate Limits
Anthropic uses a 4-tier system based on spend history and account age. Limits vary significantly by model family โ Opus models get lower RPM than Haiku at every tier.
| Tier | Requirement | Opus RPM | Sonnet RPM | Haiku RPM | Max TPM |
|---|---|---|---|---|---|
| Free | โ | 200 | 500 | 1,000 | 40Kโ100K |
| Tier 1 | $5 credit | 1,000 | 2,000 | 4,000 | 200Kโ400K |
| Tier 2 | $40+ / 7 days | 2,000 | 4,000 | 8,000 | 500Kโ1M |
| Tier 3 | $200+ / 14 days | 4,000 | 8,000 | 15,000 | 1Mโ2M |
Models covered: Opus 5, Opus 4.8, Fable 5 (flagship tier); Sonnet 5, Sonnet 4.6 (mid-tier); Haiku 4.5 (budget). Anthropic's Batch API also offers separate limits at 50% cost.
Google Gemini Rate Limits
Google uses a simpler free vs. paid model. Free tier limits are generous for experimentation. Paid tier limits are among the highest in the industry โ especially for TPM.
| Model | Free RPM | Paid RPM | Free TPM | Paid TPM |
|---|---|---|---|---|
| Gemini 3.1 Pro | 5 | 2,000 | 32K | 4M |
| Gemini 2.5 Pro | 5 | 2,000 | 32K | 4M |
| Gemini 2.5 Flash | 15 | 4,000 | 1M | 8M |
| Gemini 2.5 Flash-Lite | 30 | 6,000 | 1M | 8M |
๐ก Google's Secret Weapon: Flash-Lite
Gemini 2.5 Flash-Lite offers 6,000 RPM and 8M TPM on the paid tier at just $0.075/$0.30 per million tokens. That's the highest throughput-per-dollar of any provider โ perfect for high-volume classification, extraction, or routing tasks.
Flat-Rate Providers: DeepSeek, Mistral, xAI, and More
Most smaller providers offer flat rate limits โ no tier progression, no spend-based upgrades. What you get on day one is what you keep.
| Provider | RPM | TPM | Notes |
|---|---|---|---|
| DeepSeek | 60 | 100Kโ200K | Extremely cheap ($0.07โ$0.49/M). High quality. Low limits are the trade-off. |
| Mistral | 60 | 200K | European hosting option. 60 RPM is the hard ceiling. |
| xAI | 60 | 100K | Grok models. 1M context window is the standout feature, not RPM. |
| Cohere | 100 | 200K | Best flat-rate RPM among small providers. Good for RAG workloads. |
| Together.ai | 60 | 200K | Open-source models (Llama 4). Limits apply per model. |
| Moonshot | 60 | 100K | Kimi models. Competitive pricing, limited throughput. |
| AI21 | 60 | 200K | Jamba models. 60 RPM is standard for this tier. |
5 Strategies to Handle Rate Limits
1. Request Queuing
Add a queue layer (BullMQ, AWS SQS, Redis) to buffer requests when you hit RPM limits. This smooths out traffic spikes without dropping requests. The queue retries automatically when capacity opens up.
2. Multi-Key Rotation
Use multiple API keys and round-robin between them. This effectively multiplies your RPM limit. Most providers allow multiple keys per account โ check their terms of service first.
3. Smart Model Routing
Route simple requests to high-RPM budget models (Gemini Flash-Lite: 6K RPM, Haiku 4.5: 15K RPM) and complex requests to flagships. This reduces load on expensive, rate-limited endpoints while keeping quality high where it matters.
4. Batch APIs
For non-urgent workloads, use Batch APIs (available from OpenAI and Anthropic). They have separate, higher rate limits AND cost 50% less. Perfect for bulk processing, evaluation, and fine-tuning data prep.
5. Exponential Backoff
When you do hit a 429, don't retry immediately. Use exponential backoff: wait 1s, then 2s, then 4s, then 8s. Add jitter (random ยฑ30%) to prevent thundering herd problems when many clients retry simultaneously.
โ ๏ธ Don't Over-Engineer This
Most apps never hit rate limits. If you're processing fewer than 1,000 requests per minute, any provider will work. Focus on cost and quality first. Add rate-limit handling only when you actually see 429 errors in production.
How Many Concurrent Users Can Each Provider Handle?
Here's a practical translation: how many simultaneous users can each provider support, assuming each user makes one request every 30 seconds (typical for chat apps)?
| Provider | Model | RPM (paid) | Concurrent Users |
|---|---|---|---|
| Flash-Lite | 6,000 | ~180,000 | |
| OpenAI | GPT-5 Mini (T5) | 20,000 | ~600,000 |
| Anthropic | Haiku 4.5 (T3) | 15,000 | ~450,000 |
| Gemini 3.1 Pro | 2,000 | ~60,000 | |
| Cohere | Command R+ | 100 | ~3,000 |
| DeepSeek | V4 Pro | 60 | ~1,800 |
| Mistral | Large 3 | 60 | ~1,800 |
| xAI | Grok 4.3 | 60 | ~1,800 |
Calculation: RPM รท 2 (one request every 30 seconds per user). Tier 5 OpenAI and Tier 3 Anthropic assumed for the high-end figures.
Frequently Asked Questions
What happens when I hit a rate limit?
You'll receive a 429 Too Many Requests HTTP response. The response header usually includes a Retry-After value (in seconds) telling you when to try again. Most provider SDKs handle this automatically with built-in retries.
Can I get higher rate limits?
Yes โ for tiered providers (OpenAI, Anthropic, Google), limits increase automatically as you spend more. For enterprise needs, all major providers offer custom rate limit agreements. Contact their sales teams for dedicated capacity.
Do rate limits apply per API key or per account?
It depends on the provider. OpenAI applies limits at the organization level (shared across all keys). Google applies limits at the project level. Anthropic applies limits per API key. Check your provider's documentation for specifics.
Which provider has the highest rate limits?
For RPM: OpenAI (20,000 RPM on budget models at Tier 5). For TPM: Google Gemini (8M TPM on Flash/Flash-Lite). For flat-rate simplicity: Cohere (100 RPM, no tiers to manage).
Are Batch API rate limits separate?
Yes. OpenAI and Anthropic both offer Batch APIs with separate, higher rate limits at 50% of the standard price. Batch requests are typically processed within 24 hours, making them ideal for non-urgent bulk workloads.
Check if your provider can handle your traffic. Enter your expected RPM and tokens per request to see which models work โ and what it costs.
Rate Limit Calculator โ or Calculate Full Costs