AI API rate limits, OpenAI rate limits, Anthropic RPM, Gemini rate limits, DeepSeek rate limits, API 429 errors, TPM limits, LLM rate limits 2026, Mistral rate limits, xAI rate limits">
โ† Blog

AI API Rate Limits Compared 2026: RPM & TPM for Every Provider

A comprehensive breakdown of requests per minute (RPM), tokens per minute (TPM), and tier-based limits across all 11 major AI API providers โ€” with strategies for handling 429 errors.

At 2 requests per user per minute, DeepSeek can only handle 30 concurrent users. That's fine for internal tools and small chatbots, but not for production apps with real traffic. Google Gemini Flash Lite handles 100x more users at 1/10th the cost.

If you're building anything beyond a hobby project, rate limits matter. A 429 error at peak traffic means lost revenue and frustrated users. This guide breaks down exactly what limits you'll hit โ€” and how to work around them.

The Bottom Line

For most developers, rate limits are not the bottleneck โ€” cost is. But if you're building a high-traffic chatbot or real-time API, rate limits become critical. Google Gemini Flash Lite offers the best combination of high RPM (6,000), high TPM (8M), and low cost ($0.075/$0.30). OpenAI scales well with spend. DeepSeek and Mistral are limited to 60 RPM โ€” fine for low-traffic apps, but you'll hit the ceiling fast.

Rate Limits by Provider: Quick Comparison

This table shows the maximum RPM and TPM available at each provider's highest paid tier. For tiered providers (OpenAI, Anthropic), limits increase as you spend more.

Provider Max RPM Max TPM Tiered? Cheapest Model
OpenAI 10,000โ€“20,000 2Mโ€“4M 5 tiers GPT-5 Mini ($0.30/$1.25)
Anthropic 4,000โ€“15,000 1Mโ€“2M 4 tiers Haiku 4.5 ($1/$5)
Google 2,000โ€“6,000 4Mโ€“8M Free + paid Flash-Lite ($0.075/$0.30)
DeepSeek 60 100Kโ€“200K Flat V4 Flash ($0.07/$0.27)
Mistral 60 200K Flat Ministral 3 3B ($0.10/$0.10)
Cohere 100 200K Flat Command R ($0.50/$1.50)
xAI 60 100K Flat Grok Build 0.1 ($0.30/$0.50)
Together.ai 60 200K Flat Llama 4 Scout ($0.08/$0.30)
Moonshot 60 100K Flat Kimi K2.6 ($0.30/$1.00)
AI21 60 200K Flat Jamba 1.5 Large ($0.50/$1.50)

๐Ÿ’ก Key Takeaway

Google Gemini models have the highest TPM limits (up to 8M tokens/min on Flash Lite) โ€” ideal for long-context or high-throughput workloads. OpenAI has the highest RPM (20K on budget models at Tier 5). Most smaller providers cap at 60 RPM regardless of spend.

OpenAI Rate Limits

OpenAI uses a 5-tier system based on cumulative API spend. You start at Tier 1 ($5+ spent) and progress automatically. Limits are per-model โ€” budget models (GPT-5 Mini, GPT-4o mini) get 2x the RPM of flagship models.

Tier Spend Required Flagship RPM Budget RPM Flagship TPM Budget TPM
Tier 1$55001,00040K80K
Tier 2$502,0004,000150K300K
Tier 3$1005,00010,000400K800K
Tier 4$2508,00015,000800K1.5M
Tier 5$1,00010,00020,0002M4M

Flagship models: GPT-5.5, GPT-5, GPT-5.4, GPT-5.6 Sol/Terra/Luna, GPT-oss 120B/20B. Budget models: GPT-5 Mini, GPT-4o mini. The Batch API has separate, higher limits and costs 50% less.

Anthropic (Claude) Rate Limits

Anthropic uses a 4-tier system based on spend history and account age. Limits vary significantly by model family โ€” Opus models get lower RPM than Haiku at every tier.

Tier Requirement Opus RPM Sonnet RPM Haiku RPM Max TPM
Freeโ€”2005001,00040Kโ€“100K
Tier 1$5 credit1,0002,0004,000200Kโ€“400K
Tier 2$40+ / 7 days2,0004,0008,000500Kโ€“1M
Tier 3$200+ / 14 days4,0008,00015,0001Mโ€“2M

Models covered: Opus 5, Opus 4.8, Fable 5 (flagship tier); Sonnet 5, Sonnet 4.6 (mid-tier); Haiku 4.5 (budget). Anthropic's Batch API also offers separate limits at 50% cost.

Google Gemini Rate Limits

Google uses a simpler free vs. paid model. Free tier limits are generous for experimentation. Paid tier limits are among the highest in the industry โ€” especially for TPM.

Model Free RPM Paid RPM Free TPM Paid TPM
Gemini 3.1 Pro52,00032K4M
Gemini 2.5 Pro52,00032K4M
Gemini 2.5 Flash154,0001M8M
Gemini 2.5 Flash-Lite306,0001M8M

๐Ÿ’ก Google's Secret Weapon: Flash-Lite

Gemini 2.5 Flash-Lite offers 6,000 RPM and 8M TPM on the paid tier at just $0.075/$0.30 per million tokens. That's the highest throughput-per-dollar of any provider โ€” perfect for high-volume classification, extraction, or routing tasks.

Flat-Rate Providers: DeepSeek, Mistral, xAI, and More

Most smaller providers offer flat rate limits โ€” no tier progression, no spend-based upgrades. What you get on day one is what you keep.

Provider RPM TPM Notes
DeepSeek 60 100Kโ€“200K Extremely cheap ($0.07โ€“$0.49/M). High quality. Low limits are the trade-off.
Mistral 60 200K European hosting option. 60 RPM is the hard ceiling.
xAI 60 100K Grok models. 1M context window is the standout feature, not RPM.
Cohere 100 200K Best flat-rate RPM among small providers. Good for RAG workloads.
Together.ai 60 200K Open-source models (Llama 4). Limits apply per model.
Moonshot 60 100K Kimi models. Competitive pricing, limited throughput.
AI21 60 200K Jamba models. 60 RPM is standard for this tier.

5 Strategies to Handle Rate Limits

1. Request Queuing

Add a queue layer (BullMQ, AWS SQS, Redis) to buffer requests when you hit RPM limits. This smooths out traffic spikes without dropping requests. The queue retries automatically when capacity opens up.

2. Multi-Key Rotation

Use multiple API keys and round-robin between them. This effectively multiplies your RPM limit. Most providers allow multiple keys per account โ€” check their terms of service first.

3. Smart Model Routing

Route simple requests to high-RPM budget models (Gemini Flash-Lite: 6K RPM, Haiku 4.5: 15K RPM) and complex requests to flagships. This reduces load on expensive, rate-limited endpoints while keeping quality high where it matters.

4. Batch APIs

For non-urgent workloads, use Batch APIs (available from OpenAI and Anthropic). They have separate, higher rate limits AND cost 50% less. Perfect for bulk processing, evaluation, and fine-tuning data prep.

5. Exponential Backoff

When you do hit a 429, don't retry immediately. Use exponential backoff: wait 1s, then 2s, then 4s, then 8s. Add jitter (random ยฑ30%) to prevent thundering herd problems when many clients retry simultaneously.

โš ๏ธ Don't Over-Engineer This

Most apps never hit rate limits. If you're processing fewer than 1,000 requests per minute, any provider will work. Focus on cost and quality first. Add rate-limit handling only when you actually see 429 errors in production.

How Many Concurrent Users Can Each Provider Handle?

Here's a practical translation: how many simultaneous users can each provider support, assuming each user makes one request every 30 seconds (typical for chat apps)?

Provider Model RPM (paid) Concurrent Users
GoogleFlash-Lite6,000~180,000
OpenAIGPT-5 Mini (T5)20,000~600,000
AnthropicHaiku 4.5 (T3)15,000~450,000
GoogleGemini 3.1 Pro2,000~60,000
CohereCommand R+100~3,000
DeepSeekV4 Pro60~1,800
MistralLarge 360~1,800
xAIGrok 4.360~1,800

Calculation: RPM รท 2 (one request every 30 seconds per user). Tier 5 OpenAI and Tier 3 Anthropic assumed for the high-end figures.

Frequently Asked Questions

What happens when I hit a rate limit?

You'll receive a 429 Too Many Requests HTTP response. The response header usually includes a Retry-After value (in seconds) telling you when to try again. Most provider SDKs handle this automatically with built-in retries.

Can I get higher rate limits?

Yes โ€” for tiered providers (OpenAI, Anthropic, Google), limits increase automatically as you spend more. For enterprise needs, all major providers offer custom rate limit agreements. Contact their sales teams for dedicated capacity.

Do rate limits apply per API key or per account?

It depends on the provider. OpenAI applies limits at the organization level (shared across all keys). Google applies limits at the project level. Anthropic applies limits per API key. Check your provider's documentation for specifics.

Which provider has the highest rate limits?

For RPM: OpenAI (20,000 RPM on budget models at Tier 5). For TPM: Google Gemini (8M TPM on Flash/Flash-Lite). For flat-rate simplicity: Cohere (100 RPM, no tiers to manage).

Are Batch API rate limits separate?

Yes. OpenAI and Anthropic both offer Batch APIs with separate, higher rate limits at 50% of the standard price. Batch requests are typically processed within 24 hours, making them ideal for non-urgent bulk workloads.

Check if your provider can handle your traffic. Enter your expected RPM and tokens per request to see which models work โ€” and what it costs.

Rate Limit Calculator โ†’ or Calculate Full Costs

Save money: ๐Ÿ“Š Live API Pricing ยท Cost Optimizer โ€” find out how much you could save by switching models. Free tool.