GPT-6 Astra API Pricing & Cost Calculator

GPT-6 Astra costs $10/M input, $1/M cached input, $12.50/M cache writes and $50/M output at Standard rates. Batch and Flex are half price. Requests above 272K total input tokens use higher rates for the entire request.

Published September 10, 2026 · Pricing checked against official OpenAI documentation

Quick answer

GPT-6 Astra pricing table

ProcessingInputCached inputCache writeOutput
Standard$10.00/M$1.00/M$12.50/M$50.00/M
Batch$5.00/M$0.50/M$6.25/M$25.00/M
Flex$5.00/M$0.50/M$6.25/M$25.00/M

Cache writes are priced at 1.25 times uncached input. Batch is for asynchronous jobs; Flex trades response speed and availability for lower cost. Eligibility and operational behavior still matter even when token rates match.

GPT-6 Astra cost calculator

Estimate one request

$2.00

This is a token-cost estimate, not an invoice. Tool calls, storage, search, computer use, regional processing and other billable services are outside this calculator.

How the 272K threshold changes cost

The threshold is applied to the full request, not only the tokens beyond 272K. A request with 300K uncached input and 20K output therefore costs $6.00 for input plus $1.50 for output at Standard rates: $7.50. At 272K input exactly, the multiplier does not apply; OpenAI says it applies to prompts with more than 272K input tokens.

Calculator assumption: total input is uncached input + cached input + cache-write tokens. If that sum exceeds 272,000, the calculator doubles all input/cache rates and multiplies output rates by 1.5.

Practical examples

RequestStandardBatch/Flex
100K input + 20K output$2.00$1.00
20K uncached + 80K cached + 20K output$1.28$0.64
300K input + 20K output (long context)$7.50$3.75

When to choose Standard, Batch or Flex

Frequently asked questions

What does cached input cost?

Cached input costs $1/M at Standard rates and $0.50/M with Batch or Flex. Cache-write tokens cost $12.50/M Standard or $6.25/M Batch/Flex.

Does output pricing include reasoning tokens?

Reasoning tokens are billed as output tokens when generated by the model. The calculator treats every billed output token at the selected output rate.

Is Fast mode included?

No. This calculator focuses on Standard, Batch and Flex. OpenAI documents Fast mode at twice applicable rates, but availability and regional restrictions require separate handling.

Official sources

Related APIpulse pages