GPT-6 Astra API Pricing & Cost Calculator
GPT-6 Astra costs $10/M input, $1/M cached input, $12.50/M cache writes and $50/M output at Standard rates. Batch and Flex are half price. Requests above 272K total input tokens use higher rates for the entire request.
- Standard: $10 input · $1 cached input · $12.50 cache write · $50 output per 1M tokens.
- Batch and Flex: 50% of Standard rates.
- Above 272K input tokens: 2x input/cache rates and 1.5x output rates for the full request.
- Context window: 1.05M tokens; maximum output: 128K tokens.
GPT-6 Astra pricing table
| Processing | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| Standard | $10.00/M | $1.00/M | $12.50/M | $50.00/M |
| Batch | $5.00/M | $0.50/M | $6.25/M | $25.00/M |
| Flex | $5.00/M | $0.50/M | $6.25/M | $25.00/M |
Cache writes are priced at 1.25 times uncached input. Batch is for asynchronous jobs; Flex trades response speed and availability for lower cost. Eligibility and operational behavior still matter even when token rates match.
GPT-6 Astra cost calculator
Estimate one request
This is a token-cost estimate, not an invoice. Tool calls, storage, search, computer use, regional processing and other billable services are outside this calculator.
How the 272K threshold changes cost
The threshold is applied to the full request, not only the tokens beyond 272K. A request with 300K uncached input and 20K output therefore costs $6.00 for input plus $1.50 for output at Standard rates: $7.50. At 272K input exactly, the multiplier does not apply; OpenAI says it applies to prompts with more than 272K input tokens.
Practical examples
| Request | Standard | Batch/Flex |
|---|---|---|
| 100K input + 20K output | $2.00 | $1.00 |
| 20K uncached + 80K cached + 20K output | $1.28 | $0.64 |
| 300K input + 20K output (long context) | $7.50 | $3.75 |
When to choose Standard, Batch or Flex
- Standard: interactive production work where normal response behavior matters.
- Batch: asynchronous evaluations, document processing or backfills that can wait for batch completion.
- Flex: lower-priority workloads that can tolerate slower responses and occasional resource unavailability.
Frequently asked questions
What does cached input cost?
Cached input costs $1/M at Standard rates and $0.50/M with Batch or Flex. Cache-write tokens cost $12.50/M Standard or $6.25/M Batch/Flex.
Does output pricing include reasoning tokens?
Reasoning tokens are billed as output tokens when generated by the model. The calculator treats every billed output token at the selected output rate.
Is Fast mode included?
No. This calculator focuses on Standard, Batch and Flex. OpenAI documents Fast mode at twice applicable rates, but availability and regional restrictions require separate handling.