- Input $/1M
- $2.00
- Output $/1M
- $10.00
- Context
- 1.05M
- Max output
- 128K
- Reasoning
- Yes
- Included in
- 5 plans
| Rate | Input | Output | Cache read | Cache write |
|---|---|---|---|---|
| Standard | Input$2.00 | Output$10.00 | Cache read$0.10 | Cache write$2.50 |
| Above 272K input tokens; higher rates apply to the full request | Input$4.00 | Output$15.00 | Cache read$0.20 | Cache write$5.00 |
- Reasoning tokens
- Billed at the output token rate
Published rate checked 2026-09-29. Sources below.
Processing tiers
The same model, processed on a different service tier. Rates for prompts up to the long-context threshold; speed claims are the provider’s and are not modelled.
| Tier | Input | Cache read | Output | Status |
|---|---|---|---|---|
| Standard | Input$2.00 | Cache read$0.10 | Output$10.00 | StatusAvailable |
| Batch | Input$1.00 | Cache read$0.05 | Output$5.00 | StatusAvailable |
| Flex | Input$1.00 | Cache read$0.05 | Output$5.00 | StatusAvailable |
| Fast | Input$4.00 | Cache read$0.20 | Output$20.00 | StatusAvailable |
| Ultrafast | InputNo price | Cache readNo price | OutputNo price | StatusComing soonUltrafast support for GPT-6.1 Sol is coming later. No Ultrafast price is published for it. |
Prices are for global processing. OpenAI adds a 10% premium for regional (data residency) processing where available and does not offer Fast mode with EU data residency; StackReplay does not price regional processing.
Current published base rates. Context tiers, cache-write duration and other conditions may change the rate for a request. These are not a reconstructed historical invoice.
- Context window
- 1,050,000
- Maximum output
- 128,000
- Input
- text · image
- Output
- text
- Knowledge cutoff
- Apr 30, 2026
- Reasoning
- Supported
- Tool calling
- Supported
- Structured output
- Supported
Endpoints: Responses, Chat Completions (without tool calling) and Batch.
A separate model from GPT-6 Sol with its own prices; cached input is billed at 5% of the input rate.
- Usage limits
- API request and token limits depend on your account usage tier. They are separate from ChatGPT and Codex subscription allowances.
- Thinking
- Effort: low, medium, high, xhigh, max. None and minimal are not supported.
- Cached input
- 5% of the input price, half of GPT-6 Sol's cached input price.
- Long prompts
- Above 272K input tokens: 2× input and cache prices, 1.5× output prices for the full request.
- Processing options
- Batch and Flex: 50% of Standard token prices. Fast mode: 2× Standard. Ultrafast is coming later, with no published price yet.
- Tool compatibility
- Use Responses for tool calling. Chat Completions is supported without tool calling.
- Regional processing
- 10% premium where available. US and EU data residency; Fast mode is unavailable with EU data residency.
Pricing, assumptions & evidence
Published base rate
Input $2.00 · Output $10.00 · Cache read $0.10 · Cache write $2.50
Above 272K input tokens; higher rates apply to the full request: input $4.00, output $15.00, cache read $0.20.
- GPT-6.1 Sol model specifications and API availability
- GPT-6.1 Sol launch rollout in ChatGPT Work and Codex (Plus, Pro, Business, Enterprise, Edu)
Missing token-category prices are not zero. Workload pricing applies exact recorded categories and admitted routes. No benchmark score or quality ranking is inferred from prices.
Aliases, routes and identity
Catalog ID gpt-6-1-sol
Model release (kind: release)
Developer: OpenAI

