GPT-5.6 Model Tiers and API Pricing
Sol vs Terra vs LunaA source-based guide to choosing between GPT-5.6 Sol, Terra, and Luna, with current API token prices, reasoning options, caching rules, and practical workload boundaries.
This page organizes first-party documentation and editorial analysis. ModelRun Lab has not independently reproduced the published performance or hardware claims.
What changed with GPT-5.6
OpenAI released GPT-5.6 as a three-tier family rather than a single API model. Sol is the flagship tier, Terra is positioned as the balanced option for everyday work, and Luna is the fastest and least expensive tier.
This guide organizes OpenAI's published model and pricing information. ModelRun Lab has not independently benchmarked the three tiers, and the workload recommendations below are editorial starting points rather than guaranteed performance outcomes.
Current API prices
OpenAI lists the following prices per one million tokens:
| Tier | Input | Output | Positioning | | --- | ---: | ---: | --- | | GPT-5.6 Sol | $5.00 | $30.00 | Highest-capability tier | | GPT-5.6 Terra | $2.50 | $15.00 | Balanced general-purpose tier | | GPT-5.6 Luna | $1.00 | $6.00 | Fastest and lowest-cost tier |
OpenAI notes that it reduced Terra and Luna pricing on July 30, 2026. Pricing can change, so production cost models should be checked against the current official pricing record before deployment.
A practical selection rule
Choose Luna when throughput and unit cost matter more than maximum capability. Plausible starting workloads include classification, extraction, routing, short transformations, and high-volume subagent work. These are editorial examples, not an official guarantee that every task will meet a required quality bar.
Choose Terra as the default evaluation candidate for mixed production workloads. OpenAI positions it as the balanced tier for everyday work. It is a reasonable first model to test when an application needs stronger coding, research, or tool use than a cost-first model but does not justify Sol pricing on every request.
Choose Sol for the most difficult work where a higher success rate can offset higher token cost. OpenAI highlights complex coding, professional knowledge work, cybersecurity, science, design, and long-running agentic tasks. Teams should confirm the value with their own task-specific evaluation rather than relying only on provider benchmarks.
Reasoning and agent options
GPT-5.6 supports multiple reasoning-effort settings. OpenAI describes max as a setting above xhigh for more exploration and revision. The ultra setting coordinates multiple agents and therefore trades higher token use for additional capability and parallelism.
For API applications, ultra-like workflows can be built through the Responses API multi-agent beta. Programmatic Tool Calling allows a model to write and run lightweight programs that coordinate tools and filter intermediate results.
These features can materially change total cost. A model's listed token price is not the same as the final price of completing a workflow: reasoning effort, retries, tool calls, parallel agents, prompt length, and output length all matter.
Prompt caching
OpenAI states that GPT-5.6 supports explicit cache breakpoints and a minimum cache life of 30 minutes. Cache writes are billed at 1.25 times the uncached input rate, while cache reads receive a 90 percent cached-input discount.
Caching is most relevant when a stable system prompt, document set, or repeated context is reused often enough to offset the higher write rate. Applications with frequently changing prompts may see less benefit.
Cost comparison method
A simple first-pass estimate is:
request cost = input tokens ? input rate + output tokens ? output rate
Then add the effects of cache writes, cache reads, retries, and parallel agents. For model routing, compare the cost of a successful completed task rather than the cost of one call. A cheaper model that requires repeated retries may cost more overall.
Evidence boundaries
- The tier names, prices, availability, caching rules, and feature descriptions come from OpenAI.
- Workload examples and selection guidance are editorial interpretations.
- ModelRun Lab has not independently reproduced OpenAI's benchmark or cost-efficiency claims.
- API prices and availability should be rechecked before production use.