OpenAI Models Landscape β Complete Reference & Cost Estimator
A full breakdown of every model in OpenAI’s current API portfolio β organized by tier, use case, and cost β with an interactive calculator to estimate your production spend before you build.
Four Layers. One Portfolio.
OpenAI folded standalone reasoning models into the flagship line in 2026. Understanding which layer β and which reasoning effort setting β fits your workload now matters more than knowing every model name.
Frontier General
GPT-5.6 Sol β highest capability, adjustable reasoning effort (none β max), 1M+ token context, premium pricing.
Value Production
GPT-5.6 Terra, GPT-5.6 Luna β high-volume defaults, fast, cost-efficient, same 1.05M context window as Sol.
Reasoning Effort (built-in)
o3 and o4-mini were retired Feb 2026. Set reasoning_effort on any GPT-5.6 model instead of switching models.
Specialist Media
GPT Image 2, Sora 2, gpt-realtime-2.1, gpt-4o-transcribe β modality-native models for image, video, voice, transcription.
Compatibility
GPT-5.5, GPT-5.4 family β still fully priced and supported, useful for cost comparison, not recommended for new builds.
Cadence Has Stayed Aggressive
OpenAI has shipped a new flagship family roughly every one to two months since GPT-5’s August 2025 launch.
What Changed Since Cohort 10 (May 2026)
Reasoning Effort Replaces Standalone Reasoning Models
o3 and o4-mini are retired. Every GPT-5.6 model now accepts a reasoning_effort parameter (none, low, medium, high, xhigh, max) β one model line spans everyday chat and deep, multi-step reasoning.
1.05M Token Context Is Now the Default
GPT-5.6 Sol, Terra, and Luna all ship with a 1.05M token context window out of the box β no separate “long context” model needed.
Four Pricing Tiers Per Model
Standard, Batch, Flex, and the newly renamed Fast mode (formerly Priority processing, renamed 30 Jul 2026) each carry different multipliers β batch/flex run roughly half of standard, fast mode roughly double.
Regional Data Residency Carries a 10% Uplift
Models released on or after 5 Mar 2026 that support regional processing endpoints are billed a 10% premium over standard pricing for data residency.
Modality-Native Models Still Win Their Domains
Voice (gpt-realtime-2.1), image (GPT Image 2), video (Sora 2), and transcription (gpt-4o-transcribe, gpt-transcribe) remain dedicated specialist models β don’t route audio or image work through general GPT endpoints.
The Models Worth Knowing
Not the full 70+ catalog β the models that actually show up in production builds.
Frontier model for complex reasoning, coding, and agentic tool use. 1.05M context, reasoning effort noneβmax.
Balances intelligence and cost. Same 1.05M context and tool support as Sol, roughly 60% cheaper.
Cost-sensitive, high-throughput workloads β classification, tagging, lightweight chat at scale.
Still fully supported, priced comparably to GPT-5.6. No longer the default recommendation for new builds.
State-of-the-art image generation and editing. Text and image inputs, image output billed per token.
Text-to-video generation, billed per second of output. Pro adds 1024p and 1080p resolutions.
Flagship voice model for live, low-latency audio-in/audio-out agents with tool use.
High-accuracy speech-to-text for file and realtime transcription.
Best cost/quality embedding model for retrieval, RAG, and search. 1,536 dimensions.
Text & Reasoning β Standard Pricing
Prices per 1M tokens, Standard tier, short context. Long-context requests run roughly 2Γ these rates on GPT-5.6/5.5/5.4.
| Model | Input | Cached input | Output | Status |
|---|---|---|---|---|
| gpt-5.6-sol | $5.00 | $0.50 | $30.00 | Flagship β recommended |
| gpt-5.6-terra | $2.00 | $0.20 | $12.00 | Balanced β default |
| gpt-5.6-luna | $0.20 | $0.02 | $1.20 | High-volume |
| gpt-5.5 | $5.00 | $0.50 | $30.00 | Legacy β still supported |
| gpt-5.5-pro | $30.00 | β | $180.00 | Legacy pro tier |
| gpt-5.4 | $2.50 | $0.25 | $15.00 | Legacy |
| gpt-5.4-mini | $0.75 | $0.075 | $4.50 | Legacy |
| gpt-5.4-nano | $0.20 | $0.02 | $1.25 | Legacy |
| gpt-5.4-pro | $30.00 | β | $180.00 | Legacy pro tier |
Batch tier β 50% of Standard. Fast mode β 2Γ Standard. Regional data-residency endpoints add a 10% uplift on eligible models (released β₯ 5 Mar 2026).
| Model | Modality | Input | Cached input | Output |
|---|---|---|---|---|
| gpt-realtime-2.1 | Audio | $32.00 | $0.40 | $64.00 |
| β³ | Text | $4.00 | $0.40 | $24.00 |
| β³ | Image | $5.00 | $0.50 | β |
| gpt-realtime-2.1-mini | Audio | $10.00 | $0.30 | $20.00 |
| β³ | Text | $0.60 | $0.06 | $2.40 |
| gpt-realtime-translate | Audio | β | $0.034/min | |
| gpt-live-transcribe | Audio | β | $0.017/min | |
| Model | Modality | Input | Cached input | Output |
|---|---|---|---|---|
| gpt-image-2 | Image | $8.00 | $2.00 | $30.00 |
| β³ | Text | $5.00 | $1.25 | β |
| gpt-image-1.5 | Image | $8.00 | $2.00 | $32.00 |
| gpt-image-1-mini | Image | $2.50 | $0.25 | $8.00 |
| Model | Size | Price / second |
|---|---|---|
| sora-2 | 720p | $0.10 |
| sora-2-pro | 720p | $0.30 |
| β³ | 1024p | $0.50 |
| β³ | 1080p | $0.70 |
| Model | Estimated cost |
|---|---|
| gpt-transcribe | $0.0045 / minute |
| gpt-live-transcribe | $0.017 / minute |
| gpt-4o-transcribe | $0.006 / minute ($2.50 in / $10.00 out per 1M tokens) |
| gpt-4o-mini-transcribe | $0.003 / minute ($1.25 in / $5.00 out per 1M tokens) |
| Model | Price / 1M tokens | Dimensions | Best for |
|---|---|---|---|
| text-embedding-3-small | $0.02 | 1,536 | Retrieval, RAG, search β best cost/quality |
| text-embedding-3-large | $0.13 | 3,072 | Max precision use cases |
| Tool | Details | Pricing |
|---|---|---|
| Web search | All models | $10.00 / 1k calls + content tokens at model rates |
| β³ | Non-reasoning preview | $25.00 / 1k calls, content tokens free |
| Containers | Hosted Shell & Code Interpreter | $0.03β$1.92 per 20-min session (1GBβ64GB) |
| File search | Storage | $0.10/GB/day (1GB free) |
| β³ | Tool call | $2.50 / 1k calls |
Estimate Your Production Spend
Pick a model and enter your expected monthly token volume to estimate cost before you build.
Estimate only β excludes cached-input discounts, tool-call fees, and long-context multipliers. Always confirm against OpenAI’s live pricing page before committing budget.
What Should I Use For…
The fastest way to pick a model: start from the task, not the model name.
How To Choose in Five Questions
Is this a build/prototype or a production workload?
Prototyping β default to GPT-5.6 Terra and iterate. Production at scale β model your monthly token volume with the calculator above before committing.
Does the task need deep, multi-step reasoning?
If yes, use GPT-5.6 Sol with reasoning_effort set to high or xhigh. If the task is straightforward (chat, drafting, simple Q&A), leave effort at none or low β you’re paying for reasoning tokens either way.
Is the workload high-volume and cost-sensitive?
Classification, tagging, and simple summarization at scale β GPT-5.6 Luna. For non-urgent batch jobs, stack the Batch tier discount (β50% off) on top.
Is the input non-text (voice, image, video)?
Route to the modality-native specialist model β gpt-realtime-2.1 for voice, GPT Image 2 for images, Sora 2 for video, gpt-4o-transcribe for transcription β never through a general chat model.
Does latency matter more than cost right now?
Use Fast mode (formerly Priority processing, renamed 30 Jul 2026) β roughly 2Γ Standard pricing for faster response times. Otherwise, Standard tier is the right default.
Decoding Data Science
Prepared for Cohort 11 Β· AI Residency Β· Pricing as of 1 August 2026 Β· Source: OpenAI official docs & pricing page