Cohort 11 Β· Live Reference Updated 1 Aug 2026 DDS AI Residency

OpenAI Models Landscape β€” Complete Reference & Cost Estimator

A full breakdown of every model in OpenAI’s current API portfolio β€” organized by tier, use case, and cost β€” with an interactive calculator to estimate your production spend before you build.

πŸ“…Updated: 1 Aug 2026 🧠Models covered: 70+ πŸ’‘Default recommendation: GPT-5.6 Terra πŸ”—Source: OpenAI official docs + pricing
🎯Cohort 11 Default Stack GPT-5.6 Terra β†’ everyday chat/coding  Β·  GPT-5.6 Sol β†’ high-stakes agentic/document work  Β·  GPT-5.6 Luna β†’ high-volume, cost-sensitive tasks  Β·  text-embedding-3-small β†’ retrieval/search  Β·  GPT Image 2 β†’ image generation  Β·  gpt-realtime-2.1 β†’ live voice agents  Β·  gpt-4o-transcribe β†’ speech-to-text
Portfolio Architecture

Four Layers. One Portfolio.

OpenAI folded standalone reasoning models into the flagship line in 2026. Understanding which layer β€” and which reasoning effort setting β€” fits your workload now matters more than knowing every model name.

Layer 1

Frontier General

GPT-5.6 Sol β€” highest capability, adjustable reasoning effort (none β†’ max), 1M+ token context, premium pricing.

Layer 2

Value Production

GPT-5.6 Terra, GPT-5.6 Luna β€” high-volume defaults, fast, cost-efficient, same 1.05M context window as Sol.

Layer 3

Reasoning Effort (built-in)

o3 and o4-mini were retired Feb 2026. Set reasoning_effort on any GPT-5.6 model instead of switching models.

Layer 4

Specialist Media

GPT Image 2, Sora 2, gpt-realtime-2.1, gpt-4o-transcribe β€” modality-native models for image, video, voice, transcription.

Layer 5

Compatibility

GPT-5.5, GPT-5.4 family β€” still fully priced and supported, useful for cost comparison, not recommended for new builds.

Release Timeline

Cadence Has Stayed Aggressive

OpenAI has shipped a new flagship family roughly every one to two months since GPT-5’s August 2025 launch.

Aug 2025
GPT-5
Nov 2025
GPT-5.1
Dec 2025
GPT-5.2 & GPT-5.2-Codex
Feb 2026
GPT-5.3-Codex
Feb 2026
o-series (o3, o4-mini) retired β€” folded into reasoning_effort
Mar 2026
GPT-5.4
Apr 2026
GPT-5.5 & GPT Image 2
Feb–Jul 2026
GPT-5.6 family (Sol, Terra, Luna)
30 Jul 2026
“Priority processing” renamed Fast mode
Key Strategic Themes

What Changed Since Cohort 10 (May 2026)

01

Reasoning Effort Replaces Standalone Reasoning Models

o3 and o4-mini are retired. Every GPT-5.6 model now accepts a reasoning_effort parameter (none, low, medium, high, xhigh, max) β€” one model line spans everyday chat and deep, multi-step reasoning.

02

1.05M Token Context Is Now the Default

GPT-5.6 Sol, Terra, and Luna all ship with a 1.05M token context window out of the box β€” no separate “long context” model needed.

03

Four Pricing Tiers Per Model

Standard, Batch, Flex, and the newly renamed Fast mode (formerly Priority processing, renamed 30 Jul 2026) each carry different multipliers β€” batch/flex run roughly half of standard, fast mode roughly double.

04

Regional Data Residency Carries a 10% Uplift

Models released on or after 5 Mar 2026 that support regional processing endpoints are billed a 10% premium over standard pricing for data residency.

05

Modality-Native Models Still Win Their Domains

Voice (gpt-realtime-2.1), image (GPT Image 2), video (Sora 2), and transcription (gpt-4o-transcribe, gpt-transcribe) remain dedicated specialist models β€” don’t route audio or image work through general GPT endpoints.

Model Cards

The Models Worth Knowing

Not the full 70+ catalog β€” the models that actually show up in production builds.

gpt-5.6-solRecommended

Frontier model for complex reasoning, coding, and agentic tool use. 1.05M context, reasoning effort none→max.

In $5.00/1MOut $30.00/1M
gpt-5.6-terraDefault

Balances intelligence and cost. Same 1.05M context and tool support as Sol, roughly 60% cheaper.

In $2.00/1MOut $12.00/1M
gpt-5.6-lunaHigh-volume

Cost-sensitive, high-throughput workloads β€” classification, tagging, lightweight chat at scale.

In $0.20/1MOut $1.20/1M
gpt-5.5 / gpt-5.4Legacy

Still fully supported, priced comparably to GPT-5.6. No longer the default recommendation for new builds.

In $2.50–5.00/1MOut $15–30/1M
gpt-image-2Recommended

State-of-the-art image generation and editing. Text and image inputs, image output billed per token.

Image in $8.00/1MOut $30.00/1M
sora-2 / sora-2-proVideo

Text-to-video generation, billed per second of output. Pro adds 1024p and 1080p resolutions.

720p $0.10–0.30/sec
gpt-realtime-2.1Recommended

Flagship voice model for live, low-latency audio-in/audio-out agents with tool use.

Audio in $32.00/1MOut $64.00/1M
gpt-4o-transcribeRecommended

High-accuracy speech-to-text for file and realtime transcription.

β‰ˆ $0.006/min
text-embedding-3-smallRecommended

Best cost/quality embedding model for retrieval, RAG, and search. 1,536 dimensions.

In $0.02/1M
Flagship Models

Text & Reasoning β€” Standard Pricing

Prices per 1M tokens, Standard tier, short context. Long-context requests run roughly 2Γ— these rates on GPT-5.6/5.5/5.4.

ModelInputCached inputOutputStatus
gpt-5.6-sol$5.00$0.50$30.00Flagship β€” recommended
gpt-5.6-terra$2.00$0.20$12.00Balanced β€” default
gpt-5.6-luna$0.20$0.02$1.20High-volume
gpt-5.5$5.00$0.50$30.00Legacy β€” still supported
gpt-5.5-pro$30.00β€”$180.00Legacy pro tier
gpt-5.4$2.50$0.25$15.00Legacy
gpt-5.4-mini$0.75$0.075$4.50Legacy
gpt-5.4-nano$0.20$0.02$1.25Legacy
gpt-5.4-pro$30.00β€”$180.00Legacy pro tier

Batch tier β‰ˆ 50% of Standard. Fast mode β‰ˆ 2Γ— Standard. Regional data-residency endpoints add a 10% uplift on eligible models (released β‰₯ 5 Mar 2026).

Changed from Cohort 10: GPT-5.5 was the flagship recommendation in May. As of August 2026, GPT-5.6 (Sol/Terra/Luna) is the current flagship family.
Realtime & Voice β€” per 1M tokens (or per minute where noted)
ModelModalityInputCached inputOutput
gpt-realtime-2.1Audio$32.00$0.40$64.00
β€³Text$4.00$0.40$24.00
β€³Image$5.00$0.50β€”
gpt-realtime-2.1-miniAudio$10.00$0.30$20.00
β€³Text$0.60$0.06$2.40
gpt-realtime-translateAudioβ€”$0.034/min
gpt-live-transcribeAudioβ€”$0.017/min
Image Generation β€” per 1M tokens
ModelModalityInputCached inputOutput
gpt-image-2Image$8.00$2.00$30.00
β€³Text$5.00$1.25β€”
gpt-image-1.5Image$8.00$2.00$32.00
gpt-image-1-miniImage$2.50$0.25$8.00
Video Generation (Sora) β€” per second
ModelSizePrice / second
sora-2720p$0.10
sora-2-pro720p$0.30
β€³1024p$0.50
β€³1080p$0.70
Transcription
ModelEstimated cost
gpt-transcribe$0.0045 / minute
gpt-live-transcribe$0.017 / minute
gpt-4o-transcribe$0.006 / minute ($2.50 in / $10.00 out per 1M tokens)
gpt-4o-mini-transcribe$0.003 / minute ($1.25 in / $5.00 out per 1M tokens)
Embeddings
ModelPrice / 1M tokensDimensionsBest for
text-embedding-3-small$0.021,536Retrieval, RAG, search β€” best cost/quality
text-embedding-3-large$0.133,072Max precision use cases
Built-in Tools
ToolDetailsPricing
Web searchAll models$10.00 / 1k calls + content tokens at model rates
β€³Non-reasoning preview$25.00 / 1k calls, content tokens free
ContainersHosted Shell & Code Interpreter$0.03–$1.92 per 20-min session (1GB–64GB)
File searchStorage$0.10/GB/day (1GB free)
β€³Tool call$2.50 / 1k calls
πŸ’°Cost Estimator

Estimate Your Production Spend

Pick a model and enter your expected monthly token volume to estimate cost before you build.

Input cost
$0.00
Output cost
$0.00
Est. monthly total
$0.00

Estimate only β€” excludes cached-input discounts, tool-call fees, and long-context multipliers. Always confirm against OpenAI’s live pricing page before committing budget.

Task Mapping

What Should I Use For…

The fastest way to pick a model: start from the task, not the model name.

Everyday chat, coding assistant, internal tools
Best balance of quality and cost for daily-driver use
GPT-5.6 Terra
Multi-step agentic workflows, long documents
Needs the strongest reasoning + 1M+ context
GPT-5.6 Sol
High-volume classification, tagging, summarization
Cost matters more than peak reasoning
GPT-5.6 Luna
Complex math, science, multi-step logic
Set reasoning effort to high or xhigh
GPT-5.6 Sol (effort: high)
RAG / semantic search / document retrieval
Best cost-to-quality embedding model
text-embedding-3-small
Image generation and editing
Current state-of-the-art image model
GPT Image 2
Short promo / social video clips
Billed per second, 720p is enough for most social use
Sora 2
Live voice agents (support, sales, tutoring)
Purpose-built for low-latency audio in/out
gpt-realtime-2.1
Transcribing recorded calls or workshops
High accuracy, billed per minute
gpt-4o-transcribe
Non-urgent, large overnight batch jobs
Batch tier runs β‰ˆ50% of Standard pricing
Any model, Batch tier
Selection Guide

How To Choose in Five Questions

1

Is this a build/prototype or a production workload?

Prototyping β†’ default to GPT-5.6 Terra and iterate. Production at scale β†’ model your monthly token volume with the calculator above before committing.

2

Does the task need deep, multi-step reasoning?

If yes, use GPT-5.6 Sol with reasoning_effort set to high or xhigh. If the task is straightforward (chat, drafting, simple Q&A), leave effort at none or low β€” you’re paying for reasoning tokens either way.

3

Is the workload high-volume and cost-sensitive?

Classification, tagging, and simple summarization at scale β†’ GPT-5.6 Luna. For non-urgent batch jobs, stack the Batch tier discount (β‰ˆ50% off) on top.

4

Is the input non-text (voice, image, video)?

Route to the modality-native specialist model β€” gpt-realtime-2.1 for voice, GPT Image 2 for images, Sora 2 for video, gpt-4o-transcribe for transcription β€” never through a general chat model.

5

Does latency matter more than cost right now?

Use Fast mode (formerly Priority processing, renamed 30 Jul 2026) β€” roughly 2Γ— Standard pricing for faster response times. Otherwise, Standard tier is the right default.

Decoding Data Science

Prepared for Cohort 11 Β· AI Residency Β· Pricing as of 1 August 2026 Β· Source: OpenAI official docs & pricing page

Leave a Reply

Your email address will not be published. Required fields are marked *