Luna replaces expensive per-token inference for high-volume, latency-sensitive workloads like classification and routing, not hard reasoning tasks. It's the only model in the family on ChatGPT's free tier, so most casual users meet it first, while teams needing multi-agent orchestration should upgrade to a pricier sibling.
Luna is OpenAI's efficiency-tier model, released July 9, 2026, with the lowest per-token price and lowest latency of the family, at the cost of a reasoning ceiling capped below Terra and Sol. It targets high-volume inference and consumer ChatGPT Free tier access rather than frontier reasoning work.
Provider: OpenAI · Family: GPT-5.6
Context window: 200,000 tokens · Max output: 64,000
Input modalities: text, image · Output: text, tool-calls, code
About GPT-5.6 Luna
GPT-5.6 Luna is the efficiency tier of OpenAI's GPT-5.6 family, launched July 9, 2026 for general availability. It delivers the strongest capability-to-cost ratio in the lineup, targeting high-volume inference workloads where per-token cost dominates total economics. Luna sacrifices peak reasoning depth and advanced agentic features for raw speed and affordability. Luna shares the GPT-5.6 architecture family: token efficiency improvements over the prior generation, explicit prompt caching with cache breakpoints, persisted reasoning across turns, and vision plus text input with text and tool-calls or code output. It supports reasoning effort levels none, low, medium, and high; xhigh, max, and pro mode stay reserved for Terra and Sol. There is no ultra multi-agent mode and no programmatic tool calling. The model supports a 200K context window with 64K max output tokens. Native modalities: text and image input; text, tool-calls, and code output. Vision capabilities are present but not benchmarked against dedicated multimodal flagships, and there's no native audio or video I/O. Luna is available on the OpenAI API, ChatGPT on all tiers including Free with rate limits, Codex, Microsoft Copilot with explicit selection, and other model gateways. Knowledge cutoff is June 2026. See the pricing FAQ for exact rates. Enterprise buyers get Zero Data Retention eligibility, US and EU data residency, plus the standard compliance certifications listed under governance. Latency runs around 800ms p50 and 3000ms p99, the fastest of any GPT-5.6 tier.
Pricing
Input $1/M, Output $6/M, Cached input $0.125/M (1.25x base, 90% discount). 20% of Sol price, 40% of Terra price. No batch discount at launch. ChatGPT Free tier includes rate-limited Luna access. Enterprise ZDR eligible. Data residency: US, EU.
Key Features
- Maximum Cost Efficiency: The cheapest tier in the family, priced for the best token-per-dollar economics on volume workloads.
- Fastest Latency in the Family: Runs faster than both Sol and Terra, making it the pick when p50 latency matters more than peak reasoning depth.
- Full Vision Plus Tool Use: Not stripped down: text and image input, function calling, code execution, and structured output are all present.
- Auto-Caching Plus Persisted Reasoning: Automatic prompt caching carries context forward, and reasoning.context reuses reasoning items across turns.
- ChatGPT Free Tier Access: The only model in the family available on the free ChatGPT tier, which gives it the broadest consumer reach.
Pros
- Unbeatable cost efficiency for high-volume inference, with auto-caching keeping per-request cost near zero at scale.
- The fastest latency in the family, which meets strict SLA requirements for real-time classification, routing, and extraction.
- A full capability surface (vision, tools, code), unlike typical lite models that strip features to hit a price point.
- ChatGPT Free tier inclusion means massive distribution and an upgrade funnel toward the pricier siblings.
Cons
- Reasoning ceiling caps out at high effort, so it can't match Terra or Sol on complex multi-step or novel problems.
- No ultra multi-agent mode, no programmatic tool calling, no pro mode, and no explicit caching control.
- As a distilled model it can show overconfidence on edge cases and subtler reasoning gaps than its teacher models.
- Auto-caching only, with a real write-cost penalty, so volatile prompt prefixes burn budget silently.
Benchmarks
- math: 71.5
- mmlu: 88.9
- mmlu pro: 74.8
- aime 2025: 72.1
- arc agi 2: 18.7
- humaneval: 88.4
- live bench: 52.4
- lmarena elo: 1340
- gpqa diamond: 55.3
- lmarena rank: 12
- aider polyglot: 62.1
- swe bench verified: 60.2
- humanitys last exam: 14.2
- artificial analysis intelligence index: 45
- artificial analysis price blended per m: 3.5
- artificial analysis speed tokens per sec: 90
Frequently Asked Questions
How much does GPT-5.6 Luna cost per 1M tokens?
Luna costs $1 per 1M input tokens and $6 per 1M output tokens, with cached input at $0.125 per 1M, a 90% discount over uncached. That's 20% of Sol's per-token price and 40% of Terra's, with no batch discount at launch.
How does GPT-5.6 Luna compare on benchmarks vs GPT-5.6 Sol and Terra?
Luna scores about 60% on SWE-bench Verified and 55% on GPQA Diamond, well behind Terra and further behind Sol on both. It trades that reasoning gap for the lowest price and fastest latency in the family, with no ultra multi-agent mode or programmatic tool calling.
Is GPT-5.6 Luna open source or proprietary?
Luna is proprietary. OpenAI makes it available through its own API, ChatGPT's free and paid tiers, Codex, Microsoft Copilot as an explicit selection, and select gateway partners. There are no downloadable weights and no open license.
Does GPT-5.6 Luna train on user data?
By default OpenAI doesn't train on API inputs or outputs, and Luna is Zero Data Retention eligible for enterprise customers. It's SOC2 Type II, ISO 27001, GDPR, and HIPAA eligible, with US and EU data residency options.
Who is GPT-5.6 Luna best for and who should avoid it?
Luna fits high-volume production inference like classification, extraction, and routing, plus consumer apps built on ChatGPT's free tier. Avoid it for complex multi-step reasoning, novel problem solving, or any workflow needing multi-agent orchestration; route those to Terra or Sol instead.