OpenAI released flagship model GPT-5.6 in three versions — Luna, Terra and Sol
OpenAI released the GPT-5.6 family of three models — Luna, Terra and Sol — with prices from $1 to $30 per 1M tokens. On the Agents' Last Exam benchmark, the senior Sol outperformed Claude Fable 5 by 13.1 points, but lost to it on SWE-Bench Pro — 64.6% versus 80%. OpenAI claimed that almost a third of SWE-Bench Pro assignments are incorrect.
AI-processed from Simon Willison; edited by Hamidun News
OpenAI on July 17, 2026 released its flagship GPT-5.6 model family in general availability in three sizes — Luna, Terra, and Sol — and according to OpenAI's own benchmark, the senior model Sol outperformed Anthropic's Claude Fable 5 by 13.1 points in the Agents' Last Exam test for long agent tasks.
How much do the new models cost
Prices are set per "1 million tokens" separately for input and output, and the three sizes of GPT-5.6 differ noticeably in cost.
- Luna (junior model) — $1 per 1M input tokens / $6 per 1M output tokens
- Terra (middle model) — $2.50 per 1M input tokens / $15 per 1M output tokens
- Sol (senior model) — $5 per 1M input tokens / $30 per 1M output tokens
- For comparison: the Claude Opus lineup costs $5/$25 per 1M tokens, and Claude Fable 5 — $10/$50
- All three GPT-5.6 models have a knowledge cutoff of February 16, 2026, a context window of 1 million tokens, and an output limit of 128,000 tokens
How much does GPT-5.6 outperform competitors
OpenAI relies on the Agents' Last Exam benchmark — a test of long agent workflows across 55 professional domains. According to the company, GPT-5.6 Sol scored 53.6 points on this test — a new maximum, beating Claude Fable 5 (in adaptive reasoning mode) by 13.1 points. Even in medium reasoning mode, Sol leads Fable 5 by 11.4 points at roughly four times lower cost. The junior models, Terra and Luna, according to OpenAI, beat Fable 5 at roughly 16 times lower expense.
However, on another popular benchmark — SWE-Bench Pro, a test of solving real engineering problems — Claude Fable 5 scored 80%, while GPT-5.6 Sol scored only 64.6%. Possibly this is why, on the eve of launch, OpenAI published a separate article claiming that about 30% of the tasks in SWE-bench Pro are incorrect (broken), and urged other model developers to more carefully check results on this benchmark.
What the first tests say
An author of a blog who had early access to GPT-5.6 Sol describes the model as "definitely very competent," but notes that on complex code-writing tasks he works with, the model didn't seem to him better than Claude Fable 5 from Anthropic. According to him, the most interesting details are contained in OpenAI's official guide to using GPT-5.6 — it describes a number of new API capabilities that third-party tool developers have yet to explore.
What this means
The launch of GPT-5.6 in three versions continues the race between OpenAI and Anthropic for leadership in agent tasks — those where a model independently executes long chains of actions rather than simply answering a single request. At the same time, comparison across different benchmarks gives a mixed picture: OpenAI wins on its own agent workflow test, but loses on SWE-Bench Pro — even though OpenAI itself publicly questioned the quality of a third of that benchmark's tasks. For developers, this is a signal: choice of model should be verified on your own tasks, not relied upon a single averaged rating.
Frequently asked questions
How much does GPT-5.6 Sol cost?
The senior model GPT-5.6 Sol costs $5 per 1 million input tokens and $30 per 1 million output tokens — more expensive than Luna ($1/$6) and Terra ($2.50/$15), but cheaper than Claude Fable 5 ($10/$50).
How does GPT-5.6 differ from Claude Fable 5 in capabilities?
On the Agents' Last Exam benchmark, GPT-5.6 Sol leads Fable 5 by 13.1 points (53.6 versus Fable 5's score), but on SWE-Bench Pro Fable 5 scores 80% against 64.6% for Sol — meaning superiority depends on task type.
What are the context and limits for GPT-5.6 models?
All three models — Luna, Terra, and Sol — have a context window of 1 million tokens, an output limit of 128,000 tokens, and a knowledge cutoff of February 16, 2026.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.