Simon Willison→ original

OpenAI released flagship model GPT-5.6 in three versions — Luna, Terra and Sol

OpenAI released the GPT-5.6 family of three models — Luna, Terra and Sol — with prices from $1 to $30 per 1M tokens. On the Agents' Last Exam benchmark, the senior Sol outperformed Claude Fable 5 by 13.1 points, but lost to it on SWE-Bench Pro — 64.6% versus 80%. OpenAI claimed that almost a third of SWE-Bench Pro assignments are incorrect.

AI-processed from Simon Willison; edited by Hamidun News
OpenAI released flagship model GPT-5.6 in three versions — Luna, Terra and Sol
Source: Simon Willison. Collage: Hamidun News.
◐ Listen to article

OpenAI on July 17, 2026 released its flagship GPT-5.6 model family in general availability in three sizes — Luna, Terra, and Sol — and according to OpenAI's own benchmark, the senior model Sol outperformed Anthropic's Claude Fable 5 by 13.1 points in the Agents' Last Exam test for long agent tasks.

How much do the new models cost

Prices are set per "1 million tokens" separately for input and output, and the three sizes of GPT-5.6 differ noticeably in cost.

  • Luna (junior model) — $1 per 1M input tokens / $6 per 1M output tokens
  • Terra (middle model) — $2.50 per 1M input tokens / $15 per 1M output tokens
  • Sol (senior model) — $5 per 1M input tokens / $30 per 1M output tokens
  • For comparison: the Claude Opus lineup costs $5/$25 per 1M tokens, and Claude Fable 5 — $10/$50
  • All three GPT-5.6 models have a knowledge cutoff of February 16, 2026, a context window of 1 million tokens, and an output limit of 128,000 tokens

How much does GPT-5.6 outperform competitors

OpenAI relies on the Agents' Last Exam benchmark — a test of long agent workflows across 55 professional domains. According to the company, GPT-5.6 Sol scored 53.6 points on this test — a new maximum, beating Claude Fable 5 (in adaptive reasoning mode) by 13.1 points. Even in medium reasoning mode, Sol leads Fable 5 by 11.4 points at roughly four times lower cost. The junior models, Terra and Luna, according to OpenAI, beat Fable 5 at roughly 16 times lower expense.

However, on another popular benchmark — SWE-Bench Pro, a test of solving real engineering problems — Claude Fable 5 scored 80%, while GPT-5.6 Sol scored only 64.6%. Possibly this is why, on the eve of launch, OpenAI published a separate article claiming that about 30% of the tasks in SWE-bench Pro are incorrect (broken), and urged other model developers to more carefully check results on this benchmark.

What the first tests say

An author of a blog who had early access to GPT-5.6 Sol describes the model as "definitely very competent," but notes that on complex code-writing tasks he works with, the model didn't seem to him better than Claude Fable 5 from Anthropic. According to him, the most interesting details are contained in OpenAI's official guide to using GPT-5.6 — it describes a number of new API capabilities that third-party tool developers have yet to explore.

What this means

The launch of GPT-5.6 in three versions continues the race between OpenAI and Anthropic for leadership in agent tasks — those where a model independently executes long chains of actions rather than simply answering a single request. At the same time, comparison across different benchmarks gives a mixed picture: OpenAI wins on its own agent workflow test, but loses on SWE-Bench Pro — even though OpenAI itself publicly questioned the quality of a third of that benchmark's tasks. For developers, this is a signal: choice of model should be verified on your own tasks, not relied upon a single averaged rating.

Frequently asked questions

How much does GPT-5.6 Sol cost?

The senior model GPT-5.6 Sol costs $5 per 1 million input tokens and $30 per 1 million output tokens — more expensive than Luna ($1/$6) and Terra ($2.50/$15), but cheaper than Claude Fable 5 ($10/$50).

How does GPT-5.6 differ from Claude Fable 5 in capabilities?

On the Agents' Last Exam benchmark, GPT-5.6 Sol leads Fable 5 by 13.1 points (53.6 versus Fable 5's score), but on SWE-Bench Pro Fable 5 scores 80% against 64.6% for Sol — meaning superiority depends on task type.

What are the context and limits for GPT-5.6 models?

All three models — Luna, Terra, and Sol — have a context window of 1 million tokens, an output limit of 128,000 tokens, and a knowledge cutoff of February 16, 2026.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…