- Same-Day Launches: Anthropic released Claude Opus 5.5 as OpenAI introduced GPT-6 Sol and Luna on September 22, bringing lower API prices to three new models.
- API Prices: Standard input and output rates per million tokens are $4 and $20 for Opus 5.5, $2 and $10 for Sol, and $0.10 and $0.50 for Luna.
- Coding Results: Opus 5.5 and GPT-6 Astra both reach 59.6% in the Artificial Analysis Terminal-Bench 4.0 chart, at the highest-scoring configurations shown.
- Different Trade-Offs: Sol and Luna cost less to run than their predecessors in independent testing, but improvements vary across coding and knowledge-work tasks.
- Beyond Benchmarks: OpenAI adds caching controls, while Anthropic changes API behavior and routes some sensitive requests to older Claude models.
Anthropic and OpenAI released new AI models on September 22, making lower API prices a shared feature of two otherwise different launches. Anthropic introduced Claude Opus 5.5, while OpenAI expanded its GPT-6 family with GPT-6 Sol and GPT-6 Luna.
Opus 5.5 targets demanding coding and knowledge work at a lower price than Opus 5. Sol and Luna bring OpenAI’s newer model family to less expensive tiers, with GPT-6 Astra remaining an existing comparison point rather than another launch that day. Independent results show why the choice depends on the workload: lower token prices, higher benchmark scores and lower spending per task do not always move together.
Claude Opus 5.5 Updates Coding and Knowledge Work
Anthropic describes improvements in long-running coding jobs, debugging, refactoring, code review and coordinating subagents. It also emphasizes clearer explanations and stronger work with documents, spreadsheets and tasks spanning several applications. The company says Opus 5.5 approaches the more expensive Claude Fable 5.1 on most work, although that remains a vendor assessment rather than a guarantee for every application.
The release is also a migration consideration for developers. Anthropic’s documentation says reasoning cannot be disabled and forced tool-use settings return errors. Applications that require a particular tool call need to account for those changes instead of treating the upgrade as a model-name swap.
GPT-6 Sol and Luna Extend OpenAI’s Lower-Cost Options
OpenAI positions GPT-6 Sol for complex coding and agent workflows, and GPT-6 Luna for focused, high-volume tasks. Both support adjustable reasoning effort, text and image inputs, and a 1,050,000-token context window. They are distinct from GPT-5.6 Sol and Luna, whose results remain predecessor comparisons.
OpenAI also reports fewer factual mistakes by Sol on an internal test built from conversations where users had flagged earlier errors. That is a targeted stress test, not a measurement of the error rate in ordinary use. Its launch report makes the same distinction for challenging alignment evaluations.
API Pricing: Opus 5.5, GPT-6 Sol and Luna
The Anthropic launch pricing and OpenAI’s Sol and Luna documentation put the three models at substantially different price points.
Standard API prices in US dollars per million tokens. Opus 5 is included as a predecessor reference.
| Billed Usage | Opus 5 | Opus 5.5 | GPT-6 Sol | GPT-6 Luna |
|---|---|---|---|---|
| Input | $5.00 | $4.00 | $2.00 | $0.10 |
| Output | $25.00 | $20.00 | $10.00 | $0.50 |
| Cached Input Reads | $0.50 | $0.20 | $0.20 | $0.01 |
Opus 5.5’s input and output rates are 20% below Opus 5’s, while cached input reads cost 60% less. Anthropic’s larger claim of 40% savings combines prices with token use on typical workloads at default settings. It should not be read as a guaranteed reduction for every job or reasoning setting.
Against OpenAI’s listed GPT-5.6 promotional rates, Sol’s input and output prices are halved. Luna’s input rate falls from $0.20 to $0.10, while output drops from $1.20 to $0.50. The table excludes cache writes, tool charges and alternative processing modes. OpenAI applies higher rates to requests above 272,000 input tokens.
Opus 5.5 is available through Anthropic and cloud providers including Amazon Web Services, Google Cloud and Microsoft Azure. OpenAI’s launch covers ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, alongside API access. Free and Go users can access Luna in the desktop app. OpenAI explicitly said the new models were not yet available in Chat at launch.
Coding Benchmarks: Opus 5.5 Versus GPT-6 Astra, Sol and Luna
The charts below reproduce selected models and effort settings, not live rankings. Company-reported results and independent evaluations are separated in the detailed table. Benchmark version, scoring rules, tools and fallback models matter; matching effort labels do not establish equal computational budgets.
In the Artificial Analysis Terminal-Bench 4.0 comparison, Opus 5.5 reaches 59.6% at max or xhigh effort with fallback enabled. Astra also reaches 59.6% at xhigh. Sol records 43.9% at max and Luna 12.6% at max in the displayed chart. Equal rounded scores do not establish identical performance on other tasks.
These are not the same runs as Anthropic’s launch comparison, which lists 66.4% for Opus 5.5, 52.3% for Opus 5 and 57.9% for Astra. That table also reports 54.4% for Opus 5.5 on FrontierCode v1.1 Main, against 48.0% for Opus 5 and 53.3% for Astra. Mixing those numbers with the independent chart would change the apparent gaps.
OpenAI’s separate DeepSWE v1.1 results put Sol at 68.8% and Luna at 66.6%, both at max effort. Astra’s published comparison lists 74.1% for Astra and 73.7% for Opus 5. No Opus 5.5 result is supplied for that row, so it cannot support a complete ranking of the new releases.
Business Work and Reasoning: Performance Versus Cost
The AutomationBench-AA chart shows 69.5% for Opus 5.5 at max with fallback, 68.5% for Astra at max, 61.7% for Sol at xhigh and 53.2% for Luna at max. Sol’s max-effort result is slightly lower, at 61.6%, illustrating why the setting needs to accompany the score.
This headline metric measures the average share of task objectives completed without guardrail violations. Zapier’s hosted AutomationBench leaderboard instead reports tasks completed fully. The provider-reported 40.0% for Opus 5.5 and 41.4% for Astra belong to that separate comparison, not the chart above. Fallback handling also differs, so scoring definitions and deployment settings both need to remain visible.
GDPval-AA Compares Professional Deliverables
On GDPval-AA v2.1, Opus 5.5 at max reaches 1,846 Elo, compared with 1,735 for Fable 5.1 and 1,708 for Opus 5. The launch comparison lists Astra at 1,542. These are ratings derived from comparisons of work products, not percentages of tasks completed or productivity gains.
The chart places Sol and Luna at lower-cost, lower-rated positions than Opus 5.5’s max-effort configuration. It also shows that changing effort can substantially change the same model’s cost and rating. Exact Sol and Luna ratings are not estimated from the plotted positions for the table.
Artificial Analysis’s launch-day report found GDPval-AA regressions for both new OpenAI models at max effort. Its Coding Agent Index was mixed: Sol scored 57 versus 55 for GPT-5.6 Sol, while Luna scored 41 versus 43 for GPT-5.6 Luna. A newer model name does not imply improvement on every evaluation.
Humanity’s Last Exam Adds a Different Reasoning Test
Artificial Analysis reports 61.4% for Opus 5.5 at max effort with default fallback on Humanity’s Last Exam. That independent result is separate from Anthropic’s 67.7% tools-enabled launch score. The two should not be presented as interchangeable measurements.
The chart’s vertical axis measures score, while its horizontal axis shows average evaluation cost per task on a logarithmic scale. Higher scores and lower costs are preferable; the chart’s “Lower is better” wording concerns cost, not the exam score. Average task cost is also not the same as spending per successful result after retries and human review.
Astra retains a higher company-reported score on Terminal-Bench-Science 0.1: 64.6% versus Opus 5.5’s 58.7%, although the latter is substantially above Opus 5’s 29.0%. Reported standard errors of several percentage points warrant caution about treating that gap as a definitive research advantage.
Full Benchmark Comparison
The first group preserves Anthropic’s launch compilation. The second draws from OpenAI’s Astra and Sol/Luna reports. The final group uses the independent chart snapshots and the separately identified Artificial Analysis Coding Agent Index report. A dash means no score is supplied here for that source and configuration, not zero.
Company-reported and independent benchmark results, grouped by source. Higher is better within each row; configurations and scoring methods differ.
| Benchmark / evaluator | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|
| Anthropic launch comparison – September 22 report | |||||||
| Agentic codingTerminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% | – | – | 37.3% |
| Agentic codingFrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | – | – | 47.5% |
| Agentic codingCursorBench 4.0 | 57.8% | 51.8% | 46.6% | – | – | – | 41.7% |
| Knowledge workGDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | – | – | 1588 |
| Business workflowsAutomationBench | 40.0% | 31.4% | 26.9% | 41.4% | – | – | 28.8% |
| Multidisciplinary reasoningHumanity’s Last Exam | 67.7% with tools | 65.6% with tools | 63.6% with tools | 57.2% with tools | – | – | – |
| Agentic scientific researchTerminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | – | – | 22.4% |
| Computer useOSWorld 2.0 | 81.8%partial | 80.7% partial | 74.0% partial | – | – | – | – |
| Visual chart recognitionChartography | 89.0%with tools | 88.4% with tools | 83.4% with tools | – | – | – | – |
| OpenAI reports – Sol and Luna and Astra | |||||||
| Software engineeringDeepSWE v1.1 – OpenAI reports | – | 67.4% | 73.7% | 74.1% | 68.8% max | 66.6% max | 72.7% |
| Professional workflowsAgents’ Last Exam V1 – OpenAI reports | – | – | 55.5% | 59.3% | 56.4% max | – | 53.6% |
| Computer useOSWorld 2.0 – offline set, partial score | – | – | 70.2% | 72.6% | 60.5% xhigh | – | 65.7% |
| Business workflowsAutomationBench 1.0.6 – OpenAI launch settings | – | 31.4% max; Opus 5 fallback | 26.9% max | 30.3% low | 33.2% xhigh | – | – |
| Artificial Analysis – chart snapshots and separately identified September 22 report | |||||||
| Agentic codingTerminal-Bench 4.0 – Artificial Analysis | 59.6% max or xhigh; fallback | 55.1% xhigh; fallback | – | 59.6% xhigh | 43.9% max | 12.6% max | 40% max; rounded report |
| Business workflowsAutomationBench-AA – objectives-based score | 69.5% max; fallback | 59.4% max; fallback | – | 68.5% max | 61.7% xhigh | 53.2% max | – |
| Coding indexAA Coding Agent Index – September 22 report | – | – | – | – | 57 max | 41 max | 55 max |
Caching, Safeguards and the Cost of Real Workflows
OpenAI Adds Controls for Reusing Context
Alongside the models, OpenAI introduced improved prompt caching, a monitoring dashboard, diagnostics for cache misses and explicit breakpoints for reusable prompt prefixes. Developers can also change reasoning effort through the documented configuration-update mechanism while preserving earlier cached context.
These changes matter for agents that repeatedly carry forward instructions, tool definitions and documents. Reusing that material can reduce processing and expense, but realized savings depend on how much context remains eligible for caching. A lower advertised token rate is only one part of the workflow’s bill.
Anthropic’s Safeguards Can Change Which Model Answers
Anthropic routes most cybersecurity tasks to Opus 4.8 and biology requests flagged by safeguards to Opus 5, while ordinary bug-fixing remains available on Opus 5.5. The company planned to extend its Cyber Verification Program to Opus 5.5 in the following weeks; vetted life-science organizations could already apply for more permissive access.
In its behavioral testing, Anthropic reports about 85% fewer attempts to cross containment boundaries than Opus 5 or Mythos 5.1, with the remaining attempts low severity and self-reported. This is a company-reported result on a particular evaluation, not an overall safety rating or a safety comparison with Sol and Luna.
Lower Token Prices Do Not Guarantee Lower Task Costs
Artificial Analysis found that Opus 5.5 at max effort used about 1.6 times Opus 5’s output tokens and had roughly similar cost per task across its Intelligence Index. That does not contradict savings at Anthropic’s default settings; it shows that the configuration changes the economics.
For OpenAI’s models, the evaluator reported average Intelligence Index task costs of $1.06 for Sol at max, versus $1.99 for its predecessor, and $0.07 for Luna at max, versus $0.18. Those savings accompanied mixed capability results rather than uniform gains.
Application-specific testing can differ again. In CodeRabbit’s review pipeline, the Standard Opus 5.5 configuration found 51 of 80 known bugs versus 49 for the production baseline, but used 49.2% more tokens. It also missed some bugs the baseline caught. Token counts alone did not establish the dollar-cost difference.
The September 22 releases therefore offer different combinations of capability and cost, not a universal replacement order. The practical comparison is the quality of the completed work, its total expense and the behavior of the deployed system – including reasoning effort, tools, caching and safeguards – on the tasks an organization actually needs to run.


