- Cheaper AI: Anthropic’s small Claude Haiku 5.5 model cuts token prices up to 90%; tokens are billable text units.
- Worker Role: It handles summaries, information extraction and smaller assignments for larger AI agents, with stronger professional-work results than Haiku 4.5.
- Effort Control: Adjustable thinking changes capability and cost; independent tests favor Haiku overall but GPT-6 Luna on business-software workflows.
- Billing Boundary: Rates rise above 100,000 prompt tokens, including previously processed input reused from a cache.
- API Credits: Monthly credits for calling Claude from software are rolling out to subscribers of its paid Max and Team plans.
AI developer Anthropic has released Claude Haiku 5.5, a cheaper small model for developers building document summaries, information-processing tools and assistants that carry out tasks. The upgrade adds adjustable thinking effort, allowing applications to spend more reasoning on harder assignments.
The model is available through Anthropic’s Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure. Its context window, the material it can hold in one conversation, reaches one million tokens, compared with Haiku 4.5’s 200,000. The lower rates sit alongside capability gains, while the new reasoning controls change how many tokens a job consumes.
A Small Worker for Larger AI Jobs
Anthropic positions Haiku as a supporting model for its larger Sonnet 5.5 and Opus 5.5 models. In a financial-presentation example supplied by Rogo, a financial AI company, the larger model assembles the presentation while Haiku retrieves the revenue figure for one slide. Rogo’s Alex Wang, who works in applied AI, attributes the appeal to sufficient accuracy for that narrow assignment, combined with speed and low cost that permit repeated use.
Anthropic already had envisioned a parallel-worker role for Haiku 4.5, in which Sonnet 4.5 would break a larger job into subtasks for multiple Haiku instances. Haiku 5.5 also adds browser-tool support on the Claude API and Google Cloud, enabling work inside webpages.
The professional-work results in Anthropic’s model system cardshow a substantial change from the predecessor. On GDPval-AA v2.1, independently run by evaluation firm Artificial Analysis, Haiku 5.5 scored 1,620 at maximum effort against Haiku 4.5’s 735. The evaluation asks models to produce professional work such as documents, slides and spreadsheets across 220 tasks in 44 occupations. Its Elo ratings rank the comparative quality of outputs judged against one another. Haiku’s default medium-effort score was 1,277.
For narrower financial-document work, Anthropic reports 60.3% at maximum effort on OfficeQA Pro, up from Haiku 4.5’s 47.1%. That harder 133-question test supplies preselected historical US Treasury documents as extracted text.
Anthropic’s computer-use evaluation also shows gains: Haiku 5.5 earned 72.4% partial credit on OSWorld 2.1’s 82-task offline subset, versus 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna. Haiku 5.5 operated an Ubuntu computer through screenshots, mouse and keyboard actions at maximum effort, without internet access. Haiku 4.5 used a fixed thinking budget. Haiku 5.5 fully passed 37.1% of tasks; the higher partial-credit score rewards completed checkpoints within unfinished tasks.
Larger models retain more headroom on demanding terminal work. Anthropic reports Haiku at 39.2% on Terminal-Bench 4.0, a 66-task command-line evaluation, against Sonnet 5.5’s 70.6% and Opus 5.5’s 66.4%. Haiku ran at maximum effort without internet access or a fallback model when safeguards blocked a request; the larger-model results include fallback effects, and Opus used xhigh, the effort level below maximum.
Cognition’s FrontierCode Main evaluation put Haiku and Sonnet almost level at maximum effort: 46.4% and 46.2%. It assesses patches for difficult open-source software changes. Sonnet’s best result was 52.1% at xhigh.
Effort Changes Both Results and Cost
Haiku 5.5 is the first Haiku model with adjustable effort. Adaptive thinking is enabled by default, with medium effort as the starting setting. Lower effort reduces reasoning and can skip it on simple requests; thinking consumes output tokens as well as time.
OpenAI’s GPT-6 Luna is a current alternative built for focused, high-volume work. Artificial Analysis’s matched medium-effort comparison gives Haiku the higher Intelligence Index score, 34 against 30. The index combines ten evaluations, so different task results can pull in different directions. Luna scored 40% against Haiku’s 29% on AutomationBench-AA, which evaluates agents carrying out workflows in online business software.
The effort comparison also exposes a cost difference despite matching low-tier token rates:
Artificial Analysis: Matched Effort Settings
| Effort | Haiku 5.5 Index | GPT-6 Luna Index | Haiku Cost per Index Task | Luna Cost per Index Task |
|---|---|---|---|---|
| Medium | 34 | 30 | $0.05 | $0.02 |
| Maximum | 43 | 38 | $0.21 | $0.07 |
Task cost is AA’s weighted benchmark calculation from measured token use and listed prices. Its published method leaves the new Haiku long-prompt surcharge unspecified.
Haiku used more output in those benchmark workloads. At medium effort, AA recorded a weighted 33,000 output tokens per index task against Luna’s 11,000; at maximum, the figures were 162,000 and 50,000. These are totals across a task’s model calls. Higher effort improved Haiku’s aggregate result while substantially expanding its generated reasoning and answers.
Faster generation can coexist with a longer wait for an answer. In AA’s maximum-effort measurements, Haiku generated about 243 tokens per second against Luna’s 128, but took about 415 seconds to its first answer token against 109. That timing includes thinking; AA’s response-time method uses reasoning averaged over 60 diverse prompts. Haiku’s medium profile reached its first answer token in about 13 seconds.
Customer application results describe a different experience. In a statement supplied by Anthropic, Asana engineer Aaron Vinh reports task-completion latency falling more than 30% against an unnamed incumbent model. Its AI Teammates tests covered bug triage, project setup and searches for overdue or high-risk work.
Google also offers stable alternatives with different roles: Gemini 3.5 Flash-Lite targets high-volume translation and simple data processing, while Gemini 3.8 Flash targets longer software-engineering and agent workflows. Both are stable models available through the Gemini API.
Which Prompts Get the Biggest Price Cut
Haiku 5.5’s first-party dollar rates are 90% below Haiku 4.5 through 100,000 prompt tokens, including exactly 100,000. Above that length, the listed reduction is 50%, with rates five times the shorter-prompt tier. The prompt’s length determines both input and output rates for the request.
Prompt caching saves processed input for reuse: cache writes store it, and cache reads reuse it on later requests. The table’s write rates apply to a five-minute cache lifetime.
US Dollars per Million Tokens
| Charge | Haiku 5.5: Up to 100,000 Prompt Tokens | Haiku 5.5: Over 100,000 Prompt Tokens | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output, Including Thinking | $0.50 | $2.50 | $5.00 |
| Cache Reads | $0.01 | $0.05 | $0.10 |
| Five-Minute Cache Writes | $0.125 | $0.625 | $1.25 |
First-party base rates with global routing. Cache writes shown use a five-minute lifetime; one-hour writes have separate rates. Partner cloud schedules and regional premiums can differ.
Cached content remains part of the prompt even though reading it costs less. The input total includes newly processed content, cache writes and cache reads, including cached thinking. A large reused document can therefore push a request into the higher tier.
The updated tokenizer, which divides text into billable units, produces approximately 30% more input tokens for the same text than Haiku 4.5, with the increase varying by content. That changes both token budgets and the tier a prompt can enter. Anthropic’s estimated 75% average running-cost reduction accounts for changed token consumption and the prior request mix: about 90% of Haiku 4.5 requests fell into the shorter tier.
GPT Luna’s standard prices match all four of Haiku’s shorter-tier rates, including cache writes and reads. Its higher-price threshold starts above 272,000 input tokens. Google’s paid Gemini 3.5 Flash-Lite rates are $0.30 input and $2.50 output per million tokens. Gemini 3.8 Flash’s promotional rates, $0.75 and $3.75, run through December 31, 2026, before doubling on January 1.
Reusing Context and Moving Existing Applications
Haiku’s minimum cacheable prompt falls from 4,096 tokens to 512, allowing smaller repeated inputs to qualify for caching. It also preserves cached thinking across ordinary user turns, where Haiku 4.5 stripped earlier thinking and removed subsequent messages from the cache. Changing the effort setting can invalidate message caches, so switching reasoning depth can also change which input is reused.
Sonnet 5.5’s cache-read price drops from $0.20 to $0.10 per million tokens. The cut reduces that repeated-input component of a larger agent’s bill; Anthropic estimates roughly 20% savings on most agentic work, with the benefit depending on how much context is reused.
Applications using Haiku 4.5’s explicit thinking budgets need to move to adaptive effort, and computer-use integrations need a new toolset, as the migration guide explains. Haiku 4.5 capacity commitments under Priority Tier, Anthropic’s reserved-capacity service, do not carry over: Haiku 5.5 does not support that tier.
Safeguards Still Constrain Agent Work
Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s. They allow some defensive work but block penetration testing and other attacker-associated techniques. A blocked first-party request ends without being sent to a fallback model. Organizations seeking broader cybersecurity or biology access can apply through Anthropic’s verification programs.
Anthropic’s pre-deployment tests report improved resistance to prompt injection, where hostile instructions hidden in material an agent reads try to redirect its actions. On the Gray Swan benchmark, without additional prompt-injection protections, attack success fell from 83.2% for Haiku 4.5 to 7.1% after up to 15 attempts. The graphical computer-use subset still reached 24.4%, leaving a material weak spot for agents operating desktop interfaces.
Subscription Credits Fund First-Party API Work
Anthropic is also rolling out monthly API credits for its paid Max and Team plans. Max 5x subscribers receive $100 and Max 20x subscribers $200. Team credits pool $20 per Standard seat and $100 per Premium seat, capped at $500 for the whole team. For example, three Standard seats and two Premium seats produce a $260 monthly pool.
Eligible plans need to be active for seven days, and the offer is appearing over several days. Credits go to one linked Claude Console organization, the shared account from which its API users spend the balance. They cover any available model on the Claude Platform, including API and Agent SDK use, while interactive Claude Code and partner cloud platforms remain outside the credit’s scope.
Unused credits expire at the end of each billing cycle. Once the allowance runs out, requests draw on purchased credits or auto-reload; without another balance, API requests stop until new credits arrive. Organizations invoiced through Anthropic sales are billed for excess usage as usual.


