Anthropic Introduces Haiku 5.5, Its Fastest and Most Affordable Claude Model
Preface
Anthropic’s Claude Haiku 5.5 is a small model designed for tasks where speed, scale, and cost matter. Released on Wednesday, it is positioned for high-volume work such as summarizing documents and querying databases, as well as time-sensitive applications including live customer support and browser operation. Its headline change is price: for prompts up to 100,000 tokens, developers pay $0.10 per million input tokens and $0.50 per million output tokens, substantially less than Haiku 4.5. At the same time, reported benchmark results suggest that Haiku 5.5 can handle some computer-use and command-line tasks effectively. This article reviews its pricing, intended uses, reported evaluations, limitations, and availability. The central consideration is not just whether the model is fast or inexpensive, but whether it is reliable enough for the work it is assigned.
Lazy bag
Haiku 5.5 offers lower token prices than Haiku 4.5 and is intended for repetitive, high-volume, or speed-sensitive tasks. It scored 72.4% on OSWorld 2.1 and 39.2% on Terminal-Bench 4.0, exceeding the cited results for OpenAI’s GPT-6 Luna on both tests. Its GDPval-AA v2.1 score was 1620. However, a hands-on logic question produced a fast but incorrect answer, underscoring that benchmark strength does not guarantee correctness in every interaction. Users should verify important outputs rather than treating quick responses as proof of accuracy. Anthropic also introduced an adjustable effort setting for Haiku and reduced Sonnet 5.5’s cache-read price.
Main Body
Anthropic released Claude Haiku 5.5 on Wednesday and described it as the cheapest, fastest, and most capable small model it has made. The model is aimed at workloads that need to process many requests economically, as well as applications where a short response time is important. Examples include summarizing documents, querying databases, handling customer-support conversations, and operating a web browser on a user’s behalf. These use cases share a practical emphasis: completing defined tasks efficiently, often at high volume, rather than producing unusually creative or open-ended work.
For prompts up to 100,000 tokens, developers using the model through an API pay $0.10 per million input tokens and $0.50 per million output tokens. Tokens are the small pieces of text that a model processes and generates; they are also the units providers use to calculate many API charges. The stated rates match those OpenAI set for its competing small model, GPT-6 Luna, when it launched on September 22. Haiku 4.5, by comparison, cost $1 per million input tokens and $5 per million output tokens. The new rates are therefore 90% lower for those prompt sizes.
Prompts above 100,000 tokens also receive a 50% price reduction. Anthropic estimates the average saving at about 75%, taking into account that about 90% of requests to the older model fell below the 100,000-token threshold and that Haiku 5.5 divides text into slightly more tokens. The average saving is consequently an estimate based on the provider’s description of request patterns, rather than a claim that every individual request will cost 75% less. Actual charges depend on prompt length, generated output, and the applicable rate tier.
The model’s target workloads help explain the emphasis on cost and speed. Customer-support chats and long-email summaries are examples of repetitive, relatively straightforward tasks for which a lower price per request can matter when usage is substantial. Database queries and browser operation also benefit when a model can act quickly within a defined workflow. These characteristics make Haiku 5.5 potentially useful as one component in a larger product or service. They do not, by themselves, establish that it is suitable for every task or that it can operate without supervision.
Benchmark results provide a more specific picture of its capabilities. On OSWorld 2.1, which tests whether an AI can operate a real computer through long, multi-step tasks, Haiku 5.5 scored a 72.4% success rate. The source compares this result with 48.9% for OpenAI’s GPT Luna. OSWorld reports partial-credit percentages, so the score reflects performance across evaluated tasks rather than a guarantee that the model will complete every computer task successfully. The result nevertheless suggests that Haiku 5.5 can perform meaningfully on some workflows involving interfaces and sequences of actions.
Terminal-Bench 4.0 assesses AI agents on professional tasks they must complete by entering commands on their own. It measures the share of tasks completed correctly on the first attempt. Haiku 5.5 scored 39.2%, compared with 16.4% for OpenAI’s Luna and 0% for Haiku 4.5. Anthropic’s Claude Sonnet 5.5 scored 70.6% on the same coding test. The comparison places Haiku 5.5 ahead of the cited small-model results while also showing a substantial gap between its score and Sonnet 5.5’s. Scores from a particular benchmark indicate performance under that test’s conditions and should not be read as universal measures of coding ability.
On GDPval-AA v2.1, an evaluation of real professional work across 44 occupations, Haiku 5.5 scored 1620. The assessment uses an Elo rating system, a head-to-head rating approach borrowed from chess. The reported figures were 1437 for Luna and 735 for Haiku 4.5. This offers another point of comparison across models, although an aggregate rating cannot capture every occupation, organization, or real-world requirement. Taken together, the benchmark results portray Haiku 5.5 as a small model with notable performance on selected computer-use, command-line, and professional-work evaluations, not as a model that is uniformly superior in every setting.
A hands-on example illustrates why that distinction matters. The model was given a simple logic question and responded almost instantly, but its answer was wrong. This observation is limited to one interaction and is not a substitute for a systematic evaluation. Still, it is a practical reminder that speed and benchmark scores do not eliminate the possibility of errors. For consequential tasks, users should check outputs, test the model on representative examples, and establish appropriate human review before relying on its answers or actions.
Haiku 5.5 is the first Haiku model to include an adjustable effort setting. The setting allows users to trade cost for smarter answers by choosing how much effort the model applies. This gives developers another way to manage performance and spending for different task types. A simple, repetitive request may not need the same level of effort as a complex one. As with any adjustable setting, the useful balance depends on the application: teams need to evaluate whether added effort improves results enough to justify its cost and response-time implications.
Anthropic also cut the cache-read price for Sonnet 5.5 in half, to $0.10 per million tokens. Cache reads apply to text that a model has already processed, making the discounted rate relevant when applications reuse context. This pricing change is separate from Haiku 5.5’s token rates and may matter to developers who use Sonnet 5.5 in workflows with repeated text. Together, the announcements address API costs across more than one model in the Claude family.
The release completes a sequence of Claude 5.5 model launches. Haiku 5.5 arrived 15 days after Opus 5.5 on September 22 and nine days after Sonnet 5.5 on September 28. It was the last of the three Claude 5.5 models Anthropic had promised. Opus 5.5 was the first release since CEO Dario Amodei published an essay urging the industry to slow gains in AI capabilities. The timing situates Haiku 5.5 within a broader product rollout, while the model’s positioning focuses particularly on affordability and rapid execution.
Haiku 5.5 is available on the Claude website, Amazon Web Services, Google Cloud, and Microsoft Azure under the model name claude-haiku-5-5. Anthropic is also rolling out monthly API credits this week for some subscribers: $100 for users on the Max 5x plan, $200 for Max 20x, and up to $500, pooled across users, for Team plans. These credits are a separate offering from the model’s listed token prices. For developers and organizations considering adoption, the decision will depend on API costs, the fit between the model and intended tasks, and the safeguards needed to manage incorrect results.
Key Insights Table
| Aspect | Description |
|---|---|
| API pricing | For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, 90% below Haiku 4.5’s $1 and $5 rates. |
| Estimated savings | Anthropic estimates average savings near 75%, noting that about 90% of requests to the old model were under 100,000 tokens and Haiku 5.5 uses slightly more tokens. |
| OSWorld 2.1 | Haiku 5.5 scored 72.4%, compared with 48.9% for OpenAI’s GPT Luna. |
| Terminal-Bench 4.0 | Haiku 5.5 scored 39.2%, versus 16.4% for OpenAI’s Luna, 0% for Haiku 4.5, and 70.6% for Sonnet 5.5. |
| GDPval-AA v2.1 | Haiku 5.5 scored 1620, compared with Luna at 1437 and Haiku 4.5 at 735, across an evaluation covering 44 occupations. |
| Capabilities and caution | The model adds an adjustable effort setting and performs well on selected benchmarks, but a reported logic test produced a fast, incorrect answer. |
| Availability and credits | Available through the Claude website, Amazon Web Services, Google Cloud, and Microsoft Azure as claude-haiku-5-5. Monthly credits include $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team plans. |
Last edited at:2026/10/8
