Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Anthropic: Claude Haiku 5.5

Fastest Claude model, built for high-volume, latency-sensitive work such as classification, extraction, routing, and subagent tasks.
Compare models
claude-haiku-5-5
AvailableAnthropicChatCached input 90% lessNew
Input / Output-70%
$0.10 / $0.50$0.03 / $0.15
Context
1M
Released
Oct 7, 2026
Max output
128K
Modalities
VisionChat
Capabilities
Tool usePrompt CacheReasoning

About Claude Haiku 5.5

Claude Haiku 5.5 is Anthropic's small, fast Claude model, released on October 7, 2026 for high-volume, latency-sensitive work such as classification, extraction, routing, and subagent tasks. It is the first Haiku with adjustable effort, so one model can answer simple requests quickly and think longer on harder ones. Anthropic calls it its fastest model at standard speed, with a knowledge cutoff of June 2026.

Where it works well

  • Anthropic calls it its fastest model at standard speed, which suits live customer support and browser use.
  • It is the first Haiku with effort levels, so one model can serve quick chat and longer agent tasks.
  • Anthropic reports large gains over Haiku 4.5 on benchmarks including OSWorld 2.1, GDPval-AA, and Terminal-Bench 4.0.
  • It accepts text and image input and works as a subagent beside Opus 5.5 or Sonnet 5.5 on coding work, per Anthropic.
  • Adaptive thinking is on by default, and the effort setting controls how much the model thinks.

When to choose another model

  • Anthropic says Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding and open-ended work.
  • Safety classifiers can decline a request with a refusal stop reason, and cybersecurity safeguards still block penetration testing.
  • Manual thinking budgets are not supported; adaptive thinking with the effort setting replaces them.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textMessages API
    POST/v1/messages
    API format:
    curl https://api.tokenlab.sh/v1/messages \
      -H "Content-Type: application/json" \
      -H "x-api-key: sk-xxx" \
      -H "anthropic-version: 2023-06-01" \
      -d '{
        "model": "claude-haiku-5-5",
        "max_tokens": 1024,
        "messages": [
          {"role": "user", "content": "Hello!"}
        ]
      }'

Pricing

TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.

Input tokens <= 100K

per 1M tokens
Official price
Input $0.10 / Output $0.50 / Cache read $0.01 / Cache write $0.125
TokenLab price
Input $0.03 / Output $0.15 / Cache read $0.003 / Cache write $0.0375
Discount
-70%

Input tokens 100K-1M

per 1M tokens
Official price
Input $0.50 / Output $2.50 / Cache read $0.05 / Cache write $0.625
TokenLab price
Input $0.15 / Output $0.75 / Cache read $0.015 / Cache write $0.1875
Discount
-70%
Prompt cache pricing

Cache read

Official price
$0.01
TokenLab price
$0.003
Discount
-70%

Cache write

Official price
$0.125
TokenLab price
$0.0375
Discount
-70%
Tool fees

Charged per call, only when the model uses the tool.

Web search

Official price
$0.01/search
TokenLab price
$0.003/search
Discount
-70%

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Claude Haiku 5.5 in Console with a prompt ready to edit or send.

Help me try claude-haiku-5-5 with a short message at /v1/messages. Show the reply, latency, and cost.

Use cases

Best for
  • Reasoning
  • Vision
  • High-volume classification and routing

    Label tickets, route requests, and pull fields out of documents at volume, where response speed and steady throughput matter most.

  • Subagents in coding workflows

    Run narrow side tasks such as searching a codebase or summarizing files while a larger model leads the work, a pairing Anthropic highlights.

  • Live support and browser tasks

    Answer customers in real time or drive browser steps from screenshots, where the wait between turns shapes the experience.

  • Summaries and compaction

    Condense long conversations, filings, or logs into short briefs, and compact an agent's context between steps.

Prompt examples

Classify each of these 50 support tickets as billing, bug, or feature request, and return a JSON list with the ticket ID and category.

Read this filing and pull the company name, fiscal year, revenue, and net income into a table. Mark any value you could not find.

Summarize this conversation so far in six bullets that keep every decision, open question, and file name, so a new agent can continue the work.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is Claude Haiku 5.5 best at?

Narrowly scoped, high-volume work: classification, extraction, routing, summaries, compaction, and subagent tasks. Anthropic markets it as its fastest model and says Sonnet 5.5 or Opus 5.5 remain better for complex agentic coding.

Which effort level should I start with on Haiku 5.5?

Anthropic recommends medium, the default, for most work including agentic coding. Use low for chat, short tool tasks, and simple high-volume requests, though in long agent prompts it may skip a search or stop early. Use high for strict instruction following, and xhigh or max only where your own tests show a gain.

Can I turn off thinking on Haiku 5.5?

Partly. Thinking is adaptive and on by default. Sending thinking disabled works at low, medium, and high effort, and returns an error at xhigh and max. Anthropic's preferred way to get less thinking is a lower effort level. Thinking counts toward the output limit, so leave room for it.

How does Haiku 5.5 differ from Haiku 4.5?

Anthropic reports large benchmark gains, a longer context window and output limit, adjustable effort, and adaptive thinking in place of manual thinking budgets. It uses the newer tokenizer, so the same text counts as more tokens, and its safety classifiers can return refusals that Haiku 4.5 did not.

When should I pick Sonnet 5.5 instead of Haiku 5.5?

Choose Sonnet 5.5 or Opus 5.5 for complex agentic coding, long terminal sessions, and tasks that need sustained judgment, where Anthropic says they stay stronger. Choose Haiku 5.5 when speed and volume matter and the task is narrowly scoped, such as classification, summaries, or subagent work.

How much does Claude Haiku 5.5 cost?

On TokenLab, Claude Haiku 5.5 costs Input $0.03 / Output $0.15 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of Claude Haiku 5.5?

Claude Haiku 5.5 accepts up to 1,000,000 tokens of context and returns up to 128,000 tokens in one response.

Which endpoint should Claude Haiku 5.5 use?

Use https://api.tokenlab.sh/v1/messages for Claude Haiku 5.5. The request example below shows the matching code shape.

Which operations does Claude Haiku 5.5 support?

Claude Haiku 5.5 supports Text to text. Select an operation above to see its endpoint and request example.

Compare Claude Haiku 5.5

Sources

Reviewed Oct 8, 2026

More from Claude 5

Related models