Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Anthropic: Claude Haiku 4.5

Fast Claude lane for lightweight agents, office tasks, and responsive chat
Compare models
claude-haiku-4-5
AvailableAnthropicChatCached input 90% less
Input / Output-70%
$1.00 / $5.00$0.30 / $1.50
Context
200K
Max output
64K
Modalities
VisionChat
Capabilities
Tool usePrompt CacheReasoning

About Claude Haiku 4.5

Claude Haiku 4.5 is Anthropic's small, fast Claude model, released in October 2025 for high-volume and latency-sensitive work. Anthropic describes it as the fastest model in its lineup with near-frontier intelligence. It accepts text and images, supports tool use and prompt caching, and uses manual extended thinking with a token budget rather than the adaptive thinking of newer Claude models.

Where it works well

  • Lowest latency among the Claude models listed on Anthropic's overview page, which suits chat front ends and real-time assistants.
  • Reads images next to text, so it can classify screenshots, receipts, and charts at volume.
  • Extended thinking is optional and set with an explicit token budget, so you decide per request how much reasoning to pay for.
  • Supports tool use, structured JSON output, and prompt caching, which keeps repeated system prompts cheap in agent loops.

When to choose another model

  • Its knowledge cutoff is much earlier than the Claude 5 generation, so it knows less about recent libraries and events.
  • It has a smaller context window and output cap than the Opus and Sonnet 5.x models, which limits whole-repository or book-length jobs.
  • For long autonomous coding runs or ambiguous multi-step planning, a Sonnet or Opus model fails less often.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textMessages API
    POST/v1/messages
    API format:
    curl https://api.tokenlab.sh/v1/messages \
      -H "Content-Type: application/json" \
      -H "x-api-key: sk-xxx" \
      -H "anthropic-version: 2023-06-01" \
      -d '{
        "model": "claude-haiku-4-5",
        "max_tokens": 1024,
        "messages": [
          {"role": "user", "content": "Hello!"}
        ]
      }'

Pricing

TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.

Default spec

per 1M tokens
Official price
Input $1.00 / Output $5.00 / Cache read $0.10 / Cache write $1.25
TokenLab price
Input $0.30 / Output $1.50 / Cache read $0.03 / Cache write $0.375
Discount
-70%
Prompt cache pricing

Cache read

Official price
$0.10
TokenLab price
$0.03
Discount
-70%

Cache write

Official price
$1.25
TokenLab price
$0.375
Discount
-70%
Tool fees

Charged per call, only when the model uses the tool.

Web search

Official price
$0.01/search
TokenLab price
$0.003/search
Discount
-70%

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
30-day success rate
100.0%
7-day median latency
1.6 sn=364
7-day P95 latency
15.8 sn=364
30-day requests
100+
Last active
7 minutes ago

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Claude Haiku 4.5 in Console with a prompt ready to edit or send.

Help me try claude-haiku-4-5 with a short message at /v1/messages. Show the reply, latency, and cost.

Use cases

Best for
  • Reasoning
  • Vision
  • Real-time chat and support

    Put it behind a customer-facing widget where reply speed matters more than depth, with tool calls for order lookups and a handoff when the question gets hard.

  • Classification and extraction

    Tag tickets, pull fields from invoices or screenshots, and return strict JSON for a downstream pipeline that runs thousands of times a day.

  • Sub-agent workers

    Run it as the cheap worker in a multi-agent setup: a larger model plans, and several Haiku calls search, summarize, or check files in parallel.

Prompt examples

Classify this support email as billing, bug, feature request, or other, and return JSON with the label and a one-sentence reason.

Summarize this chat transcript in three bullet points and list any action item with its owner.

Look at this screenshot of an error dialog and tell me which setting the user should change.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is Claude Haiku 4.5 best used for?

Fast, high-volume jobs: chat assistants, classification, data extraction, summarization, and worker agents. Anthropic positions it as the fastest Claude model with near-frontier intelligence, so it is the default pick when response time and throughput matter more than the deepest reasoning.

Does Claude Haiku 4.5 support extended thinking?

Yes, but in the manual form: you enable thinking and set a token budget. It does not use the adaptive thinking mode or the effort parameter that the newer Sonnet and Opus models use, so you control reasoning depth through the budget you pass.

Can Claude Haiku 4.5 process images?

Yes. It takes images alongside text and answers in text. That covers reading screenshots, charts, and scanned documents. It cannot generate or edit images, so use an image model when you need picture output.

When should I move up from Haiku 4.5 to Sonnet?

Move up when Haiku misses instructions on long prompts, drops steps in multi-tool agent runs, or needs knowledge of recent releases. Sonnet 5.5 is the next step and keeps a fast profile, with a longer context window and a much later knowledge cutoff.

How much does Claude Haiku 4.5 cost?

On TokenLab, Claude Haiku 4.5 costs Input $0.30 / Output $1.50 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of Claude Haiku 4.5?

Claude Haiku 4.5 accepts up to 200,000 tokens of context and returns up to 64,000 tokens in one response.

Which endpoint should Claude Haiku 4.5 use?

Use https://api.tokenlab.sh/v1/messages for Claude Haiku 4.5. The request example below shows the matching code shape.

Which operations does Claude Haiku 4.5 support?

Claude Haiku 4.5 supports Text to text. Select an operation above to see its endpoint and request example.

Compare Claude Haiku 4.5

Guides that use Claude Haiku 4.5

Sources

Reviewed Oct 2, 2026

More from Claude 4

Related models