Anthropic: Claude Haiku 4.5
claude-haiku-4-5- Input / Output-70%
- $1.00 / $5.00$0.30 / $1.50
- Context
- 200K
- Max output
- 64K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt CacheReasoning
About Claude Haiku 4.5
Claude Haiku 4.5 is Anthropic's small, fast Claude model, released in October 2025 for high-volume and latency-sensitive work. Anthropic describes it as the fastest model in its lineup with near-frontier intelligence. It accepts text and images, supports tool use and prompt caching, and uses manual extended thinking with a token budget rather than the adaptive thinking of newer Claude models.
Where it works well
- Lowest latency among the Claude models listed on Anthropic's overview page, which suits chat front ends and real-time assistants.
- Reads images next to text, so it can classify screenshots, receipts, and charts at volume.
- Extended thinking is optional and set with an explicit token budget, so you decide per request how much reasoning to pay for.
- Supports tool use, structured JSON output, and prompt caching, which keeps repeated system prompts cheap in agent loops.
When to choose another model
- Its knowledge cutoff is much earlier than the Claude 5 generation, so it knows less about recent libraries and events.
- It has a smaller context window and output cap than the Opus and Sonnet 5.x models, which limits whole-repository or book-length jobs.
- For long autonomous coding runs or ambiguous multi-step planning, a Sonnet or Opus model fails less often.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textMessages APIPOST/v1/messagesAPI format:curl https://api.tokenlab.sh/v1/messages \ -H "Content-Type: application/json" \ -H "x-api-key: sk-xxx" \ -H "anthropic-version: 2023-06-01" \ -d '{ "model": "claude-haiku-4-5", "max_tokens": 1024, "messages": [ {"role": "user", "content": "Hello!"} ] }'
Pricing
TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.
Default spec
per 1M tokens- Official price
- Input $1.00 / Output $5.00 / Cache read $0.10 / Cache write $1.25
- TokenLab price
- Input $0.30 / Output $1.50 / Cache read $0.03 / Cache write $0.375
- Discount
- -70%
Cache read
- Official price
- $0.10
- TokenLab price
- $0.03
- Discount
- -70%
Cache write
- Official price
- $1.25
- TokenLab price
- $0.375
- Discount
- -70%
Charged per call, only when the model uses the tool.
Web search
- Official price
- $0.01/search
- TokenLab price
- $0.003/search
- Discount
- -70%
| Official priceper 1M tokens | TokenLab priceper 1M tokens | Discount | |
|---|---|---|---|
| Default spec | Input $1.00 / Output $5.00 / Cache read $0.10 / Cache write $1.25 | Input $0.30 / Output $1.50 / Cache read $0.03 / Cache write $0.375 | -70% |
| Prompt cache pricing | |||
| Cache read | $0.10 | $0.03 | -70% |
| Cache write | $1.25 | $0.375 | -70% |
| Tool feesCharged per call, only when the model uses the tool. | |||
| Web search | $0.01/search | $0.003/search | -70% |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
- 30-day success rate
- 100.0%
- 7-day median latency
- 1.6 sn=364
- 7-day P95 latency
- 15.8 sn=364
- 30-day requests
- 100+
- Last active
- 7 minutes ago
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Claude Haiku 4.5 in Console with a prompt ready to edit or send.
Help me try claude-haiku-4-5 with a short message at /v1/messages. Show the reply, latency, and cost.
Use cases
Best for- Reasoning
- Vision
Real-time chat and support
Put it behind a customer-facing widget where reply speed matters more than depth, with tool calls for order lookups and a handoff when the question gets hard.
Classification and extraction
Tag tickets, pull fields from invoices or screenshots, and return strict JSON for a downstream pipeline that runs thousands of times a day.
Sub-agent workers
Run it as the cheap worker in a multi-agent setup: a larger model plans, and several Haiku calls search, summarize, or check files in parallel.
Prompt examples
Classify this support email as billing, bug, feature request, or other, and return JSON with the label and a one-sentence reason.
Summarize this chat transcript in three bullet points and list any action item with its owner.
Look at this screenshot of an error dialog and tell me which setting the user should change.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is Claude Haiku 4.5 best used for?
Fast, high-volume jobs: chat assistants, classification, data extraction, summarization, and worker agents. Anthropic positions it as the fastest Claude model with near-frontier intelligence, so it is the default pick when response time and throughput matter more than the deepest reasoning.
Does Claude Haiku 4.5 support extended thinking?
Yes, but in the manual form: you enable thinking and set a token budget. It does not use the adaptive thinking mode or the effort parameter that the newer Sonnet and Opus models use, so you control reasoning depth through the budget you pass.
Can Claude Haiku 4.5 process images?
Yes. It takes images alongside text and answers in text. That covers reading screenshots, charts, and scanned documents. It cannot generate or edit images, so use an image model when you need picture output.
When should I move up from Haiku 4.5 to Sonnet?
Move up when Haiku misses instructions on long prompts, drops steps in multi-tool agent runs, or needs knowledge of recent releases. Sonnet 5.5 is the next step and keeps a fast profile, with a longer context window and a much later knowledge cutoff.
How much does Claude Haiku 4.5 cost?
On TokenLab, Claude Haiku 4.5 costs Input $0.30 / Output $1.50 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Claude Haiku 4.5?
Claude Haiku 4.5 accepts up to 200,000 tokens of context and returns up to 64,000 tokens in one response.
Which endpoint should Claude Haiku 4.5 use?
Use https://api.tokenlab.sh/v1/messages for Claude Haiku 4.5. The request example below shows the matching code shape.
Which operations does Claude Haiku 4.5 support?
Claude Haiku 4.5 supports Text to text. Select an operation above to see its endpoint and request example.
Compare Claude Haiku 4.5
Guides that use Claude Haiku 4.5
Sources
Reviewed Oct 2, 2026