Anthropic: Claude Opus 4.8
claude-opus-4-8- Input / Output-70%
- $5.00 / $25.00$1.50 / $7.50
- Context
- 1M
- Released
- May 28, 2026
- Max output
- 128K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt CacheReasoning
About Claude Opus 4.8
Claude Opus 4.8 is Anthropic's Opus-tier model released on May 28, 2026, sitting between Opus 4.7 and Opus 5 in the lineup. It handles the hardest reasoning, coding, and long-horizon agent work of its generation, with adaptive thinking, image input, and tool use. Anthropic lists it as legacy and recommends Opus 5.5 for new work.
Where it works well
- Adaptive thinking with a high default effort, so hard prompts get reasoning without setting a manual token budget.
- Accepts text and images and answers in text, which covers screenshot review and diagram questions in an agent loop.
- Uses the tokenizer introduced with Opus 4.7, so token counts match later Claude models more closely than Opus 4.6 does.
- Fable 5 falls back to this model for requests its safeguards decline, so Anthropic treats it as the top safe general model beneath Fable.
When to choose another model
- Opus 5 delivers markedly better results on Anthropic's reported evaluations, so 4.8 is hard to justify for new builds.
- Its reliable knowledge cutoff is January 2026, earlier than Opus 5.5 and Sonnet 5.5.
- Anthropic marks it legacy, with a stated recommendation to migrate to Opus 5.5.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textMessages APIPOST/v1/messagesAPI format:curl https://api.tokenlab.sh/v1/messages \ -H "Content-Type: application/json" \ -H "x-api-key: sk-xxx" \ -H "anthropic-version: 2023-06-01" \ -d '{ "model": "claude-opus-4-8", "max_tokens": 1024, "messages": [ {"role": "user", "content": "Hello!"} ] }'
Pricing
TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.
Default spec
per 1M tokens- Official price
- Input $5.00 / Output $25.00 / Cache read $0.50 / Cache write $6.25
- TokenLab price
- Input $1.50 / Output $7.50 / Cache read $0.15 / Cache write $1.88
- Discount
- -70%
Cache read
- Official price
- $0.50
- TokenLab price
- $0.15
- Discount
- -70%
Cache write
- Official price
- $6.25
- TokenLab price
- $1.88
- Discount
- -70%
Charged per call, only when the model uses the tool.
Web search
- Official price
- $0.01/search
- TokenLab price
- $0.003/search
- Discount
- -70%
| Official priceper 1M tokens | TokenLab priceper 1M tokens | Discount | |
|---|---|---|---|
| Default spec | Input $5.00 / Output $25.00 / Cache read $0.50 / Cache write $6.25 | Input $1.50 / Output $7.50 / Cache read $0.15 / Cache write $1.88 | -70% |
| Prompt cache pricing | |||
| Cache read | $0.50 | $0.15 | -70% |
| Cache write | $6.25 | $1.88 | -70% |
| Tool feesCharged per call, only when the model uses the tool. | |||
| Web search | $0.01/search | $0.003/search | -70% |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
- 30-day success rate
- 100.0%
- 30-day requests
- 100+
- Last active
- 3 days ago
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Claude Opus 4.8 in Console with a prompt ready to edit or send.
Help me try claude-opus-4-8 with a short message at /v1/messages. Show the reply, latency, and cost.
Use cases
Best for- Reasoning
- Vision
Existing agent pipelines
Keep a validated coding or research agent on a fixed model version while you schedule the move to Opus 5.5 and re-run your own evaluations.
Fallback target
Use it as the retry model when a request to a Fable model is declined by a safeguard classifier but still needs a strong answer.
Deep code review
Run a long review over a large diff, asking for ranked risks, missing tests, and the smallest safe fix for each.
Prompt examples
Audit this module for race conditions and explain each one with the exact interleaving that triggers it.
Our incident timeline and the relevant logs follow. Write a root-cause analysis and a list of follow-up changes.
Plan a staged rollout for replacing this billing library, including how to confirm each stage worked.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
Where does Claude Opus 4.8 fit among Claude models?
It is the Opus release between 4.7 and Opus 5, launched on May 28, 2026. It targets the hardest reasoning and coding work of the 4.x generation. Anthropic now lists it as legacy and points to Opus 5.5 as the current Opus.
Does Claude Opus 4.8 support adaptive thinking?
Yes. Adaptive thinking is its thinking mode, and the default effort is high on Anthropic's API. Lower the effort for faster, cheaper turns, or raise it on tasks where a missed edge case is expensive.
Should I switch from Opus 4.8 to Opus 5.5?
For new work, yes: Anthropic recommends Opus 5.5 and publishes a migration guide for 4.8. Keep 4.8 only where a validated pipeline depends on its exact behavior, and note that 5.5 changes thinking and tool-use rules.
Can Claude Opus 4.8 read images?
Yes. It takes text and images as input and returns text. It cannot create images, so for picture generation you need a separate image model alongside it.
How much does Claude Opus 4.8 cost?
On TokenLab, Claude Opus 4.8 costs Input $1.50 / Output $7.50 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Claude Opus 4.8?
Claude Opus 4.8 accepts up to 1,000,000 tokens of context and returns up to 128,000 tokens in one response.
Which endpoint should Claude Opus 4.8 use?
Use https://api.tokenlab.sh/v1/messages for Claude Opus 4.8. The request example below shows the matching code shape.
Which operations does Claude Opus 4.8 support?
Claude Opus 4.8 supports Text to text. Select an operation above to see its endpoint and request example.
Compare Claude Opus 4.8
Guides that use Claude Opus 4.8
Sources
Reviewed Oct 2, 2026