OpenAI: GPT-5 mini
gpt-5-mini- Input / Output-30%
- $0.25 / $2.00$0.175 / $1.40
- Context
- 400K
- Max output
- 128K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt CacheReasoning
About GPT-5 mini
GPT-5 mini is OpenAI's smaller GPT-5 model, described as a faster and more cost-efficient version of GPT-5 for well-defined, high-throughput work. It accepts text and images, returns text, and keeps reasoning tokens, function calling, and structured outputs. Pick it over the full model when latency and volume matter more than peak accuracy.
Where it works well
- Positioned by OpenAI as the faster version of GPT-5 for time-sensitive, high-throughput applications.
- Keeps reasoning tokens, so it can still work through multi-step problems at a smaller scale.
- Handles function calling, web search, file search, and structured outputs like its larger sibling.
- Takes image input, which allows screenshot and document questions without a separate vision model.
When to choose another model
- Its knowledge cutoff of May 2024 is earlier than the full GPT-5, so it knows less about recent events.
- On the hardest reasoning and large refactors the full GPT-5 or a Pro tier is the safer pick.
- It returns text only, with no audio or image output.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "gpt-5-mini", "input": "Hello!" }'
Pricing
The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.
Default spec
per 1M tokens- Official price
- Input $0.25 / Output $2.00 / Cache read $0.025
- Verified price
- Input $0.175 / Output $1.40 / Cache read $0.0175
- Discount
- -30%
Cache read
- Official price
- $0.025
- Verified price
- $0.0175
- Discount
- -30%
Charged per call, only when the model uses the tool.
Web search
- Official price
- $0.01/search
- Verified price
- $0.007/search
- Discount
- -30%
| Official priceper 1M tokens | Verified priceper 1M tokens | Discount | |
|---|---|---|---|
| Default spec | Input $0.25 / Output $2.00 / Cache read $0.025 | Input $0.175 / Output $1.40 / Cache read $0.0175 | -30% |
| Prompt cache pricing | |||
| Cache read | $0.025 | $0.0175 | -30% |
| Tool feesCharged per call, only when the model uses the tool. | |||
| Web search | $0.01/search | $0.007/search | -30% |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open GPT-5 mini in Console with a prompt ready to edit or send.
Help me try gpt-5-mini with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Best for- Reasoning
- Vision
Customer-facing assistants
Answer product questions with tool calls into an order system where response time shapes the experience.
Batch classification
Label thousands of tickets or reviews with a structured output schema and consistent categories.
Code explanation and small edits
Explain functions, write tests, and make contained edits when a full-size model is more than the task needs.
Agent sub-steps
Run routine lookups and summaries inside a larger workflow while a stronger model handles planning.
Prompt examples
Classify this support ticket as billing, bug, or feature request and return the label with one sentence of reasoning.
Summarize this call transcript in five bullet points and list any action items with owners.
Write unit tests for this function and point out any input it fails to handle.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is GPT-5 mini?
GPT-5 mini is a smaller version of GPT-5 from OpenAI. It is aimed at fast, cost-conscious work with high request volume, and it still supports reasoning, tools, and image input.
When should I use GPT-5 mini instead of GPT-5?
Use mini for well-scoped tasks such as classification, summarization, and routine tool calls where speed and volume count. Use the full GPT-5 when answers are wrong too often on your own evaluation set or the task needs deeper reasoning.
Does GPT-5 mini accept images?
Yes. Text and images go in and text comes out. That covers reading screenshots, forms, and charts for questions or extraction, but the model cannot create or edit images itself.
What is the knowledge cutoff of GPT-5 mini?
OpenAI lists May 31, 2024 as the knowledge cutoff for this model. Anything newer has to come from the prompt, from attached files, or from a web search tool call.
Is GPT-5 mini smaller than GPT-5 nano?
No. Nano is the smaller and cheaper of the two and is aimed at summarization and classification. Mini sits between nano and the full GPT-5 in capability and cost.
How much does GPT-5 mini cost?
On TokenLab, GPT-5 mini costs Input $0.175 / Output $1.40 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of GPT-5 mini?
GPT-5 mini accepts up to 400,000 tokens of context and returns up to 128,000 tokens in one response.
Which endpoint should GPT-5 mini use?
Use https://api.tokenlab.sh/v1/responses for GPT-5 mini. The request example below shows the matching code shape.
Which operations does GPT-5 mini support?
GPT-5 mini supports Text to text. Select an operation above to see its endpoint and request example.
Compare GPT-5 mini
Sources
Reviewed Oct 2, 2026