Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

OpenAI: GPT-5 mini

Small GPT-5 for responsive agents, coding help, and everyday automation
Compare models
gpt-5-mini
AvailableOpenAIChatCached input 90% less
Input / Output-30%
$0.25 / $2.00$0.175 / $1.40
Context
400K
Max output
128K
Modalities
VisionChat
Capabilities
Tool usePrompt CacheReasoning

About GPT-5 mini

GPT-5 mini is OpenAI's smaller GPT-5 model, described as a faster and more cost-efficient version of GPT-5 for well-defined, high-throughput work. It accepts text and images, returns text, and keeps reasoning tokens, function calling, and structured outputs. Pick it over the full model when latency and volume matter more than peak accuracy.

Where it works well

  • Positioned by OpenAI as the faster version of GPT-5 for time-sensitive, high-throughput applications.
  • Keeps reasoning tokens, so it can still work through multi-step problems at a smaller scale.
  • Handles function calling, web search, file search, and structured outputs like its larger sibling.
  • Takes image input, which allows screenshot and document questions without a separate vision model.

When to choose another model

  • Its knowledge cutoff of May 2024 is earlier than the full GPT-5, so it knows less about recent events.
  • On the hardest reasoning and large refactors the full GPT-5 or a Pro tier is the safer pick.
  • It returns text only, with no audio or image output.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textResponses API
    POST/v1/responses
    API format:
    curl https://api.tokenlab.sh/v1/responses \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "gpt-5-mini",
        "input": "Hello!"
      }'

Pricing

The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.

Default spec

per 1M tokens
Official price
Input $0.25 / Output $2.00 / Cache read $0.025
Verified price
Input $0.175 / Output $1.40 / Cache read $0.0175
Discount
-30%
Prompt cache pricing

Cache read

Official price
$0.025
Verified price
$0.0175
Discount
-30%
Tool fees

Charged per call, only when the model uses the tool.

Web search

Official price
$0.01/search
Verified price
$0.007/search
Discount
-30%

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open GPT-5 mini in Console with a prompt ready to edit or send.

Help me try gpt-5-mini with a short message at /v1/responses. Show the reply, latency, and cost.

Use cases

Best for
  • Reasoning
  • Vision
  • Customer-facing assistants

    Answer product questions with tool calls into an order system where response time shapes the experience.

  • Batch classification

    Label thousands of tickets or reviews with a structured output schema and consistent categories.

  • Code explanation and small edits

    Explain functions, write tests, and make contained edits when a full-size model is more than the task needs.

  • Agent sub-steps

    Run routine lookups and summaries inside a larger workflow while a stronger model handles planning.

Prompt examples

Classify this support ticket as billing, bug, or feature request and return the label with one sentence of reasoning.

Summarize this call transcript in five bullet points and list any action items with owners.

Write unit tests for this function and point out any input it fails to handle.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is GPT-5 mini?

GPT-5 mini is a smaller version of GPT-5 from OpenAI. It is aimed at fast, cost-conscious work with high request volume, and it still supports reasoning, tools, and image input.

When should I use GPT-5 mini instead of GPT-5?

Use mini for well-scoped tasks such as classification, summarization, and routine tool calls where speed and volume count. Use the full GPT-5 when answers are wrong too often on your own evaluation set or the task needs deeper reasoning.

Does GPT-5 mini accept images?

Yes. Text and images go in and text comes out. That covers reading screenshots, forms, and charts for questions or extraction, but the model cannot create or edit images itself.

What is the knowledge cutoff of GPT-5 mini?

OpenAI lists May 31, 2024 as the knowledge cutoff for this model. Anything newer has to come from the prompt, from attached files, or from a web search tool call.

Is GPT-5 mini smaller than GPT-5 nano?

No. Nano is the smaller and cheaper of the two and is aimed at summarization and classification. Mini sits between nano and the full GPT-5 in capability and cost.

How much does GPT-5 mini cost?

On TokenLab, GPT-5 mini costs Input $0.175 / Output $1.40 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of GPT-5 mini?

GPT-5 mini accepts up to 400,000 tokens of context and returns up to 128,000 tokens in one response.

Which endpoint should GPT-5 mini use?

Use https://api.tokenlab.sh/v1/responses for GPT-5 mini. The request example below shows the matching code shape.

Which operations does GPT-5 mini support?

GPT-5 mini supports Text to text. Select an operation above to see its endpoint and request example.

Compare GPT-5 mini

Sources

Reviewed Oct 2, 2026

More from GPT 5

Related models