StepFun: Step 3.7 Flash
step-3.7-flash- Input / Output
- $0.20 / $1.15
- Context
- 256K
- Max output
- 256K
- Modalities
- Chat
- Capabilities
- Tool usePrompt Cache
About Step 3.7 Flash
Step 3.7 Flash is StepFun's sparse mixture-of-experts model for coding agents, search workflows, and long-context productivity work. Its card describes a 198B-parameter design with about 11B active per token, and three selectable reasoning levels. It follows Step 3.5 Flash, adds structured-output support, and is released under Apache 2.0.
Where it works well
- Search and verification: the card describes an agent that checks claims against retrieved sources and resists adversarial pages.
- Agent reliability: the card reports 67.1 on ClawEval-1.1, a test of multi-turn workflows with adversarial traps.
- Repository work: it scores 56.3 on SWE-Bench Pro, second place in the card's comparison.
- Reasoning depth can be set to low, medium, or high, trading latency for accuracy per request.
- JSON mode, response-schema output, tool calling, and prompt caching are all supported.
When to choose another model
- StepFun itself flags Terminal-Bench at 59.5 and GDPVal-AA at 45.8 as areas needing improvement.
- It is a text model without image input, so screenshots and scanned pages need a vision-capable model first.
- The full 198B weights are heavy to self-host, so most teams use it through an API.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textMessages APIPOST/v1/messagesAPI format:curl https://api.tokenlab.sh/v1/messages \ -H "Content-Type: application/json" \ -H "x-api-key: sk-xxx" \ -H "anthropic-version: 2023-06-01" \ -d '{ "model": "step-3.7-flash", "max_tokens": 1024, "messages": [ {"role": "user", "content": "Hello!"} ] }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Default
per 1M tokens- Official price
- Input $0.20 / Output $1.15 / Cache read $0.04 / Text Input $0.20 / Text Output $1.15
- Official
- Input $0.20 / Output $1.15 / Cache read $0.04 / Text Input $0.20 / Text Output $1.15
- Discount
- —
Cache read
- Official price
- $0.04
- Official
- $0.04
- Discount
- —
| Official priceper 1M tokens | Officialper 1M tokens | Discount | |
|---|---|---|---|
| Default | Input $0.20 / Output $1.15 / Cache read $0.04 / Text Input $0.20 / Text Output $1.15 | Input $0.20 / Output $1.15 / Cache read $0.04 / Text Input $0.20 / Text Output $1.15 | — |
| Prompt cache pricing | |||
| Cache read | $0.04 | $0.04 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Step 3.7 Flash in Console with a prompt ready to edit or send.
Help me try step-3.7-flash with a short message at /v1/messages. Show the reply, latency, and cost.
Use cases
Concurrent coding agents
Run several agents against one repository, each tracing code paths and producing patches, with a lower reasoning level for routine edits.
Search and verification loops
Have an agent search, read sources, and check its own claims before answering, using the model's resistance to adversarial pages.
Structured extraction
Return typed JSON from invoices, tickets, or filings with response-schema enforcement instead of parsing free text.
Long-context summarization
Condense long reports, transcripts, and ticket histories into briefs, using a lower reasoning level for routine summaries and a higher one for analysis.
Prompt examples
Search the repository for every caller of billing.charge(), list those that ignore the return value, and propose a patch for each.
Extract vendor, invoice number, line items, and total from this text as JSON matching the schema I gave you.
Check each claim in this draft against the sources you can retrieve and mark every statement you could not confirm.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is Step 3.7 Flash?
It is StepFun's sparse mixture-of-experts model aimed at agentic coding, search, and long-context work. The model card lists 198B total parameters, about 11B active per token, and an Apache 2.0 license. StepFun positions it for developers scaling agent workflows.
What are the reasoning levels?
StepFun provides three selectable levels: low, medium, and high. Use low for routine tool calls and short answers, and high for debugging, planning, or multi-step verification where accuracy matters more than response time.
Does Step 3.7 Flash support structured output?
Yes. It supports response schemas alongside JSON mode and tool calling, so you can constrain replies to a schema. That sets it apart from Step 3.5 Flash, which has no response-schema support.
Can it read images?
No. It is a text model, so image input is not supported. For screenshots or scanned pages, use Step 5 Preview, which takes images and video, or extract the text first.
When should I pick Step 5 Preview instead?
Choose Step 5 Preview when you need StepFun's flagship, a context window in the million-token range, or video input. Step 3.7 Flash is the faster, smaller option for high-volume agent loops.
How much does Step 3.7 Flash cost?
On TokenLab, Step 3.7 Flash costs Input $0.20 / Output $1.15 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Step 3.7 Flash?
Step 3.7 Flash accepts up to 262,144 tokens of context and returns up to 262,144 tokens in one response.
Which endpoint should Step 3.7 Flash use?
Use https://api.tokenlab.sh/v1/messages for Step 3.7 Flash. The request example below shows the matching code shape.
Which operations does Step 3.7 Flash support?
Step 3.7 Flash supports Text to text. Select an operation above to see its endpoint and request example.
Compare Step 3.7 Flash
Sources
Reviewed Oct 2, 2026