Stability: Stable Audio 3 Small SFX Base
stable-audio-3-small-sfx-base-text-to-audio- Price
- From $0.0283
- Modalities
- Audio
About Stable Audio 3 Small SFX Base
Stable Audio 3 Small SFX Base is the unmodified foundation checkpoint behind Stable Audio 3 Small SFX, released by Stability AI as a starting point for fine-tuning. Stability says it has not had the adversarial post-training that the standard model uses for faster, better output, and recommends the standard checkpoint for direct generation. It still produces sound effects from text.
Where it works well
- It is the pre-trained base, so outputs reflect the unadjusted model.
- Useful for comparing raw output with the post-trained Small SFX.
- Same sound-effects focus and compact size as its sibling.
When to choose another model
- Stability recommends the standard Small SFX for direct generation.
- Without post-training it may be slower or rougher than the standard model.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
MusicTokenLab endpointPOST/v1/music/generationscurl -X POST "https://api.tokenlab.sh/v1/music/generations" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "model": "stable-audio-3-small-sfx-base-text-to-audio", "prompt": "Warm ambient piano with subtle rain textures.", "title": "Evening Drift" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Music
per request- Official price
- $0.0283
- Official
- $0.0283
- Discount
- —
| Official priceper request | Officialper request | Discount | |
|---|---|---|---|
| Music | $0.0283 | $0.0283 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Stable Audio 3 Small SFX Base in Console with a prompt ready to edit or send.
Help me create with stable-audio-3-small-sfx-base-text-to-audio using music at /v1/music/generations. Show the result, status, and cost.
Use cases
Quality comparison
Generate the same prompt on this base checkpoint and on the standard Small SFX model to hear what post-training changes in speed and polish.
Baseline evaluation
Measure how the pre-trained model behaves before any customisation, so later fine-tuned versions can be judged against a known starting point.
Research prompts
Probe what the pre-trained model produces for unusual sound descriptions, noting which kinds of effects it handles well and which it misses.
Pipeline testing
Check that an integration handles more than one checkpoint, with identical request shapes and different expectations about the quality of the audio.
Prompt examples
Glass bottle rolling across a tile floor and stopping, 4 seconds.
Distant helicopter passing over a forest, fading out.
Cartoon spring boing, exaggerated, 1 second.
FAQ
What is the Stable Audio 3 Small SFX Base model?
The base, pre-trained checkpoint of Stable Audio 3 Small SFX. Stability released it as the unmodified starting point for fine-tuning, and it generates sound effects from text prompts like its post-trained sibling.
Should I use the Base model for generation?
Usually not. Stability's model card says that if you want to generate audio directly you should use Stable Audio 3 Small SFX instead, and keeps the base checkpoint for people who intend to fine-tune.
What does the Base checkpoint lack?
It has not had the adversarial post-training that the standard model uses. According to the model card, that step accelerates inference and improves quality, so expect the standard model to be faster and more polished.
Can I fine-tune the Base model through this API?
This listing only generates audio; it does not train anything. Fine-tuning needs the open weights, which Stability publishes on Hugging Face under its Community License, and your own training setup.
What does the Base model output?
Sound effects generated from English text prompts, in the same style of request as the standard Small SFX model. Quality and speed differ because the base checkpoint skips the final post-training step described by Stability.
How much does Stable Audio 3 Small SFX Base cost?
On TokenLab, Stable Audio 3 Small SFX Base costs $0.0283 per request. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Stable Audio 3 Small SFX Base use?
Use https://api.tokenlab.sh/v1/music/generations for Stable Audio 3 Small SFX Base. The request example below shows the matching code shape.
Which operations does Stable Audio 3 Small SFX Base support?
Stable Audio 3 Small SFX Base supports Music. Select an operation above to see its endpoint and request example.
Compare Stable Audio 3 Small SFX Base
Sources
Reviewed Oct 2, 2026