Stability: Stable Audio 3.0
stable-audio-3- Price
- From $0.26
- Modalities
- Audio
About Stable Audio 3.0
Stable Audio 3.0 is Stability AI's family of audio generation models, producing music and sound from written descriptions with variable-length output. It returns MP3 or WAV. Stability lists Large for enterprise sound production, Medium for full songs, and Small and Small SFX for generation on a mobile device. You choose the duration to fit a soundtrack, an instrumental idea, or an effect.
Where it works well
- Duration is variable, so a clip can match the length a project needs.
- Covers both music and sound effects within one family.
- Output can be delivered as MP3 or WAV.
When to choose another model
- Output is instrumental music and sound effects from a description; it is not tuned for sung lyrics or a specific voice.
- For lyrics and vocal songs, a song-focused music service is a better fit.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
MusicTokenLab endpointPOST/v1/music/generationscurl -X POST "https://api.tokenlab.sh/v1/music/generations" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "duration": 30, "output_format": "mp3", "model": "stable-audio-3", "prompt": "Warm ambient piano with subtle rain textures.", "title": "Evening Drift" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Pricing
per request- Official
- $0.26per request
- Official price
- $0.26per request
| Official priceper request | Officialper request | Discount | |
|---|---|---|---|
| Price | $0.26per request | $0.26per request | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Stable Audio 3.0 in Console with a prompt ready to edit or send.
Help me create with stable-audio-3 using music at /v1/music/generations. Show the result, status, and cost.
Use cases
Soundtrack beds
Generate background music of a chosen length to sit under a video, a game menu, or a product demo, then trim or loop it in your editor.
Instrumental sketches
Try ideas for genre, tempo, and mood quickly and cheaply before commissioning a composer or choosing a licensed track for a project.
Sound effects
Create impacts, ambience, and foley for scenes and interfaces from a short written description, picking the length that matches the cue you need.
Prototype audio
Fill an app or game demo with placeholder sound so reviewers can judge pacing and feel before final audio is produced by a sound designer.
Prompt examples
Calm lo-fi hip hop, soft piano and vinyl crackle, 90 BPM, 45 seconds.
Epic orchestral trailer build with brass swells and a hard stop at 30 seconds.
Rain on a tin roof with distant thunder, 15 seconds.
FAQ
What is Stable Audio 3.0?
It is Stability AI's family of text-to-audio models for music and sound effects. Output has variable length, so you choose the duration, and the family spans sizes from large studio-oriented models to small ones built for mobile devices.
Which Stable Audio 3 sizes exist?
Stability's page lists Large for enterprise sound production, Medium for full song composition with open weights, and Small and Small SFX for generation on a mobile device. The Small checkpoints are also offered as separate models.
What formats does it return?
MP3 and WAV. MP3 gives a compact file for previews and web use, while WAV keeps uncompressed audio that suits editing in a digital audio workstation or a video timeline.
Can I set the length of the audio?
Yes. Variable-length generation is a stated feature of the family, so you specify the duration you need instead of cutting a fixed clip afterward. Match it to the cue, scene, or loop you are building.
Does Stable Audio 3.0 sing lyrics?
Stability describes the family as generating music and sound effects from text, not lyric-led songs. If you need a sung vocal with words, choose a dedicated song-generation model such as Suno instead.
How much does Stable Audio 3.0 cost?
On TokenLab, Stable Audio 3.0 costs $0.26 per request. The pricing table above shows the full breakdown.
Which endpoint should Stable Audio 3.0 use?
Use https://api.tokenlab.sh/v1/music/generations for Stable Audio 3.0. The request example below shows the matching code shape.
Which operations does Stable Audio 3.0 support?
Stable Audio 3.0 supports Music. Select an operation above to see its endpoint and request example.
Compare Stable Audio 3.0
Sources
Reviewed Oct 2, 2026