Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Stability: Stable Audio 3.0

Stable Audio 3 generates music and sound from written descriptions. Its variable-length audio generation is useful for matching a soundtrack, an instrumental idea, or a sound effect to the duration needed by a project.
Compare models
stable-audio-3
AvailableStabilityMusicAsync
Price
From $0.26
Modalities
Audio

About Stable Audio 3.0

Stable Audio 3.0 is Stability AI's family of audio generation models, producing music and sound from written descriptions with variable-length output. It returns MP3 or WAV. Stability lists Large for enterprise sound production, Medium for full songs, and Small and Small SFX for generation on a mobile device. You choose the duration to fit a soundtrack, an instrumental idea, or an effect.

Where it works well

  • Duration is variable, so a clip can match the length a project needs.
  • Covers both music and sound effects within one family.
  • Output can be delivered as MP3 or WAV.

When to choose another model

  • Output is instrumental music and sound effects from a description; it is not tuned for sung lyrics or a specific voice.
  • For lyrics and vocal songs, a song-focused music service is a better fit.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    MusicTokenLab endpoint
    POST/v1/music/generations
    curl -X POST "https://api.tokenlab.sh/v1/music/generations" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "duration": 30,
      "output_format": "mp3",
      "model": "stable-audio-3",
      "prompt": "Warm ambient piano with subtle rain textures.",
      "title": "Evening Drift"
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Pricing

per request
Official
$0.26per request
Official price
$0.26per request

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Stable Audio 3.0 in Console with a prompt ready to edit or send.

Help me create with stable-audio-3 using music at /v1/music/generations. Show the result, status, and cost.

Use cases

  • Soundtrack beds

    Generate background music of a chosen length to sit under a video, a game menu, or a product demo, then trim or loop it in your editor.

  • Instrumental sketches

    Try ideas for genre, tempo, and mood quickly and cheaply before commissioning a composer or choosing a licensed track for a project.

  • Sound effects

    Create impacts, ambience, and foley for scenes and interfaces from a short written description, picking the length that matches the cue you need.

  • Prototype audio

    Fill an app or game demo with placeholder sound so reviewers can judge pacing and feel before final audio is produced by a sound designer.

Prompt examples

Calm lo-fi hip hop, soft piano and vinyl crackle, 90 BPM, 45 seconds.

Epic orchestral trailer build with brass swells and a hard stop at 30 seconds.

Rain on a tin roof with distant thunder, 15 seconds.

FAQ

What is Stable Audio 3.0?

It is Stability AI's family of text-to-audio models for music and sound effects. Output has variable length, so you choose the duration, and the family spans sizes from large studio-oriented models to small ones built for mobile devices.

Which Stable Audio 3 sizes exist?

Stability's page lists Large for enterprise sound production, Medium for full song composition with open weights, and Small and Small SFX for generation on a mobile device. The Small checkpoints are also offered as separate models.

What formats does it return?

MP3 and WAV. MP3 gives a compact file for previews and web use, while WAV keeps uncompressed audio that suits editing in a digital audio workstation or a video timeline.

Can I set the length of the audio?

Yes. Variable-length generation is a stated feature of the family, so you specify the duration you need instead of cutting a fixed clip afterward. Match it to the cue, scene, or loop you are building.

Does Stable Audio 3.0 sing lyrics?

Stability describes the family as generating music and sound effects from text, not lyric-led songs. If you need a sung vocal with words, choose a dedicated song-generation model such as Suno instead.

How much does Stable Audio 3.0 cost?

On TokenLab, Stable Audio 3.0 costs $0.26 per request. The pricing table above shows the full breakdown.

Which endpoint should Stable Audio 3.0 use?

Use https://api.tokenlab.sh/v1/music/generations for Stable Audio 3.0. The request example below shows the matching code shape.

Which operations does Stable Audio 3.0 support?

Stable Audio 3.0 supports Music. Select an operation above to see its endpoint and request example.

Compare Stable Audio 3.0

Sources

Reviewed Oct 2, 2026

More from Stable Audio

Related models