Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Stability: Stable Audio 3 Small Music

Stable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes from text prompts, lightweight enough for on-device deployment.
Compare models
stable-audio-3-small-music-text-to-audio
AvailableStabilityMusicAsync
Price
From $0.0217
Modalities
Audio

About Stable Audio 3 Small Music

Stable Audio 3 Small Music is the music-focused small checkpoint of Stability AI's Stable Audio 3 family. It is a 459 million parameter latent diffusion model that makes full stereo compositions up to two minutes long from text prompts, light enough for on-device deployment. Its siblings Small SFX and Small SFX Base concentrate on sound effects instead.

Where it works well

  • Stereo music, not only loops, from a text description.
  • Small enough for on-device use according to Stability, which also keeps generation time low.
  • Accepts English text prompts with a specified length.

When to choose another model

  • Output is limited to about two minutes, so longer pieces need to be stitched.
  • Larger Stable Audio 3 models exist for more demanding production work.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    MusicTokenLab endpoint
    POST/v1/music/generations
    curl -X POST "https://api.tokenlab.sh/v1/music/generations" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "model": "stable-audio-3-small-music-text-to-audio",
      "prompt": "Warm ambient piano with subtle rain textures.",
      "title": "Evening Drift"
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Music

per request
Official price
$0.0217
Official
$0.0217
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Stable Audio 3 Small Music in Console with a prompt ready to edit or send.

Help me create with stable-audio-3-small-music-text-to-audio using music at /v1/music/generations. Show the result, status, and cost.

Use cases

  • Game and app music

    Generate loops and themes of the required length for menus, levels, and onboarding, then pick the takes that fit your product's mood.

  • Video background tracks

    Make a stereo music bed for a short film, explainer, or social video, specifying the duration so the track ends with the edit.

  • Idea sketching

    Prototype a mood, genre, or instrumentation in seconds, before deciding whether to commission a composer or license a finished track.

  • Lightweight creative tools

    Build creative apps around a model that Stability designed to run on modest hardware, with short waits between prompts and results.

Prompt examples

Warm acoustic folk with fingerpicked guitar and light percussion, 100 BPM, 60 seconds.

Dark ambient drone with slow pulsing bass and distant chimes.

Upbeat 8-bit chiptune level theme, bright and playful.

FAQ

What is Stable Audio 3 Small Music?

A compact Stable Audio 3 model that generates full stereo music compositions from English text prompts. It is a latent diffusion model light enough for on-device deployment.

How long can Small Music tracks be?

Compositions run up to two minutes. Longer pieces need to be built from several generations, or from a larger Stable Audio 3 model such as Medium, which Stability positions for full songs.

How is Small Music different from Small SFX?

Small Music targets musical compositions, while Small SFX is tuned for sound effects. They share the compact size and the same text-prompt approach, so choose by the kind of audio you need, not by the interface.

Is Small Music open-weight?

Stability publishes the Small checkpoints on Hugging Face under its Community License, with terms for commercial use on its licence page. Using the model through an API is separate from running the weights yourself.

What data was Small Music trained on?

The model card describes recordings from AudioSparx and Freesound, with copyrighted music filtered out by automated detection. It also notes a pre-trained T5Gemma text encoder that carries its own separate terms of use.

How much does Stable Audio 3 Small Music cost?

On TokenLab, Stable Audio 3 Small Music costs $0.0217 per request. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Stable Audio 3 Small Music use?

Use https://api.tokenlab.sh/v1/music/generations for Stable Audio 3 Small Music. The request example below shows the matching code shape.

Which operations does Stable Audio 3 Small Music support?

Stable Audio 3 Small Music supports Music. Select an operation above to see its endpoint and request example.

Compare Stable Audio 3 Small Music

Sources

Reviewed Oct 2, 2026

More from Stable Audio

Related models