Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Stability: Stable Audio 3 Small SFX Base

Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.
Compare models
stable-audio-3-small-sfx-base-text-to-audio
AvailableStabilityMusicAsync
Price
From $0.0283
Modalities
Audio

About Stable Audio 3 Small SFX Base

Stable Audio 3 Small SFX Base is the unmodified foundation checkpoint behind Stable Audio 3 Small SFX, released by Stability AI as a starting point for fine-tuning. Stability says it has not had the adversarial post-training that the standard model uses for faster, better output, and recommends the standard checkpoint for direct generation. It still produces sound effects from text.

Where it works well

  • It is the pre-trained base, so outputs reflect the unadjusted model.
  • Useful for comparing raw output with the post-trained Small SFX.
  • Same sound-effects focus and compact size as its sibling.

When to choose another model

  • Stability recommends the standard Small SFX for direct generation.
  • Without post-training it may be slower or rougher than the standard model.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    MusicTokenLab endpoint
    POST/v1/music/generations
    curl -X POST "https://api.tokenlab.sh/v1/music/generations" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "model": "stable-audio-3-small-sfx-base-text-to-audio",
      "prompt": "Warm ambient piano with subtle rain textures.",
      "title": "Evening Drift"
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Music

per request
Official price
$0.0283
Official
$0.0283
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Stable Audio 3 Small SFX Base in Console with a prompt ready to edit or send.

Help me create with stable-audio-3-small-sfx-base-text-to-audio using music at /v1/music/generations. Show the result, status, and cost.

Use cases

  • Quality comparison

    Generate the same prompt on this base checkpoint and on the standard Small SFX model to hear what post-training changes in speed and polish.

  • Baseline evaluation

    Measure how the pre-trained model behaves before any customisation, so later fine-tuned versions can be judged against a known starting point.

  • Research prompts

    Probe what the pre-trained model produces for unusual sound descriptions, noting which kinds of effects it handles well and which it misses.

  • Pipeline testing

    Check that an integration handles more than one checkpoint, with identical request shapes and different expectations about the quality of the audio.

Prompt examples

Glass bottle rolling across a tile floor and stopping, 4 seconds.

Distant helicopter passing over a forest, fading out.

Cartoon spring boing, exaggerated, 1 second.

FAQ

What is the Stable Audio 3 Small SFX Base model?

The base, pre-trained checkpoint of Stable Audio 3 Small SFX. Stability released it as the unmodified starting point for fine-tuning, and it generates sound effects from text prompts like its post-trained sibling.

Should I use the Base model for generation?

Usually not. Stability's model card says that if you want to generate audio directly you should use Stable Audio 3 Small SFX instead, and keeps the base checkpoint for people who intend to fine-tune.

What does the Base checkpoint lack?

It has not had the adversarial post-training that the standard model uses. According to the model card, that step accelerates inference and improves quality, so expect the standard model to be faster and more polished.

Can I fine-tune the Base model through this API?

This listing only generates audio; it does not train anything. Fine-tuning needs the open weights, which Stability publishes on Hugging Face under its Community License, and your own training setup.

What does the Base model output?

Sound effects generated from English text prompts, in the same style of request as the standard Small SFX model. Quality and speed differ because the base checkpoint skips the final post-training step described by Stability.

How much does Stable Audio 3 Small SFX Base cost?

On TokenLab, Stable Audio 3 Small SFX Base costs $0.0283 per request. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Stable Audio 3 Small SFX Base use?

Use https://api.tokenlab.sh/v1/music/generations for Stable Audio 3 Small SFX Base. The request example below shows the matching code shape.

Which operations does Stable Audio 3 Small SFX Base support?

Stable Audio 3 Small SFX Base supports Music. Select an operation above to see its endpoint and request example.

Compare Stable Audio 3 Small SFX Base

Sources

Reviewed Oct 2, 2026

More from Stable Audio

Related models