Text prompts
Write what you want in ordinary language — the second speaker, the rain, the bassline. There is no class list to memorise and no fixed set of stems to choose between.
Promptable audio source separation
SAM Audio Lab pulls a single sound out of a crowded recording from a plain-text prompt — no stem presets, no manual spectral editing. You get the sound you asked for and the residual mix beside it.
Works on audio files and on the soundtrack of a video.
Target
Residual
Capabilities
Write what you want in ordinary language — the second speaker, the rain, the bassline. There is no class list to memorise and no fixed set of stems to choose between.
When a sound is easier to show than to describe, mark it on the video frame where you hear it and the model takes the object you pointed at.
Select a stretch of the timeline that contains the sound. A few clean seconds are often all the model needs to identify the target.
Text, a visual pick and a time range stack together. Two weak cues usually resolve a target that neither one settles on its own.
A single model for every layer
One model handles the whole mixture instead of a chain of specialised tools. The same prompt interface applies whether the target is a voice, an instrument or a passing car.
How it works
Nothing to install and nothing to configure before the first run.
Upload an audio file, or a video whose soundtrack you want to work on. The picture is treated as a cue, not as decoration.
Type the target, click it in the frame, or drag over the seconds where it plays. Refine the prompt and run it again as often as you like.
Download the isolated target and the residual mix as separate tracks, ready to drop straight back into an editor.
Where it gets used
Most audio worth saving was recorded once, in a room nobody controlled.
Take a guest's voice away from the café behind them without flattening the room into a tunnel.
Recover usable dialogue from a location take, or lift an effect out of a mixed track when the original stems are gone.
Solo one instrument from a recording to learn the part, or mute it and play along with the rest of the take.
Pull one bird, one engine or one voice out of a landscape that was never actually quiet.
Feed a cleaner speech track into transcription so captions stop guessing at overlapped words.
Build single-source clips out of messy in-the-wild recordings instead of discarding them.
Inside a run
Separation is not subtraction. Each run returns the sound you prompted for together with the residual mixture it came from, so nothing is thrown away and either side can be the one you keep.
Target stem
The prompted sound, on its own track.
Residual mix
Everything the prompt did not select, still in sync.
Original
Untouched, for A/B against both outputs.
An illustrative view of a run. The working interface lives in the tool itself.
FAQ
What promptable separation does, and where it gets difficult.
Instead of choosing from a fixed menu of stems — vocals, drums, bass, other — you describe the sound you want in words, or point at it, and that is what gets separated. The set of possible targets is open rather than baked into the model's outputs.
Bring a recording, describe what to isolate, and take both sides away.
No plugin chain, no stem presets.