Skip to content
Euphona

Guide

What stem separation actually does

Stem separation estimates the sources inside a finished mix — vocals, drums, bass and the rest. It recovers what is present, never what was never there.

What it is

Stem separation takes a finished stereo mix and estimates the individual sources that went into it — typically vocals, drums, bass and everything else. You put in one file and get back several, each containing its best guess at one part of the recording.

The word “estimates” is doing real work in that sentence. When a mix was made, the separate tracks were summed into two channels, and summing is a lossy, irreversible operation: given only the sum, there are infinitely many combinations that could have produced it. A separator does not undo the summing. It uses what it has learned about how instruments look and behave to produce the most plausible decomposition, and plausible is not the same as original.

That distinction is the whole difference between what separation can do and what people hope it does. It is very good and it is not the multitrack.

How it works, roughly

Most modern separators work on a time-frequency representation — a spectrogram, which shows how much energy is present at each frequency at each moment. A neural network, trained on many thousands of songs where both the mix and the true stems were available, learns to predict which portion of the energy at each point belongs to each source.

It then applies that prediction as a kind of mask, keeping the parts of the mix it attributes to a source and suppressing the rest, before converting back to audio. Newer models work partly or entirely on the waveform directly, which helps with the phase information a spectrogram approach has to reconstruct.

What it has learned is essentially a very detailed sense of what a snare looks like when it is not alone, what a voice does across a phrase, and how bass energy is distributed. Nothing about the process involves recovering a file that still exists somewhere. There is no such file.

Two, four or five stems?

More stems is not better, and this is the single most common mistake people make with these tools.

  • Two stems — vocals and instrumental. The cleanest result available, because the model only has to make one decision per point in the spectrogram. If you want an instrumental or an a cappella, ask for this and nothing more — our free vocal remover is exactly this split, so you can hear the quality on your own track before deciding anything.
  • Four stems — vocals, drums, bass, other. The standard split, and the one worth using if you are rebuilding a mix. “Other” is a genuine catch-all: guitars, keys, strings and everything the model was not asked to identify.
  • Five stems — adds piano. Useful when the track has a prominent piano part, and actively worse when it does not, because the piano head still has to take its material from somewhere and it will find something to call a piano.

Every additional output is another chance for material to be assigned to the wrong place. Ask for the fewest stems that answer your question.

Where it struggles, and why

  • Reverb tails. A vocal drenched in reverb is acoustically part vocal and part room. The model has to decide whether the tail belongs to the vocal stem, and there is no correct answer — the reverb was printed onto the vocal, so both answers are defensible and neither matches what you imagined.
  • Heavy bus compression. When the mix was glued together with a compressor across the whole thing, every source modulates every other source’s level. The sources are no longer independent, and a separator assuming independence inherits that entanglement.
  • Dense low-mids. Bass, kick drum, low piano and the bottom of a guitar all occupy the same territory. This is exactly the region where a human mix engineer has to work hardest, and for the same reason.
  • Doubles and wide panning. A vocal doubled and panned hard left and right is two performances the model may treat as one source in an odd stereo position.
  • Anything unusual. Separators are trained on the distribution of music they were shown. An instrument that is rare in that distribution — a sitar, an unusual synthesiser, an untrained voice recorded badly — gets a worse estimate, because the model has less idea what it is looking at.

The artefacts, when they appear, are recognisable: a watery, phasing quality in the high end; brief burbling sounds where the mask changed abruptly between frames; and bleed, where a faint ghost of one source haunts another’s stem.

How to judge a separator

Do not test it on a song that separates easily. Sparse recordings with a clear vocal and a simple arrangement make every tool look excellent, which is exactly why demonstrations use them.

Test it on the hard cases, and listen for specific things:

  • Solo each stem and listen for what should not be there. Bleed is the most common failure and the easiest to hear in isolation.
  • Sum the stems back together and compare to the original. They should add up to something very close to what you started with. If material has vanished, the tool discarded whatever it could not classify.
  • Check the sample rate of what comes back. Some tools separate at 44.1 kHz internally regardless of what you gave them, then upsample. The result is band-limited in a way that is obvious on a spectrum analyser and subtle in the room — a hard shelf where the content simply stops is the giveaway.
  • Listen to the vocal’s consonants and the cymbals. High-frequency transients are where masking artefacts show up first.

What people actually use it for

  • Instrumentals and a cappellas for live performance, karaoke, or practice.
  • Remixing and sampling, where a usable drum or vocal part matters more than perfect fidelity.
  • Rescuing a mix when the session is gone and only the bounce survives. This is more common than it sounds — hard drives fail, and studios close.
  • Feeding other processing. Separation is the enabling step for spatialisation, which needs individual sources to place; for tuning a vocal that only exists inside a mix; and for treating one damaged element without touching the rest.
  • Study. Hearing how a record you love was actually balanced, one part at a time, teaches more than reading about it.

The part about rights

Separating a recording does not create any rights in the result. A vocal extracted from a commercial release is still that release’s vocal, owned by whoever owned it before you ran the tool, and using it in something you distribute needs the same permission it always did.

This is worth saying plainly because the technology makes it feel otherwise. Something you made with software on your own computer intuitively feels like yours. Copyright does not work that way, and the intuition has ended a number of releases.

Practising on a record, learning from it, or making something you never distribute is a different matter and generally fine. The line is publication. Our acceptable use policy says where we draw it, and rights holders can reach us directly.

Keep reading