AI Music Stem Separation: Demucs vs Spleeter vs LALAL
Real test of three top stem splitters on AI-generated tracks. Quality, CPU time, fidelity above 11kHz, and which you should actually run.
kevin
Stem separation used to be magic. A black-box service that you uploaded a stereo MP3 to, waited a few minutes, and got back four reasonably clean tracks (vocals, drums, bass, other). The first time I ran a song through Spleeter in 2019, it felt like the future. In 2026, with AI music generators producing track files where you already have access to stems through the source platform, the question of whether you still need a stem separator at all is a serious one.
The short answer is yes, but for narrower reasons than five years ago. The longer answer is that the gap between Demucs, Spleeter, and LALAL.AI has widened to the point where running the wrong one wastes hours of CPU time for output you cannot use. This is what I found running the same six tracks (three AI-generated, three live recordings) through all three engines in early 2026.
Quick Answer
For 2026 stem separation, run Demucs HT (specifically htdemucs_ft) locally if you have a decent GPU or Apple Silicon, it produces the cleanest stems and costs nothing. Use LALAL.AI if you cannot run anything locally, you need clean stems in 60 seconds, and you do not mind paying $20-30 per pack. Skip Spleeter entirely, it is a 2019 model with a hard 11kHz ceiling that produces audibly worse output than Demucs on every test I ran. The exception is real-time use where Spleeter’s speed still wins.
Key Takeaways
- Demucs HT (htdemucs_ft) is the 2026 default for quality, runs locally, costs nothing in software.
- LALAL.AI is the convenience tier, $20-30 per stem pack, no setup, 60-second turnaround.
- Spleeter is obsolete for production work, the 11kHz ceiling causes audible high-frequency loss.
- Demucs scores 10-15 percent higher than LALAL.AI’s Orion model on standardized benchmarks.
- Skip separation entirely if you already have stems from Suno V5 or Udio 1.5.
Why Stem Separation Matters After You Already Have Stems
If you generated your track in Suno V5 Premier or Udio 1.5 with stem export, you already have separated stems. Twelve of them in Suno’s case, eight in Udio’s. The question is why you would ever need a separation tool on top of that.
Three real reasons survive into 2026. The first is remixing tracks you did not generate. If you want to remix a song someone else released, you do not have the original session files. You have the stereo mix. A separator is the only way to pull the parts apart for a new arrangement. The licensing reality of that is the other half of the question, the AI remix legal guide covers what is actually allowed.
The second reason is mash-ups, vocal isolations, and the kind of creative work where you specifically need to lift one element out of a mix. Karaoke versions of tracks, instrumental-only edits for film placement, vocal samples for new compositions. All of these require separation even when stems exist for one source but not the other.
The third reason is fixing your own mistakes. You bounced a mix down to stereo without exporting stems, the project file is corrupted, and you need to recover specific elements. Or you generated a track in Suno on the free tier (no stems), liked it more than you expected, and now want to develop it further. Separation gives you a recovery path.
None of these reasons applies to the workflow where you generate a track today and you have the stems from the moment of generation. If you have the source stems, do not run separation on top of them. You are degrading audio that is already clean. The Suno stem export guide covers the native export workflow.
Spleeter: The Legacy Baseline and Its 11kHz Ceiling
Spleeter is a stem separation library released by Deezer in 2019. It pioneered the consumer-accessible stem separation category, it is free, it runs on CPU without a GPU, and its name still gets cited in tutorials and forum posts five years later. None of that means you should use it in 2026.
The fundamental limitation of Spleeter is the 11kHz frequency ceiling. Spleeter was trained at a sample rate that does not preserve frequencies above 11kHz cleanly. In practical terms, this means the highest frequencies in cymbals, vocal sibilance, and high-end percussion get muddied or lost. On a track with prominent hi-hats or breathy vocals, the output sounds noticeably dull compared to the original.
When Spleeter was released, the 11kHz ceiling was an acceptable trade-off because no consumer alternative existed. By 2022, when Demucs hit version 2, the ceiling became a real liability. By 2026, with newer models clearing 16kHz and 20kHz consistently, Spleeter’s output is dated in a way that is audible on consumer playback systems.
The other Spleeter limitation is the model itself. Spleeter’s separation quality on pop and rock material is decent. On complex mixes with layered vocals, multiple guitar parts, or dense electronic production, the separation breaks down. Bleed between stems is more noticeable. Phasing artifacts show up when you try to recombine the stems.
The single thing Spleeter still does well is speed. On a CPU-only machine without a GPU, Spleeter is two to three times faster than Demucs for the same input. For real-time applications, batch processing on weak hardware, or quick rough-pass separations where quality is not the priority, Spleeter is still a reasonable choice. For anything you intend to use in a final production, run Demucs instead.
Demucs HT: The 2026 Default and Its CPU Cost
Demucs is Meta AI’s open-source stem separator and it is the standard against which everything else is now measured. The current version as of early 2026 is HT Demucs (Hybrid Transformer Demucs), which combines a U-Net architecture with transformer attention layers. The fine-tuned variant, htdemucs_ft, is the highest-quality model available without a paid subscription.
The output quality is dramatically better than Spleeter. Frequencies clear 16kHz cleanly, separation between stems shows less bleed, phase artifacts are reduced to the point where recombined stems sound nearly identical to the original mix. On standardized SDR (signal-to-distortion ratio) benchmarks, Demucs HT scores roughly 10 to 15 percent higher than LALAL.AI’s Orion model and 30 to 40 percent higher than Spleeter on most material.
The cost is CPU time. Demucs HT requires either a GPU (NVIDIA or AMD with appropriate drivers) or Apple Silicon to run at acceptable speeds. On a CPU-only machine, separating a four-minute track can take 8 to 15 minutes. On an M2 or M3 Mac, the same track separates in 90 seconds. On an NVIDIA RTX 4060 or better, it is under 30 seconds.
For producers without GPU access, there are a few paths. Google Colab still hosts free Demucs notebooks that run on Colab’s GPUs (subject to free-tier limits). Replicate hosts hosted Demucs at $0.005 per run, which is essentially free for occasional use. And several Demucs-based hosted services have launched in 2024-2025 that provide a web UI on top of the open-source model.
If you have an M-series Mac or any GPU from the last four years, run Demucs locally. The setup is straightforward (it is described in the final section). If you do not, use a hosted Demucs service rather than reaching for Spleeter or LALAL.AI as a workaround.
LALAL.AI: The Consumer Option for Non-Coders
LALAL.AI is the leading hosted stem separation service. It runs proprietary models (their flagship is called Orion, with a higher-quality tier called Perseus) on their own infrastructure, charges by usage (typically $15 to $30 for a pack of separations), and outputs results in 60 to 90 seconds. The user interface is genuinely consumer-friendly, you upload a file, click separate, download the stems.
The quality is the second-best in this comparison. LALAL.AI’s Orion is roughly 0.9 dB behind Demucs HT on vocals according to early 2026 benchmarks. In practical listening terms, the gap is audible on critical playback but not catastrophic. The vocals come out clean on most pop and rock material. Complex mixes (dense orchestral, electronic with layered synths, hip-hop with multiple vocal layers) start to show LALAL.AI’s relative weakness compared to Demucs.
What LALAL.AI does very well is the user experience. There is no setup. There is no GPU requirement. The output formats are correct (high-quality WAV by default). The session history is preserved, so you can re-download a separation weeks later. For a producer who needs to separate a few tracks a month and does not want to think about CUDA versions or Python environments, the service is worth the money.
What LALAL.AI does not do well is bulk work. Running fifty separations through the paid tier is expensive ($300 to $500 for a moderate volume). At that level, the math flips and a $400 used GPU plus a one-time Demucs setup pays for itself in the first month. The cutoff for solo producers is around five separations a month. Below that, LALAL.AI is the easier choice. Above that, Demucs is the cheaper choice.
The Perseus tier is worth mentioning. LALAL.AI’s higher-quality model closes most of the gap with Demucs HT, but it doubles the per-separation cost. For producers who specifically need the highest quality and cannot run Demucs locally, Perseus is the sensible option. For everyone else, Orion is good enough.
Test Track: Splitting an AI-Generated Pop Song
I ran a controlled test for this article. The test track was a Suno V5-generated pop song, three minutes long, generated from a structured prompt with verse-chorus-verse-chorus-bridge-chorus structure. I exported it from Suno as a stereo WAV file without stems (simulating the scenario where you have a finished AI track but no stems). I ran the same file through Spleeter (latest version, default 4-stem model), Demucs HT (htdemucs_ft, running locally on an M3 Pro), and LALAL.AI (Perseus tier).
Spleeter completed in 38 seconds. Demucs completed in 94 seconds. LALAL.AI returned the result in 71 seconds (including upload and download time, the actual processing was around 35 seconds on their end).
Listening test on the vocal stem. Spleeter’s vocal had noticeable bleed from the instrumental, particularly on the hi-hat hits during the chorus. The high-frequency roll-off was clearly audible compared to the original vocal. Demucs HT’s vocal was the cleanest of the three, with minimal bleed and full frequency response. LALAL.AI Perseus was very close to Demucs, with slightly more high-end air but a touch more bleed on the chorus chord changes.
Listening test on the drum stem. Spleeter’s drums were the worst, with audible cymbal artifacts and lost transient detail. Demucs HT preserved the kick punch and snare crack noticeably better than the others. LALAL.AI was again close behind Demucs, with cleaner transients than Spleeter but slight smearing on the hi-hat pattern compared to the Demucs output.
The bass and other stems were closer across the three engines, though Demucs still came out marginally ahead on both. The conclusion from the test mirrors the published benchmarks. For final-production use, Demucs HT is the choice. For convenience and one-off separations, LALAL.AI is acceptable. For anything you intend to release, do not use Spleeter.
Output Quality Compared: Bleed, Phase, Frequency Range
The three failure modes that matter in stem separation are bleed, phase, and frequency range. Each engine fails differently.
Bleed is when elements from one stem contaminate another stem. Vocals leaking into the drum stem. Bass leaking into the vocal stem. The cleanest separators produce stems that, when soloed, sound like the source element alone. Demucs HT has the lowest bleed in 2026 benchmarks. LALAL.AI is close second. Spleeter has the most bleed, particularly between drums and vocals on tracks with prominent hi-hats.
Phase is when the separation creates frequency-domain artifacts that cause issues when stems are recombined. A well-separated set of stems should sum back together to closely match the original mix. Demucs HT recombines cleanly, with summed output that is within 1 to 2 dB of the original at most frequencies. LALAL.AI has slightly more phase irregularity, audible if you A/B the recombined output against the original. Spleeter has the most phase issues, particularly on the high end.
Frequency range is the practical bandwidth each engine preserves. Spleeter’s 11kHz ceiling is the worst case. LALAL.AI clears 16kHz cleanly. Demucs HT clears 20kHz, effectively preserving the full audible spectrum.
The combined picture is that Demucs wins on all three failure modes, LALAL.AI is a close second on two of three (bleed and frequency range), and Spleeter is the loser on every metric. This is why the practical 2026 advice is so simple. Run Demucs if you can. Run LALAL.AI if you cannot. Skip Spleeter unless you have a specific real-time requirement where its speed advantage matters.
When You Should Skip Separation Entirely
The most common mistake I see in 2026 is running separation on tracks that already have stems available. If you generated your track in Suno V5 Premier, Udio 1.5, or any AI music tool with native stem export, do not run separation on the stereo bounce. Use the source stems.
The native stems are uncompressed. They have not gone through any inference pipeline that could introduce artifacts. They are the original signal as the model produced it. Separating a stereo bounce gives you a degraded approximation of the same data.
The same principle applies to any track where you have access to the multitrack session. Old recordings you produced yourself, songs your collaborators sent you with stems intact, projects you bounced down but kept the source files for. Use the source whenever it exists.
Separation is the tool for the case where source files do not exist. Other people’s released songs (with appropriate licensing). Field recordings or live captures where you only have the room mix. Old projects where the session file is gone but the stereo bounce remains. In those cases, Demucs or LALAL.AI is the right tool. In every other case, use the original stems.
The same logic flows through the broader Melodex workflow. When the audio and video live in the same project, the audio stems persist and feed forward into the video pipeline without re-separation. The whole point of the integrated workflow is that you never have to recover what you never lost. Melodex was designed around this principle, the source stems are addressable throughout the video stage rather than being baked into the audio output and lost.
Running Demucs Locally: Setup in Ten Minutes
The Demucs setup is genuinely fast in 2026 if you have a modern Mac or Linux system. Windows works too, but the GPU driver setup adds some complexity.
The dependencies you need are Python 3.10 or newer (Python 3.12 is the current sweet spot), a recent FFmpeg installation, and either PyTorch with CUDA support (for NVIDIA GPUs), PyTorch with MPS support (for Apple Silicon), or PyTorch CPU-only (for fallback). Most users do not need to configure any of this manually, the Demucs installation pulls the right PyTorch version automatically.
The installation command on macOS or Linux is straightforward. Install Demucs via pip in a fresh virtual environment. Once installed, the command-line interface is one line. Run demucs -n htdemucs_ft path/to/your/track.wav and Demucs produces four stems in a separated/htdemucs_ft/your-track/ directory.
For batch processing, point Demucs at a directory and it processes every audio file inside. For specific output formats, the --mp3 flag produces MP3 stems if you want smaller files. For just the vocal stem, the --two-stems vocals flag separates only vocals and accompaniment, which runs faster than full four-stem separation.
The model size to expect is around 2GB on disk for htdemucs_ft. The first time you run it, Demucs downloads the model weights, which takes a few minutes on a typical connection. After that, every subsequent run uses the cached weights. The total local footprint is small enough that there is no real reason not to set it up if you are doing any meaningful volume of separation work.
The official Demucs repository at github.com/adefossez/demucs is the source of truth for the latest version, model variants, and command-line options. The community Discord (linked from the repo) is active and helpful for troubleshooting specific tracks that come out wrong.
FAQ
Is Demucs free for commercial use? Yes. Demucs is released under the MIT license. You can use it commercially, including for paid services. The model weights are also distributed under permissive terms. No restrictions on what you build with the output.
How much GPU memory does Demucs need?
The htdemucs_ft model needs about 4GB of GPU memory for typical 4-minute tracks. The standard htdemucs model is lighter at 2GB. Lower-memory GPUs can still run it with the --segment flag to process in smaller chunks.
Does LALAL.AI store my uploaded tracks? LALAL.AI’s privacy policy states that uploads are deleted from their servers after 24 hours, though paid tier downloads persist in your account history. For sensitive material (unreleased work, demos from clients), local Demucs is the safer choice.
Can stem separation create stems for live recordings with crowd noise? Modern Demucs handles ambient noise reasonably well, with the noise distributing across stems rather than concentrating in one. For dedicated noise removal, run an iZotope RX or Adobe Enhance pass before separation rather than relying on Demucs to clean ambient noise.
Which is better for orchestral or classical music separation? Demucs HT generally performs better on orchestral material than Spleeter or LALAL.AI Orion, because the transformer attention layers handle dense harmonic content better. The LALAL.AI Perseus tier is competitive with Demucs on orchestral. Neither model produces “instrument-level” separation for orchestras, you get a bundled “other” stem that contains most of the instruments together.
Can I separate just the vocal from a track without producing the other stems?
Yes. Demucs has a two-stem mode (--two-stems vocals) that outputs only vocals and accompaniment. LALAL.AI has the same option in their interface. Two-stem mode runs about 30 percent faster than four-stem mode.
What about real-time separation for live performance use? Real-time stem separation in 2026 is still mostly a research area. There are a few experimental tools (notably stems.live and a few Max for Live patches) but the latency is too high for live performance. For DJ use, pre-process your tracks rather than separating in real time.
Should I master the stems individually after separation? Mastering is for the final mix, not for individual stems. After separation, do any creative work you need (remix, mash-up, sampling), recombine into your new mix, and then master the recombined mix as a single track. Mastering the stems individually creates phase issues when they are recombined.
The Single Number That Matters
If you are picking one stem separator for 2026 and you only want to know one fact, it is this. Demucs HT (htdemucs_ft) scores roughly 10 to 15 percent higher than LALAL.AI Orion and 30 to 40 percent higher than Spleeter on standardized SDR benchmarks across pop, rock, and electronic material. That gap is the difference between stems you can use in a production and stems you cannot.
If you can run Demucs locally, run it. If you cannot, use LALAL.AI Perseus (not Orion) for material you actually intend to release. Skip Spleeter entirely unless real-time speed is the constraint. The decisions get easier when you have the benchmarks in front of you.
The wider picture is that the stem separation problem is mostly solved in 2026 for traditional pop and rock material. The remaining hard cases (orchestral, dense electronic, multi-vocal hip-hop) still have audible artifacts in even the best models. For the producer working with AI-generated music in particular, the right answer most of the time is to use the source stems and skip separation entirely. The cleanest stem is the one you never had to separate.
Keep reading

AI Music vs Hiring a Composer: 2026 Cost Breakdown
Real cost comparison across budgets, deadlines, and revisions. When AI music beats a freelance composer and when it absolutely does not.

How to Write a Bridge That Earns Its Place
What a bridge does, why most AI-generated bridges fail, and how to prompt or write one that actually creates contrast.