ComparisonPublished September 5, 2026

Speech Denoising Models Compared by Ear

Noise removal models are usually compared with numbers. Scores are useful, but they do not tell you what a model sounds like when a car passes, or how much of the speaker survives. This post does the opposite: ten speech denoisers, one recording, no scores, just audio.

The test clip

The input is 28 seconds of speech recorded against loud street traffic. Every model received the exact same file. We did not trim it, change its loudness, or run anything after the model. Each model worked at its native sample rate, from 16 kHz for most of them up to 48 kHz for the full-band ones. Files are stored as they came out of each network.

The recording was made from a clean voice track, so we also know what the speech sounds like with no noise at all. That clean reference is included below. It helps you separate two questions: how much noise is left, and how much the model changed the voice.

The models

Short notes on each network before you listen. Order is roughly by weight, from the smallest real-time models to the bigger offline ones. One of them is ours; we say so where it appears.

1. Fullband Standardours

The default model behind our own service. It works on full-band 48 kHz audio and runs on a server, so the browser never has to carry the model. We include it here for the same reason everyone else is included: so you can judge it against the open alternatives. Full disclosure: this is our product, and we did not tune anything for this clip.

2. RNNoiseMozilla, 2018

A small C library that mixes classic speech processing with a tiny gated recurrent network. It was designed to run inside real-time VoIP software on weak hardware, and its footprint is a fraction of a percent of most deep models here. On steady noise it holds up well for its size.

3. NSNet2Microsoft, DNS challenge baseline

The official baseline of Microsoft's Deep Noise Suppression challenge. A causal recurrent network that processes the signal as it arrives, frame by frame, which makes it simple to run on a CPU in real time. It was the yardstick that later models were measured against.

4. DTLNInterspeech 2020

Two stacked networks, one working on a short-time Fourier transform and one on a learned time-domain representation, both driven by small LSTMs. The whole model stays under a million parameters. DTLN became one of the most reused open baselines after the DNS challenge.

5. GTCRNXiaobin Rong et al., ICASSP 2024

Built for devices with almost no compute budget. It groups the frequency bands, runs a lightweight convolutional recurrent block over them, and adds a gammatone-style filterbank front end. The authors report about 48 thousand parameters in total.

6. DPCRNLe et al., 2023

A dual-path design: one path works inside short time frames, the other across frames, and the two share information. The causal checkpoint we used came from the DNS 3 challenge and reduces noise with a magnitude mask.

7. aTENNuateBrainChip, 2024

A recent raw-waveform model from BrainChip that combines convolutional blocks with attention and state-space layers. It is meant to be configurable for real-time use and is distributed with pre-trained 16 kHz weights for denoising.

8. FullSubNetHao et al., ICASSP 2021

A fusion model. A full-band network looks at the whole spectrum and hands useful context to a set of sub-band networks, one per frequency. It was one of the first models to make that combination practical for single-channel speech.

9. CleanUNetNVIDIA, ICASSP 2022

A U-Net that works directly on the waveform and inserts self-attention blocks between the encoder and decoder. It is larger than the real-time models here and is not built for streaming, which gives it more freedom to look at the signal.

10. Demucs denoiserMeta, 2020

The speech enhancement version of Demucs, an encoder-decoder network known from music separation. We used the dns48 checkpoint, one of the smallest of the three released variants. An older model, but still a common reference point in comparisons.

The audio

All players are stacked so you can move between them without scrolling the descriptions. Start with the two references, then listen to each model in the same order as above.

0a. Baseline: noisy original

0b. Clean reference: voice only

1. Fullband Standard result

2. RNNoise result

3. NSNet2 result

4. DTLN result

5. GTCRN result

6. DPCRN result

7. aTENNuate result

8. FullSubNet result

9. CleanUNet result

10. Demucs denoiser result

What to listen for

Notice three things per model. How much traffic disappears. How natural the voice stays, in tone and in level. And what appears in place of the noise: musical tones, a watery wobble, or nothing at all. The lightweight models are allowed to leave some traffic behind, because their job is usually to keep latency low on a phone call. The big models have fewer excuses.

We are not printing a ranking. A ranking would only be true for this one clip, this one voice and this one loudness, and it would age badly the day a model is retrained. The players above are the data. Draw your own conclusion, then test the candidate on your own recording before you trust it.

Try it on your own recording

A comparison like this only means something for your material. Your noise, your voice, your microphone. Fullband runs the same class of model on files you upload, and new accounts start with 50 free tokens.

Each model was run from its official repository and released weights. Fullband Standard is our own fullband model, included with disclosure. Details of every run, including sample rates, are kept in the Fullband engineering notes.