SelfSE: Self-Supervised Speech Enhancement via Noisy Speech Refinement Online Supplement

Authors

Wenbin Jiang, Fei Wen, Kai Yu

Abstract

The majority of deep learning-based speech enhancement techniques rely on supervised training, which requires extensive pairs of noisy and clean speech. However, due to the complexity and variability of real-world environments, obtaining ground-truth clean speech can be challenging or even impractical in certain situations. To tackle this issue, we introduce a novel self-supervised speech enhancement method that eliminates the need for clean reference speech. Our method involves two training stages to develop a speech enhancement model through iterative refinement. In the first stage, we use unprocessed noisy speech and noise to create noisier-to-noisy data pairs, which are used to train the initial model. In the second stage, we iteratively generate noisier-to-noisy data pairs using speech sampled from the estimated speech and unprocessed noisy speech, along with noise sampled from the noise corpus, to train a more effective model. Further, we provide analysis to compare and deepen the understanding of various self-supervised learning methods, including NyTT, IDR-SE, and the proposed SelfSE method. Particularly, we show that noisy-target based self-supervised learning inherently introduces a bias, and this bias is SNR-dependent that it increases as the SNR decreases. We conduct extensive experiments on three popular benchmark datasets, and the results demonstrate that our approach achieves performance comparable to supervised learning methods on simulated data and surpasses them on real-world data in both speech enhancement and recognition tasks.

Datasets

  • The VoiceBank+DEMAND dataset is used for demo.
  • Audio samples of the test set we processed are available at the repository (voicebank).

  • Setups

  • The neural network architecture is defined in model.py, with detailed configurations provided in model_arch.py.
  • The pre-trained model and experiment configurations can be found in the Example Directory (examples/voicebank).

  • Compared methods

  • OMLSA: Noise Spectrum Estimation in Adverse Environments: Improved Minima Controlled Recursive Averaging
  • DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement
  • Noise2Noise: Speech Denoising without Clean Training Data: a Noise2Noise Approach
  • Noiser2Noisy: Noisy-target Training: A Training Strategy for DNN-based Speech Enhancement without Clean Speech
  • IDR-SE: Iterative Noisy-Target Approach: Speech Enhancement Without Clean Speech

  • Audio Samples

    Model\id(noise) p257_008(cafe) p257_033(living) p257_106(bus) p232_250(office) p232_409(psquare)
    Clean
    Noisy
    OMLSA
    DCCRN
    Noise2Noise
    Noiser2Noisy
    IDR-SE
    SelfSE

    Spectrogram of the samples in third column