SelfSE: Self-Supervised Speech Enhancement via Noisy Speech Refinement Online Supplement
Authors
Wenbin Jiang, Fei Wen, Kai Yu
Abstract
The majority of deep learning-based speech enhancement techniques rely on supervised training, which requires extensive pairs of noisy and clean speech. However, due to the complexity and variability of real-world environments, obtaining ground-truth clean speech can be challenging or even impractical in certain situations. To tackle this issue, we introduce a novel self-supervised speech enhancement method that eliminates the need for clean reference speech. Our method involves two training stages to develop a speech enhancement model through iterative refinement. In the first stage, we use unprocessed noisy speech and noise to create noisier-to-noisy data pairs, which are used to train the initial model. In the second stage, we iteratively generate noisier-to-noisy data pairs using speech sampled from the estimated speech and unprocessed noisy speech, along with noise sampled from the noise corpus, to train a more effective model. Further, we provide analysis to compare and deepen the understanding of various self-supervised learning methods, including NyTT, IDR-SE, and the proposed SelfSE method. Particularly, we show that noisy-target based self-supervised learning inherently introduces a bias, and this bias is SNR-dependent that it increases as the SNR decreases. We conduct extensive experiments on three popular benchmark datasets, and the results demonstrate that our approach achieves performance comparable to supervised learning methods on simulated data and surpasses them on real-world data in both speech enhancement and recognition tasks.