| dc.description.abstract |
Speech enhancement (SE) aims to elevate the perceptual quality and intelligibil
ity of speech signals by mitigating ambient noise and distortions. Recent speech
enhancement methods struggle with preserving fine temporal details in noisy en
vironments, where low-resolution speech increases noise sensitivity. The difficulty
of conveying these details, crucial for understanding speech, is compounded by
the computational complexity of mapping-based approaches, which require learn
ing the full range of clean spectrogram values. However, this work proposes a
masking-based strategy to address these challenges in monaural speech.Two dis
tinct approaches are explored by employing the Ideal Ratio Mask (IRM) into
deep learning frameworks: a U-Net inspired architecture and a Time-Frequency
Transformer Network (TF-TransNet). The effectiveness of both approaches is
analyzed across different Signal-to-Noise Ratios (SNRs) using several objective
metrics: Short-Time Objective Intelligibility (STOI), Perceptual Evaluation of
Speech Quality (PESQ), Segmental Signal-to-Noise Ratio (SSNR) and Scale
Invariant Signal-to-Distortion Ratio (SI-SDR). In the first approach, the U-Net
model comprises an encoder for feature extraction, a decoder for reconstructing
the clean speech signal, and skip connections that enable the direct transfer of
key information between the encoder and decoder. The U-Net approach demon
strates significant improvements in speech intelligibility and quality under mod
erate to high SNRs and familiar noise types, although its performance declines
in low SNR conditions and with unseen noise types. To address these limita
tions, this work introduces the TF-TransNet, aiming to predict the IRM. This
network utilizes a Time-Frequency Attention (TFA) encoder to process the noisy
magnitude spectrogram, integrating the processed information with the decoder
through a Time-Frequency (T-F) Transformer layer. By employing Multi-Head
Self-Attention (MHSA) mechanisms and replacing the initial fully connected layer
with a Long Short-Term Memory Unit (LSTM) in the T-F transformer layer,
the model can dynamically prioritize informative features and their relationships
across both time and frequency domains, enhancing its ability to capture long
range dependencies in sequential data. The proposed TF-TransNet outperforms
existing models, such as the Convolutional Recurrent Network (CRN) and Gated
Convolutional Recurrent Network (GCRN), demonstrating enhanced speech in
telligibility and quality across both familiar and unfamiliar noise scenarios. |
en_US |