CUET DIGITAL REPOSITORY

Improved Speech Enhancement Through Attention-Driven T-F Masking Strategies

Show simple item record

dc.contributor.author Akter, Khadija
dc.contributor.author ID:, 19METE012P
dc.date.accessioned 2026-10-04T06:18:23Z
dc.date.available 2026-10-04T06:18:23Z
dc.date.issued 2024-12-12
dc.identifier.uri http://103.99.128.19:8080/xmlui/handle/123456789/596
dc.description A Master of Engineering (M.Engg) Thesis in Electronics and Telecommunication Engineering (ETE) Department at Chittagong University of Engineering and Technology (CUET). en_US
dc.description.abstract Speech enhancement (SE) aims to elevate the perceptual quality and intelligibil ity of speech signals by mitigating ambient noise and distortions. Recent speech enhancement methods struggle with preserving fine temporal details in noisy en vironments, where low-resolution speech increases noise sensitivity. The difficulty of conveying these details, crucial for understanding speech, is compounded by the computational complexity of mapping-based approaches, which require learn ing the full range of clean spectrogram values. However, this work proposes a masking-based strategy to address these challenges in monaural speech.Two dis tinct approaches are explored by employing the Ideal Ratio Mask (IRM) into deep learning frameworks: a U-Net inspired architecture and a Time-Frequency Transformer Network (TF-TransNet). The effectiveness of both approaches is analyzed across different Signal-to-Noise Ratios (SNRs) using several objective metrics: Short-Time Objective Intelligibility (STOI), Perceptual Evaluation of Speech Quality (PESQ), Segmental Signal-to-Noise Ratio (SSNR) and Scale Invariant Signal-to-Distortion Ratio (SI-SDR). In the first approach, the U-Net model comprises an encoder for feature extraction, a decoder for reconstructing the clean speech signal, and skip connections that enable the direct transfer of key information between the encoder and decoder. The U-Net approach demon strates significant improvements in speech intelligibility and quality under mod erate to high SNRs and familiar noise types, although its performance declines in low SNR conditions and with unseen noise types. To address these limita tions, this work introduces the TF-TransNet, aiming to predict the IRM. This network utilizes a Time-Frequency Attention (TFA) encoder to process the noisy magnitude spectrogram, integrating the processed information with the decoder through a Time-Frequency (T-F) Transformer layer. By employing Multi-Head Self-Attention (MHSA) mechanisms and replacing the initial fully connected layer with a Long Short-Term Memory Unit (LSTM) in the T-F transformer layer, the model can dynamically prioritize informative features and their relationships across both time and frequency domains, enhancing its ability to capture long range dependencies in sequential data. The proposed TF-TransNet outperforms existing models, such as the Convolutional Recurrent Network (CRN) and Gated Convolutional Recurrent Network (GCRN), demonstrating enhanced speech in telligibility and quality across both familiar and unfamiliar noise scenarios. en_US
dc.description.sponsorship N/A en_US
dc.language.iso en en_US
dc.publisher CUET en_US
dc.relation.ispartofseries ;TCD-94
dc.subject Speech Enhancement en_US
dc.subject Masking en_US
dc.subject U-Net en_US
dc.subject Transformer en_US
dc.subject Speech In telligibility en_US
dc.subject Speech Quality en_US
dc.title Improved Speech Enhancement Through Attention-Driven T-F Masking Strategies en_US
dc.type Thesis en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search DSpace


Advanced Search

Browse

My Account