| dc.description.abstract |
Speech Bandwidth Extension (SBE) remains a challenging task in speech pro
cessing. It involves the intricate process of estimating missing frequency com
ponents in low-resolution signals to reconstruct high-resolution transformations.
SBE is a technique used to enhance the quality of speech signals by expanding
the frequency range of the audio. This endeavor is fundamental for enhancing
speech quality, naturalness, and intelligibility, particularly in scenarios marred
by background noise and channel distortions. Low-resolution signals often lack
high-frequency components due to limited sampling rates and band-limiting, as
dictated by the Nyquist-Shannon Sampling Theorem. The inherent loss of high
frequency details stems from aliasing, quantization errors, and insufficient granu
larity in sampling, which poses significant challenges in restoring full-band speech.
Addressing these issues is critical to advancing SBE techniques. Traditional
methodologies heavily relied on statistical algorithms and learning techniques
such as Multi-Layer Perceptrons (MLP). However, the paradigm shift brought
about by Deep Learning has revolutionized this domain, opening avenues for
more efficient and effective solutions. This thesis delves into deep learning-based
approaches for addressing the complexities of BWE and speech restoration, par
ticularly emphasizing noisy environments. This research explores three distinct
augmentations of a frequency-domain deep learning network, each tailored to
tackle specific challenges encountered in bandwidth extension and speech en
hancement. The first approach introduces a joint bandwidth expansion and
speech enhancement paradigm utilizing Deep Neural Networks (DNNs). This
approach is meticulously designed to simultaneously expand the bandwidth of
speech signals, reduce noise interference, and maintain the quality and intelligibil
ity of the speech. Leveraging the inherent capabilities of DNNs, this methodology
accurately estimates missing speech components. It characterizes noise profiles
within degraded signals, synthesizing high-fidelity full-band speech from limited
bandwidth inputs. The experimentation demonstrates the superior performance
of this DNN-based approach, surpassing conventional methods and presenting
promising avenues for real-world applications. Stepping beyond conventional
methodologies, the second approach introduces an end-to-end frequency-domain
framework, aptly named the Robust extension-plus-enhancement of speech utiliz
ing Dual-former Network (RDNet). RDNet represents a paradigm shift in BWE,
aiming to recover full-band speech from noisy low-band signals directly. By intri
cately integrating speech enhancement and ideal bandwidth extension modules
ii
0.0–
Abstract
within a unified framework, RDNet demonstrates remarkable performance across
diverse noisy environments. Leveraging short-time Fourier transform (STFT) fea
tures for enhancement, RDNet achieves significant improvements in critical met
rics such as Short-Time Objective Intelligibility (STOI), Perceptual Evaluation
of Speech Quality (PESQ), and log-spectral distortion (LSD), underscoring its
efficacy and potential for practical deployment. Finally, this study’s investigation
culminates in introducing the conformer-motivated super-denoised network (CS
DNet), a novel approach tailored to mitigate the mismatch problem inherent in
domain-specific encoders/decoders. By embedding a conv Unet with T-F trans
formation layers, CSDNet reduces dependency on training data and outperforms
recent baseline methods across various objective metrics. Moreover, CSDNet ex
hibits favorable subjective performance in comparative studies, reaffirming its
suitability for real-time applications. Through these meticulously crafted ap
proaches, this thesis aims to advance the frontier of deep learning-based speech
bandwidth extension and speech restoration. By addressing the core challenges
of high-frequency component loss in low-resolution signals and proposing inno
vative solutions, this work offers insights and methodologies, proposes practical
solutions for tackling real-life problems, and contributes to the wider conversa
tion in the fields of speech processing and audio engineering |
en_US |