<?xml version="1.0" encoding="UTF-8"?><rdf:RDF xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel rdf:about="http://103.99.128.19:8080/xmlui/handle/123456789/76">
<title>Thesis in ETE</title>
<link>http://103.99.128.19:8080/xmlui/handle/123456789/76</link>
<description>Thesis published in Dept. of ETE</description>
<items>
<rdf:Seq>
<rdf:li rdf:resource="http://103.99.128.19:8080/xmlui/handle/123456789/596"/>
<rdf:li rdf:resource="http://103.99.128.19:8080/xmlui/handle/123456789/581"/>
<rdf:li rdf:resource="http://103.99.128.19:8080/xmlui/handle/123456789/574"/>
<rdf:li rdf:resource="http://103.99.128.19:8080/xmlui/handle/123456789/564"/>
</rdf:Seq>
</items>
<dc:date>2026-10-04T07:31:18Z</dc:date>
</channel>
<item rdf:about="http://103.99.128.19:8080/xmlui/handle/123456789/596">
<title>Improved Speech Enhancement Through Attention-Driven T-F Masking Strategies</title>
<link>http://103.99.128.19:8080/xmlui/handle/123456789/596</link>
<description>Improved Speech Enhancement Through Attention-Driven T-F Masking Strategies
Akter, Khadija; ID:, 19METE012P
Speech enhancement (SE) aims to elevate the perceptual quality and intelligibil&#13;
ity of speech signals by mitigating ambient noise and distortions. Recent speech&#13;
enhancement methods struggle with preserving fine temporal details in noisy en&#13;
vironments, where low-resolution speech increases noise sensitivity. The difficulty&#13;
of conveying these details, crucial for understanding speech, is compounded by&#13;
the computational complexity of mapping-based approaches, which require learn&#13;
ing the full range of clean spectrogram values. However, this work proposes a&#13;
masking-based strategy to address these challenges in monaural speech.Two dis&#13;
tinct approaches are explored by employing the Ideal Ratio Mask (IRM) into&#13;
deep learning frameworks: a U-Net inspired architecture and a Time-Frequency&#13;
Transformer Network (TF-TransNet). The effectiveness of both approaches is&#13;
analyzed across different Signal-to-Noise Ratios (SNRs) using several objective&#13;
metrics: Short-Time Objective Intelligibility (STOI), Perceptual Evaluation of&#13;
Speech Quality (PESQ), Segmental Signal-to-Noise Ratio (SSNR) and Scale&#13;
Invariant Signal-to-Distortion Ratio (SI-SDR). In the first approach, the U-Net&#13;
model comprises an encoder for feature extraction, a decoder for reconstructing&#13;
the clean speech signal, and skip connections that enable the direct transfer of&#13;
key information between the encoder and decoder. The U-Net approach demon&#13;
strates significant improvements in speech intelligibility and quality under mod&#13;
erate to high SNRs and familiar noise types, although its performance declines&#13;
in low SNR conditions and with unseen noise types. To address these limita&#13;
tions, this work introduces the TF-TransNet, aiming to predict the IRM. This&#13;
network utilizes a Time-Frequency Attention (TFA) encoder to process the noisy&#13;
magnitude spectrogram, integrating the processed information with the decoder&#13;
through a Time-Frequency (T-F) Transformer layer. By employing Multi-Head&#13;
Self-Attention (MHSA) mechanisms and replacing the initial fully connected layer&#13;
with a Long Short-Term Memory Unit (LSTM) in the T-F transformer layer,&#13;
the model can dynamically prioritize informative features and their relationships&#13;
across both time and frequency domains, enhancing its ability to capture long&#13;
range dependencies in sequential data. The proposed TF-TransNet outperforms&#13;
existing models, such as the Convolutional Recurrent Network (CRN) and Gated&#13;
Convolutional Recurrent Network (GCRN), demonstrating enhanced speech in&#13;
telligibility and quality across both familiar and unfamiliar noise scenarios.
A Master of Engineering (M.Engg) Thesis in Electronics and Telecommunication Engineering (ETE) Department at Chittagong University of Engineering and Technology (CUET).
</description>
<dc:date>2024-12-12T00:00:00Z</dc:date>
</item>
<item rdf:about="http://103.99.128.19:8080/xmlui/handle/123456789/581">
<title>Machine Learning Assisted Beam-Steering Microstrip Patch Array Antenna Design</title>
<link>http://103.99.128.19:8080/xmlui/handle/123456789/581</link>
<description>Machine Learning Assisted Beam-Steering Microstrip Patch Array Antenna Design
Hossain, Md. Farhad; ID:, 19METE025P
The rapid evolution of wireless communication technologies, notably the transi&#13;
tion from 4G to 5G, has driven the need for advanced antenna systems to meet&#13;
growing demands for seamless connectivity and enhanced performance. This the&#13;
sis focuses on designing and optimizing beam-steering microstrip patch antennas&#13;
for 5G communication, with specific emphasis on the sub-6 GHz frequency band&#13;
and millimeter-wave applications. A comprehensive methodology is presented, in&#13;
tegrating traditional antenna design principles with machine learning-based pre&#13;
dictive modeling and optimization techniques. The initial phases of the research&#13;
involve designing a 2 × 2 microstrip array antenna, simulated for beam-steering&#13;
applications at 3.5 GHz. A dataset is curated through parametric studies, and&#13;
eight machine learning models, including Random Forest, Decision Tree Regres&#13;
sion, XGBoost, and Gaussian Process Regression, have been trained to predict&#13;
antenna performance metrics such as return loss (S11) and bandwidth. Among&#13;
these, Decision Tree Regression emerged as the best overall performer, offering&#13;
the most accurate predictions, while Random Forest also demonstrated close&#13;
alignment with the actual values. Optimization strategies refined the antenna&#13;
design by combining machine learning predictions with advanced algorithms like&#13;
L-BFGS-B, Genetic Algorithms, and Simulated Annealing. The optimized S11&#13;
parameter improved significantly, reaching at-51.41 dB compared to the initial-45.21 dB, indicating the superior impedance matching at 3.5 GHz. Similarly,&#13;
the optimized S22 value reduced to-53.09 dB from-49.57 dB, further enhancing&#13;
performance. The gain of the antenna increased from 12.82 dBi to 13.58 dBi, and&#13;
the directivity improved from 12.89 dBi to 13.69 dBi, demonstrating substantial&#13;
performance enhancements. Port isolation has been maintained with S12 and S21&#13;
values below-10 dB. Additionally, the array antenna achieved a beam switching&#13;
capacity ranging from −17◦ to +17◦, making it highly effective for beam-steering&#13;
applications. This research contributes to the field by showcasing the potential&#13;
of machine learning-assisted antenna design and optimization, addressing chal&#13;
lenges such as computational complexity and design accuracy. The findings hold&#13;
promise for 5G networks and satellite communications applications, paving the&#13;
way for more efficient and reliable antenna systems
A Master of Science (M.Sc) Thesis in Electronics and Telecommunication Engineering (ETE) Department at Chittagong University of Engineering and Technology (CUET).
</description>
<dc:date>2025-01-26T00:00:00Z</dc:date>
</item>
<item rdf:about="http://103.99.128.19:8080/xmlui/handle/123456789/574">
<title>A HYBRID DEEP LEARNING MODEL FOR  HANDWRITTEN BANGLA AND ENGLISH DIGIT  RECOGNITION</title>
<link>http://103.99.128.19:8080/xmlui/handle/123456789/574</link>
<description>A HYBRID DEEP LEARNING MODEL FOR  HANDWRITTEN BANGLA AND ENGLISH DIGIT  RECOGNITION
Akbar, Md. Ali; Student ID:, 19METE004F
Convolutional Neural Network (CNN) is widely used for handwritten Bangla &#13;
and English digit recognition. The problem with the conventional CNN is that it &#13;
has an intricate and time-consuming convolution process to create a feature &#13;
map and it requires a lot of computation time during the classification phase.&#13;
This work primarily proposes substituting the convolution kernels of the classic &#13;
CNN model with dilated convolution kernels to address the limitations in &#13;
handwritten digit recognition. Although the computation time decreases, the &#13;
Dilated CNN model’s feature extraction part still performs inefficiently due to &#13;
information loss, which lowers accuracy. Observing these above-mentioned &#13;
problems, a Hybrid Dilated CNN (HDC) model is proposed by using dilated &#13;
convolution kernels with different dilation rates to eliminate the detail loss &#13;
problem of the Dilated CNN model. The proposed Hybrid Dilated CNN (HDC) &#13;
model can achieve larger receptive fields and can extract distant features. &#13;
Finally, the Support Vector Machine (SVM) and K-Nearest Neighbor (KNN) are &#13;
applied as a classifier to increase the digit classification accuracy. The &#13;
BanglaLekha-Isolated dataset of Bangla digits is used to verify the proposed &#13;
methods and demonstrate that under the same conditions, the accuracy of the&#13;
Dilated CNN, HDC, and the proposed HDC with KNN are 95.85%, 96.72%, and &#13;
96.80%, respectively. The proposed HDC model with KNN shows 100% &#13;
accuracy in the case of the MNIST dataset of English digits. The HDC – KNN &#13;
model's effectiveness in recognizing digits of both Bangla and English language &#13;
is perfectly demonstrated by the values of Precision, Recall, and F1 Score&#13;
parameters. More evidence of the HDC – KNN model's superiority comes from &#13;
the values of TPR and FPR parameters that are calculated for each of the digit &#13;
classes in the case of both languages. These experimental findings suggest that &#13;
integrating the HDC with the KNN algorithm improves the effectiveness of &#13;
recognizing handwritten Bangla and English digits.
A Master of Engineering (M.Engg) Thesis in ETE Department at Chittagong University of Engineering and Technology (CUET).
</description>
<dc:date>2024-10-07T00:00:00Z</dc:date>
</item>
<item rdf:about="http://103.99.128.19:8080/xmlui/handle/123456789/564">
<title>Deep Learning based Bandwidth Extension and Speech Restoration</title>
<link>http://103.99.128.19:8080/xmlui/handle/123456789/564</link>
<description>Deep Learning based Bandwidth Extension and Speech Restoration
Taher, Taieba; ID:, 19METE011P
Speech Bandwidth Extension (SBE) remains a challenging task in speech pro&#13;
cessing. It involves the intricate process of estimating missing frequency com&#13;
ponents in low-resolution signals to reconstruct high-resolution transformations.&#13;
SBE is a technique used to enhance the quality of speech signals by expanding&#13;
the frequency range of the audio. This endeavor is fundamental for enhancing&#13;
speech quality, naturalness, and intelligibility, particularly in scenarios marred&#13;
by background noise and channel distortions. Low-resolution signals often lack&#13;
high-frequency components due to limited sampling rates and band-limiting, as&#13;
dictated by the Nyquist-Shannon Sampling Theorem. The inherent loss of high&#13;
frequency details stems from aliasing, quantization errors, and insufficient granu&#13;
larity in sampling, which poses significant challenges in restoring full-band speech.&#13;
Addressing these issues is critical to advancing SBE techniques. Traditional&#13;
methodologies heavily relied on statistical algorithms and learning techniques&#13;
such as Multi-Layer Perceptrons (MLP). However, the paradigm shift brought&#13;
about by Deep Learning has revolutionized this domain, opening avenues for&#13;
more efficient and effective solutions. This thesis delves into deep learning-based&#13;
approaches for addressing the complexities of BWE and speech restoration, par&#13;
ticularly emphasizing noisy environments. This research explores three distinct&#13;
augmentations of a frequency-domain deep learning network, each tailored to&#13;
tackle specific challenges encountered in bandwidth extension and speech en&#13;
hancement. The first approach introduces a joint bandwidth expansion and&#13;
speech enhancement paradigm utilizing Deep Neural Networks (DNNs). This&#13;
approach is meticulously designed to simultaneously expand the bandwidth of&#13;
speech signals, reduce noise interference, and maintain the quality and intelligibil&#13;
ity of the speech. Leveraging the inherent capabilities of DNNs, this methodology&#13;
accurately estimates missing speech components. It characterizes noise profiles&#13;
within degraded signals, synthesizing high-fidelity full-band speech from limited&#13;
bandwidth inputs. The experimentation demonstrates the superior performance&#13;
of this DNN-based approach, surpassing conventional methods and presenting&#13;
promising avenues for real-world applications. Stepping beyond conventional&#13;
methodologies, the second approach introduces an end-to-end frequency-domain&#13;
framework, aptly named the Robust extension-plus-enhancement of speech utiliz&#13;
ing Dual-former Network (RDNet). RDNet represents a paradigm shift in BWE,&#13;
aiming to recover full-band speech from noisy low-band signals directly. By intri&#13;
cately integrating speech enhancement and ideal bandwidth extension modules&#13;
ii&#13;
0.0–&#13;
Abstract&#13;
within a unified framework, RDNet demonstrates remarkable performance across&#13;
diverse noisy environments. Leveraging short-time Fourier transform (STFT) fea&#13;
tures for enhancement, RDNet achieves significant improvements in critical met&#13;
rics such as Short-Time Objective Intelligibility (STOI), Perceptual Evaluation&#13;
of Speech Quality (PESQ), and log-spectral distortion (LSD), underscoring its&#13;
efficacy and potential for practical deployment. Finally, this study’s investigation&#13;
culminates in introducing the conformer-motivated super-denoised network (CS&#13;
DNet), a novel approach tailored to mitigate the mismatch problem inherent in&#13;
domain-specific encoders/decoders. By embedding a conv Unet with T-F trans&#13;
formation layers, CSDNet reduces dependency on training data and outperforms&#13;
recent baseline methods across various objective metrics. Moreover, CSDNet ex&#13;
hibits favorable subjective performance in comparative studies, reaffirming its&#13;
suitability for real-time applications. Through these meticulously crafted ap&#13;
proaches, this thesis aims to advance the frontier of deep learning-based speech&#13;
bandwidth extension and speech restoration. By addressing the core challenges&#13;
of high-frequency component loss in low-resolution signals and proposing inno&#13;
vative solutions, this work offers insights and methodologies, proposes practical&#13;
solutions for tackling real-life problems, and contributes to the wider conversa&#13;
tion in the fields of speech processing and audio engineering
A Master of Science (M.Sc) Thesis in Electronics and Telecommunication Engineering (ETE) Department at Chittagong University of Engineering and Technology (CUET).
</description>
<dc:date>2025-03-11T00:00:00Z</dc:date>
</item>
</rdf:RDF>
