CUET DIGITAL REPOSITORY

AVaTER: A Multimodal Approach of Recognizing Emotion using Cross-modal Attention Technique

Show simple item record

dc.contributor.author Das, Avishek
dc.contributor.author ID:, 21MCSE001P
dc.date.accessioned 2026-09-10T04:23:35Z
dc.date.available 2026-09-10T04:23:35Z
dc.date.issued 2025-06-02
dc.identifier.uri http://103.99.128.19:8080/xmlui/handle/123456789/587
dc.description A Master of Science (M.Sc) Thesis in Computer Science & Engineering (CSE) Department at Chittagong University of Engineering and Technology (CUET). en_US
dc.description.abstract Multimodal emotion classification involves the analysis and identification of human emo tions by integrating data from multiple sources, such as audio, video, and text. This approach leverages the complementary strengths of each modality to enhance the accu racy and robustness of emotion recognition systems. Audio data, for example, captures vocal tone and pitch, which are crucial for detecting emotions like anger or joy, while video data provides visual cues such as facial expressions and body language. Text data, often extracted from spoken words or written content, adds context and semantic depth to the emotion analysis. However, one significant challenge is effectively integrating these diverse data sources, each with unique characteristics and levels of noise. Additionally, the scarcity of large, annotated multimodal datasets in Bangla limits the training and evaluation of models. In this work, we introduced a novel multimodal Bangla dataset named MAViT-Bangla (Multimodal Audio Video Text Bangla dataset), which consists of 1002 samples incorporating audio, video, and text modalities. This dataset includes emotional categories such as anger, fear, joy, and sadness, providing a rich resource for emotion recognition studies in the Bangla language. Each sample in MAViT-Bangla was meticulously annotated to ensure high-quality labels, making it a valuable asset for researchers working in this domain. Moreover, we developed a framework for emotion recognition that utilizes a cross-modal attention mechanism among unimodal features. This mechanism facilitates the interaction and fusion of features from different modalities, enhancing the model’s ability to capture nuanced emotional cues. The proposed approach demonstrated its effectiveness by achieving an F1 score of 0.64, showcasing significant improvement over unimodal methods. This indicates that integrating multiple modalities through a cross-modal attention mechanism can substantially enhance the performance of emotion recognition systems, especially in the context of the Bangla language, where resources have historically been limited. The MAViT-Bangla dataset and our framework thus represent significant advancements in the field of multimodal emotion classification. en_US
dc.description.sponsorship N/A en_US
dc.language.iso en en_US
dc.publisher CUET en_US
dc.relation.ispartofseries ;TCD-141
dc.subject Multimodal Emotion Recognition en_US
dc.subject Bangla Emotion en_US
dc.subject Cross-Modal Attention en_US
dc.subject Natural Language Processing en_US
dc.title AVaTER: A Multimodal Approach of Recognizing Emotion using Cross-modal Attention Technique en_US
dc.type Thesis en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search DSpace


Advanced Search

Browse

My Account