CUET DIGITAL REPOSITORY

Developing an Automated Framework for Detecting Check-worthiness of Bangla Claims

Show simple item record

dc.contributor.author Rahman, Md. Rashadur
dc.date.accessioned 2026-09-06T04:00:33Z
dc.date.available 2026-09-06T04:00:33Z
dc.date.issued 2025-02-05
dc.identifier.uri http://103.99.128.19:8080/xmlui/handle/123456789/550
dc.description A Master of Science (M.Sc) Thesis in Computer Science and Engineering Department at Chittagong University of Engineering and Technology (CUET). en_US
dc.description.abstract The evolution of technology is swiftly reshaping media trends, particularly with the mainstream media’s migration to online platforms. The rapid proliferation of misinformation through digital platforms and social media has become a critical challenge in today’s information-rich society. Despite the crucial role of fact checking in combating misinformation, traditional manual approaches are time intensive and insufficient to address the sheer volume of content. Automated claim detection systems serve as a critical first step in fact-checking by identify ing potentially falsifiable claims within the text, enabling quicker response times and reducing the workload for human fact-checkers. While research on claim de tection has been conducted for resource-rich languages like English, this domain remains underexplored for resource-constrained languages such as Bangla. In this study, we introduce the first-ever multiclass check-worthy claim detection dataset for Bangla named CheckBanC, comprising 10,023 sentences annotated into three categories. To address the unique challenges of Bangla claim detection, we pro pose a novel hybrid feature fusion model that combines contextual embeddings from DistilBERT with statistical bigram features. This approach captures both local syntactic patterns and global semantic relationships, significantly enhanc ing the model’s ability to identify check-worthy claims. Experimental results demonstrate the superiority of our method, achieving a new state-of-the-art per formance on the CheckBanC dataset. It outperforms various machine learning, deep learning, and transformer-based baselines, achieving a weighted F1-score of 0.84, surpassing other existing state-of-the-art approaches in claim detection. Our findings establish a strong foundation for future research in automated claim detection for Bangla, facilitating fact-checking efforts and curbing misinformation in resource-constrained languages. en_US
dc.language.iso en en_US
dc.publisher CUET en_US
dc.relation.ispartofseries ;TCD-110
dc.subject Fact-checking en_US
dc.subject Claim detection en_US
dc.subject Distil-BERT en_US
dc.subject Bangla claim dataset en_US
dc.subject Computational journalism en_US
dc.subject Text classification en_US
dc.title Developing an Automated Framework for Detecting Check-worthiness of Bangla Claims en_US
dc.type Thesis en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search DSpace


Advanced Search

Browse

My Account