Thesis
A Multilingual Fine-Tuned LLM Framework for Smishing Detection in Bengali-Speaking Community
Masters by Research, Murdoch University
2026
DOI:
https://doi.org/10.60867/00000158
Abstract
Short Message Service (SMS) remains a widely used communication medium, particularly within the Bengali-speaking community. However, its popularity has made it a target for cybercriminals through smishing (SMS phishing) designed to deceive users into revealing sensitive information. The challenge is especially severe in Bengali-speaking regions, where users communicate in Bengali, English, Romanised Bengali (Banglish), and Bengali-English code-mixed SMS. Existing detection systems, typically trained on monolingual datasets, struggle to process such multilingual and code-mixed messages. One critical barrier is the scarcity of publicly available datasets that truly reflect the linguistic diversity of Bengali-speaking communities—no existing corpus comprehensively covers Bengali, English, Banglish, and code-mixed SMS with fine-grained annotations. Furthermore, most existing SMS security research adopts a binary classification approach that labels messages only as ham or spam, treating both promotional and phishing messages as spam despite their fundamentally different intents and impacts. Promotional SMS serve an important role in informing users about products, services, and offers and should not be conflated with malicious phishing content. This lack of fine-grained categorization limits the practical usefulness of current detection systems.
To address these gaps, this research introduces SmishDetect-LLM, a multilingual framework for smishing detection tailored to the Bengali-speaking community. The framework acquires a comprehensive dataset of 7,005 SMS messages encompassing all four linguistic varieties through LLM-based translation and synthetic generation, an approach adopted because no organically collected corpus spans these varieties under a three-class taxonomy. Pretrained language models—mBERT, DistilBERT Multilingual, XLM-RoBERTa, MuRIL (Multilingual Representations for Indian Languages) Large, and Gemma 3—are fine-tuned using Low-Rank Adaptation (LoRA), reducing trainable parameters by approximately 99%. Unlike conventional binary approaches, the system performs three-way classification—Normal, Promotional, and Smishing—enabling distinction between legitimate promotional content and malicious phishing attempts.
Experimental results demonstrate the efficiency of SmishDetect-LLM framework. Models trained on all four linguistic varieties consistently outperformed those trained only on Bengali and English, with mBERT showing the greatest improvement, achieving an 8.53% accuracy gain. Gemma-3 achieved the highest overall accuracy at 99.14%, while MuRIL Large was the strongest encoder model at 99.07% on the complete multilingual dataset.
These findings validate that incorporating Banglish and code-mixed text enhances classification performance for real-world Bengali SMS communication, presenting an efficient approach to multilingual SMS security for low-resource language communities.
Details
- Title
- A Multilingual Fine-Tuned LLM Framework for Smishing Detection in Bengali-Speaking Community
- Authors/Creators
- Shariul Islam
- Contributors
- Mohammed Kaosar (Supervisor) - Murdoch University, School of Information TechnologyGuanjin Wang (Supervisor) - Murdoch University, School of Information Technology
- Awarding Institution
- Murdoch University; Masters by Research
- Identifiers
- 991005909771807891
- Murdoch Affiliation
- School of Information Technology
- Resource Type
- Thesis
Metrics
1 Record Views