Classifying Thai SMS Messages from Scammers using Machine Learning
Main Article Content
Abstract
This research aims to: 1) classify Thai SMS messages related to scams using Machine Learning; 2) investigate the use of PyThaiNLP combined with TF-IDF for text feature extraction and apply the SMOTE technique to address class imbalance before training with the Naïve Bayes algorithm; and 3) evaluate the performance of the model in classifying Thai SMS messages into three categories: normal, scam, and promotional messages, using a dataset of 5,000 samples. The experimental results using the Naïve Bayes algorithm show that the normal class achieved a Precision (accuracy of positive predictions) of 0.94, Recall (ability to detect actual positives) of 0.84, and F1-Score (balance between Precision and Recall) of 0.88. The promotional class obtained a Precision of 0.57, Recall of 0.81, and F1-Score of 0.67, while the scam class achieved a Precision of 0.84, Recall of 0.91, and F1-Score of 0.87. These results indicate that the model performs well, particularly in detecting scam messages. Furthermore, applying the SMOTE technique improved the model performance, achieving the highest Accuracy of 85.00%, with Precision, Recall, and F1-Score at a good level. In conclusion, this approach is effective for Thai SMS scam detection and can be applied in real-world scenarios.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
References
ธ. ไทยพาณิชย์, “คลิกเดียวชีวิตเปลี่ยน! วิธีรับมือ ‘ลิงก์แปลกปลอม’ และ ‘ข้อความหลอกลวง,’”, [ออนไลน์]. https://www.scb.co.th/th/personal-banking/fraud-fighter/update-fraud/sms-one-click. (เข้าถึงเมื่อ: 9 เมษายน 2569).
W. Phatthiyaphaibun et al., "PyThaiNLP: Thai Natural Language Processing in Python," in Proc. 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023), 2023, pp. 25–36.
ศวิตา ทองขุนวงศ์ และ ภัคพล สวัสกมล, "การเปรียบเทียบประสิทธิภาพของตัวแบบการเรียนรู้ของเครื่องสำหรับการจำแนกผู้ป่วยโรคมะเร็งปอด," วารสารวิจัย มทร. กรุงเทพ, ปีที่ 18, ฉบับที่ 1, หน้า 33–42, มกราคม–มิถุนายน 2567.
N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic minority over-sampling technique," J. Artif. Intell. Res., vol. 16, pp. 321–357, Jun. 2002.
นันทวัฒน์ หล้าซิว, "กรอบพัฒนาวิธีการจำแนกประเภทข้อความสนทนาด้วยเทคนิคการเรียนรู้เชิงลึกร่วมกับเทคนิคการเพิ่มข้อมูล," วิทยานิพนธ์ วท.ม. (วิทยาการคอมพิวเตอร์), มหาวิทยาลัยธรรมศาสตร์, กรุงเทพฯ, 2566.
H. C. Wu, R. W. P. Luk, K. F. Wong, and K. L. Kwok, "Interpreting TF-IDF term weights as making relevance decisions," ACM Trans. Inf. Syst., vol. 26, no. 3, pp. 1–37, Jun. 2008.
องอาจ อุ่นอนันต์ และ พยุง มีสัจ, "การจำแนกความน่าเชื่อถือของเว็บไซต์แหล่งข่าวภาษาไทยโดยใช้เทคนิคการทำเหมืองข้อมูล," วารสารวิชาการมหาวิทยาลัยอีสเทิร์นเอเชีย ฉบับวิทยาศาสตร์และเทคโนโลยี, ปีที่ 14, ฉบับที่ 2, หน้า 101–116, พฤษภาคม–สิงหาคม 2563.
วสันต์ เจริญทองตระกูล, "พฤติกรรมการรับส่งข้อความสั้น (SMS) ทางโทรศัพท์เคลื่อนที่ของนักเรียน นิสิต นักศึกษา ในกรุงเทพมหานคร," วิทยานิพนธ์ ว.ม. (สื่อสารมวลชน), มหาวิทยาลัยธรรมศาสตร์, กรุงเทพฯ, 2546.
สุปราณี วงษ์แสงจันทร์, ประภาพร กุลลิ้มรัตน์ชัย และ พิมล จงวรนนท์, "การรับมือกับกลโกงในโลกไซเบอร์," วารสารวิชาการมหาวิทยาลัยอีสเทิร์นเอเชีย ฉบับวิทยาศาสตร์และเทคโนโลยี, ปีที่ 19, ฉบับที่ 2, หน้า 14–26, พฤษภาคม–สิงหาคม 2568.
สิริลักข์ เมืองนิล, ศุภกร ปุญญฤทธิ์ และ สุณีย์ กัลยะจิตร,"การป้องกันการตกเป็นเหยื่อแก๊งคอลเซ็นเตอร์," วารสารกระบวนการยุติธรรม, ปีที่ 18, ฉบับที่ 3, หน้า 1–22, พฤษภาคม–สิงหาคม 2568.
ชวัลวิทย์ โสภาศิริทรัพย์, สุณีย์ กัลยะจิตร และ พรชุลี นาคพงษ์, "แนวทางป้องกันการตกเป็นเหยื่อหลอกรักออนไลน์ (Romance Scam)," วารสารกระบวนการยุติธรรม, ปีที่ 17, ฉบับที่ 3, หน้า 1–16, กันยายน–ธันวาคม 2567.