Fitri Marisa, Sharifah Sakina Syed Ahmad, Deshinta Arrova Dewi, Agustinus Noertjahyana, Anastasia L. Maukar, Istiadi
This study investigates sentiment classification in Indonesian social media discussions around mental health, focusing on three methodological challenges: class imbalance, the limitations of lexicon-based automatic labeling as a silver-standard proxy, and the need for auditable explanations. The dataset contains 5,184 lexicon-labeled tweets, preprocessed and represented using TF-IDF. To prevent data leakage, imbalance handling is performed via train-only resampling (SMOTE applied only to training data) within a leakage-safe pipeline. We evaluate several classical models using accuracy and macro metrics (macro-F1/macro-recall) to ensure fairer assessment across classes, and we add a modern Indonesian Twitter transformer baseline (IndoBERTweet) under the same repeated stratified split protocol. To validate label quality, we construct a Label Audit Subset (LAS) of 600 tweets (200 per lexicon class), annotated by two annotators and finalized through adjudication to produce gold labels. Although annotation reliability is high (92.5% agreement; Cohen's κ = 0.874 ), the audit reveals substantial lexicon-label noise (53.5% mismatch; κ = 0.198 ), particularly for Positive and Neutral labels, yielding more conservative performance estimates under gold-label evaluation. To improve auditability, we implement a confidence-aware LIME protocol that separates high-confidence, borderline, and hard-case predictions using probability, margin, and entropy. Token summaries exhibit drift when case selection is based on gold labels rather than lexicon labels (Jaccard@20 = 0.176–0.290), indicating that label noise can shift interpretive conclusions. Finally, token-drift signals are mapped into operational design hypotheses for further studies, including low-pressure gamification, trusted-ties social support (silaturrahmi parameter, i.e., culturally grounded social-support ties), and friction mitigation via triage/escalation mechanisms in risk scenarios. © This article is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. License details: https://creativecommons.org/licenses/by-sa/4.0/
Informatics Engineering Department, Universitas Widya Gama Malang, Indonesia; Faculty of Artificial Intelligence and Cyber Security, Universiti Teknikal Malaysia, Melaka, Malaysia; Center for Data Science and Sustainable Technologies, INTI International University, Malaysia; Informatics Department, Petra Christian University, Indonesia; Industrial Engineering Department, President University, Indonesia
Research at a Glance
Register to unlockTopics & SDG Alignment
Register to unlockCollaboration
Register to unlockAuthor Profile (Selected)
Register to unlockReferences Overview
Register to unlockJournal & Source
Register to unlockMetadata & Integrity
Register to unlock