DOI: 10.12688/f1000research.181728.2 ISSN: 2046-1402
Sentiment Analysis of Acceptance TVET Online Courses on the Skill Academy App from Google Play: Leveraging Text Mining with Comparison Machine Learning Model
Darmono Darmono, Yanuar Agung Fadlullah, Khakam Ma’ruf, Muhamad Riyan Maulana, Apry Aditya Saputra, Ramzy Bin Sulaiman, Bagus Banjar Bagaskara, Rizal Justian Setiawan, Sahril Sahril Background Online Technical and Vocational Education and Training (TVET) has expanded rapidly in Indonesia through platforms such as Skill Academy. User reviews provide valuable insights into user acceptance, but their volume makes manual analysis impractical. Therefore, this study applies text mining to analyze user sentiment and compare eight machine learning models for sentiment classification. Methods The study used 3,000 reviews collected via web scraping using the google-play-scraper library. The data was then anonymized, cleaned, translated into English, and automatically sentiment-labeled using VADER. The validity of the labeling was verified by comparing it with TextBlob and the original star-rating categories using Cohen’s Kappa. Eight classification algorithms Naive Bayes, SVM, Logistic Regression, Random Forest, Decision Tree, KNN, XGBoost, and LightGBM were trained using a TF-IDF pipeline with GridSearchCV and five-fold cross-validation. Results The labeling validation results showed moderate agreement, with a Cohen’s Kappa of 0.500 for TextBlob and 0.533 for the original star rating categories. On the test data, Logistic Regression achieved the best performance with an accuracy of 85.48%, a macro-F1 score of 76.08%, a balanced accuracy of 80.86%, and a macro-AUC of 0.9486, followed by SVM and XGBoost. The sentiment distribution was dominated by positive reviews at 75.31%, followed by negative reviews at 15.57%, and neutral reviews at 9.12%. The SMOTETomek experiment improved balanced accuracy and recall for the minority class but reduced macro-precision. Word cloud and word frequency analysis revealed that the main complaints were related to the app being slow, resource-intensive, and prone to errors, while user appreciation centered on ease of learning and the quality of the material. Conclusions Research shows that logistic regression achieved the best classification performance. While user sentiment was largely positive, improving application stability remains essential for sustaining user acceptance, enhancing learning experiences, and supporting long-term adoption of online TVET platforms.
More from our Archive
-
DOI: 10.68381/jca02008 2026
Proximal Smoothness and the Lower-C
2
Property F. H. Clarke, R. J. Stern, P. R. Wolenski
-
DOI: 10.68381/jca13044 2026
Characterizations of Prox-Regular Sets in Uniformly Convex Banach Spaces Frédéric Bernard, Lionel Thibault, Nadia Zlateva
-
DOI: 10.68381/jca15047 2026
Brøndsted-Rockafellar Property and Maximality of Monotone Operators Representable by Convex Functions in Non-Reflexive Banach Spaces Maicon Marques Alves, Benar Fux Svaiter
-
DOI: 10.68381/jca16027 2026
Proximal Smoothness and the Exterior Sphere Condition Chadi Nour, Ron J. Stern, Jean Takche
-
DOI: 10.68381/jca16053 2026
A New Old Class of Maximal Monotone Operators Maicon Marques Alves, Benar Fux Svaiter
-
DOI: 10.68381/jca13045 2026
Maximal Monotonicity via Convex Analysis Jonathan Borwein
-
DOI: 10.68381/jca08009 2026
Variational Inequalities and Regularity Properties of Closed Sets in Hilbert Spaces Giovanni Colombo, Vladimir V. Goncharov
-
DOI: 10.68381/jca17060 2026
Existence and Uniqueness of Solutions for Non-Autonomous Complementarity Dynamical Systems Bernard Brogliato, Lionel Thibault
-
DOI: 10.68381/jca01001 2026
Variational Sum of Monotone Operators H. Attouch, J.-B. Baillon, M. Théra
-
DOI: 10.68381/jca22017 2026
Weak Convexity of Sets and Functions in a Banach Space Grigorii E. Ivanov