STUDENT DROPOUT RISK ANALYSIS USING A HYBRID K-MEANS AND CART METHOD

Authors

  • Ridwan Afriansah Program Studi Magister Manajemen Rekayasa, Fakultas Teknik, Universitas Halu Oleo, Kendari, Indonesia
  • Ahmad Syarif Sukri Program Studi Magister Manajemen Rekayasa, Fakultas Teknik, Universitas Halu Oleo, Kendari, Indonesia
  • Mansur Program Studi Magister Manajemen Rekayasa, Fakultas Teknik, Universitas Halu Oleo, Kendari, Indonesia
  • Muhammad Nadzirin Anshari Nur Program Studi Magister Manajemen Rekayasa, Fakultas Teknik, Universitas Halu Oleo, Kendari, Indonesia
  • Hasmina Tari Mokui Program Studi Magister Manajemen Rekayasa, Fakultas Teknik, Universitas Halu Oleo, Kendari, Indonesia
  • Tambi Program Studi Magister Manajemen Rekayasa, Fakultas Teknik, Universitas Halu Oleo, Kendari, Indonesia

DOI:

https://doi.org/10.51876/simtek.v11i2.1883

Keywords:

CART, Student Dropout, Educational Data Mining, K-Means Clustering, Two-Stage Modeling

Abstract

Student dropout is an important issue in higher education because it affects study continuity, resource efficiency, and institutional performance. This study aims to develop a student dropout risk analysis model using a hybrid K-Means and Classification and Regression Tree (CART) approach within a two-stage modeling framework. The data were obtained from the Academic Information System of STMIK X for the academic years 2015/2016–2018/2019. After data cleaning and selection, 631 student records were obtained, consisting of 119 dropout students and 512 active/graduated students. In the first stage, K-Means was employed to identify the underlying student segmentation structure, while in the second stage, CART was applied using the cluster labels as an additional feature. The K-Means results identified K = 10 as the statistically optimal number of clusters, with a silhouette score of 0.6330. These ten clusters were subsequently interpreted into three risk groups: low, medium, and high. On the test data, the hybrid model achieved an accuracy of 92.25%, precision of 60.87%, recall of 87.50%, and an F1-score of 71.79%, outperforming the standalone CART model, which achieved an accuracy of 88.03%, precision of 44.44%, recall of 25.00%, and an F1-score of 32.00%. Validation using Orange Data Mining also demonstrated a consistent improvement, with the hybrid model achieving an accuracy of 92.60%. Attendance was identified as the most dominant predictor, with a feature importance value of 0.71. These results indicate that K-Means segmentation can enrich the data representation prior to CART classification and improve the early detection of students at risk of dropping out

Additional Files

Published

06-10-2026

How to Cite

Afriansah, R., Sukri, A. S., Mansur, M., Nur, M. N. A., Mokui, H. T., & Tambi, T. (2026). STUDENT DROPOUT RISK ANALYSIS USING A HYBRID K-MEANS AND CART METHOD. Simtek : Jurnal Sistem Informasi Dan Teknik Komputer, 11(2), 313–318. https://doi.org/10.51876/simtek.v11i2.1883
Abstract View: 0