Klasifikasi Risiko Dropout Mahasiswa Menggunakan Algoritma Random Forest pada Dataset Predict Students' Dropout and Academic Success

  • Novita Sari Siagian Universitas HKBP Nommensen Pematang Siantar
  • Monica Sari Batubara Universitas HKBP Nommensen Pematang Siantar
  • Leony Sinaga Universitas HKBP Nommensen Pematang Siantar

Abstract

The high rate of student dropout remains a significant challenge for higher education institutions because it affects graduate quality, educational effectiveness, and student retention. Therefore, an effective approach is required to identify students at risk of dropping out at an early stage based on available academic data. This study aims to classify student dropout risk using the Random Forest algorithm and the Predict Students’ Dropout and Academic Success dataset obtained from the UCI Machine Learning Repository. A quantitative approach was employed through several stages, including dataset collection, data cleaning, binary target transformation (Dropout and Non-Dropout), data partitioning, Random Forest model development, and model evaluation using 10-Fold Cross Validation. The dataset consists of 4,424 student records with 35 attributes, including 34 predictor attributes and one target attribute. Model performance was evaluated using Accuracy, Precision, Recall, F1-Score, Area Under the Curve (AUC), and Matthews Correlation Coefficient (MCC). The experimental results achieved an Accuracy of 86.1%, Precision of 85.8%, Recall of 86.1%, F1-Score of 85.8%, AUC of 0.902, and MCC of 0.673. The AUC value indicates excellent classification capability in distinguishing students at risk of dropping out from those who are not, while the MCC value demonstrates good predictive quality despite the class imbalance in the dataset. These findings indicate that the Random Forest algorithm provides strong classification performance and has the potential to be implemented as a decision support system for the early identification of students at risk of dropping out, enabling universities to provide timely academic interventions.

Keywords: Random Forest, machine learning, student dropout, classification, academic data.

 

Published
2026-07-13
How to Cite
Siagian, N. S., Batubara, M. S., & Sinaga, L. (2026). Klasifikasi Risiko Dropout Mahasiswa Menggunakan Algoritma Random Forest pada Dataset Predict Students’ Dropout and Academic Success. IKRA-ITH Informatika : Jurnal Komputer Dan Informatika, 10(2), 629-640. https://doi.org/10.37817/ikraith-informatika.v10i2.7102