Ketahanan Model Machine Learning terhadap Ketidaklengkapan Data Survei untuk Prediksi Kinerja Akademik
Abstract
Survey-based prediction models are commonly evaluated on complete responses, although missingness can occur randomly or structurally. This study examines when survey information remains sufficient to predict self-reported changes in academic performance using a public dataset of 361 responses, 28 predictors, and three imbalanced classes. Three feature configurations—all predictors (F1), predictors excluding three high-risk proxy items (F2), and contextual predictors only (F3)—were evaluated with CatBoost using five-times repeated stratified five-fold cross-validation. Baseline Macro F1 was 0.414±0.047, 0.384±0.042, and 0.351±0.046 for F1, F2, and F3, respectively. Recall for the decreased class was only 0.023 for F1 and 0.010 for both F2 and F3, indicating inadequate absolute performance for identifying at-risk students. For F2, 50% random item loss retained 94.8% of baseline Macro F1 under test-time missingness and 97.7% when both training and test data were incomplete. In contrast, monotone dropout with only the first 50% or 25% of items observed reduced retention to 79.3% and 78.8% (Holm-adjusted p<0.001). Thus, relative information sufficiency depends on the missingness pattern, but high retention does not imply deployability when baseline performance is weak and the minority class is nearly undetected.

