Deteksi Anomali Lalu Lintas Jaringan Menggunakan Algoritma K-Means dengan Optimasi Jumlah Klaster Berbasis Elbow Method dan Silhouette Score
Abstract
The increasing volume of computer network traffic makes manual detection of anomalous activity inefficient, while signature-based approaches are unable to recognize new attack patterns. This study aims to cluster network traffic patterns using the K-Means algorithm to identify anomalous traffic without utilizing labels during training. The dataset consists of 10,005 network connection records with attributes of duration, packet count, bytes sent, bytes received, bytes per packet, and protocol. Preprocessing stages include missing value and duplication checks, feature selection, one-hot encoding of the protocol attribute, and feature standardization using StandardScaler. The data was then reduced to 5,000 samples through random sampling for computational efficiency. The optimal number of clusters was determined in the range K=1 to K=5 using the Elbow Method and Silhouette Score. The results show that the optimal K value is 2 with a Silhouette Score of 0.6347. The first cluster contains 4,807 records representing normal traffic, while the second cluster contains 193 records characterized by much higher connection duration and volume of bytes sent. Validation against the original labels shows that all members of the second cluster (100%) are anomalous traffic, with 100% precision and 18.7% recall. These results prove that K-Means with cluster number optimization is able to isolate high-volume anomalous traffic patterns in an unsupervised manner.

