EVALUASI KUANTITATIF RELIABILITAS DAN KECEPATAN INFERENSI PADA TIGA ARSITEKTUR LARGE LANGUAGE MODEL (ChatGPT, DeepSeek, Gemini) MENGGUNAKAN METRIK NLP BERBASIS PYTHON

  • Abubakar Basri universitas Djuanda
Keywords: Large Language Model, halusinasi, ambiguitas prompt, cosine similarity, latensi inferensi

Abstract

The massive adoption of Large Language Models (LLMs) across sectors demands an in-depth
evaluation of the systems' intrinsic capabilities, particularly in responding to ambiguous
instructions. This study presents a structured comparative analysis of three leading LLM
architectures ChatGPT, DeepSeek, and Gemini focusing on factual reliability and computational
efficiency (inference latency) as a direct consequence of prompt-ambiguity manipulation. The
methodology adopts a Black-Box Testing approach combined with quantitative analysis,
systematically applied to 100 prompts classified into eight complexity categories. Factual reliability
was measured using a TF-IDF-based Cosine Similarity metric against ground truth, while latency
was recorded through end-to-end stopwatch measurement and processed with Python scripts on
Google Colab. Results show that Gemini achieved the highest average accuracy (66.91%), followed
by DeepSeek (64.89%) and ChatGPT (62.10%), while DeepSeek led in average inference speed
(2.84 seconds) owing to its Mixture-of-Experts architecture. Pearson correlation tests for all three
models produced negative coefficients (ChatGPT −0.200; Gemini −0.194; DeepSeek −0.030),
disproving the traditional speed–accuracy trade-off hypothesis. These findings are expected to serve
as an empirical compass for users selecting models with precision and as evaluative literature for
developers building more robust and efficient AI systems.

Published
2026-09-12
How to Cite
Basri, A. (2026). EVALUASI KUANTITATIF RELIABILITAS DAN KECEPATAN INFERENSI PADA TIGA ARSITEKTUR LARGE LANGUAGE MODEL (ChatGPT, DeepSeek, Gemini) MENGGUNAKAN METRIK NLP BERBASIS PYTHON. IKRA-ITH Informatika : Jurnal Komputer Dan Informatika, 10(2), 1274-1282. Retrieved from https://journals.upi-yai.ac.id/index.php/ikraith-informatika/article/view/7483