Deteksi Domain DGA menggunakan Random Forest Berdasarkan Karakteristik String dengan Evaluasi Unseen Family
DOI:
https://doi.org/10.55606/jupti.v5i3.7994Keywords:
Characteristic String, DGA Detection, Domain Generation Algorithm, Random Forest, Unseen FamilyAbstract
Domain Generation Algorithm (DGA) is a mechanism used by malware to algorithmically generate domain names as candidate destinations for communication with command and control servers. Domain string characteristics can be used to distinguish algorithmically generated domains from legitimate domains. This study aims to identify DGA domains based on string characteristics using Random Forest and evaluate its performance against DGA families not observed during training. The dataset consists of 20,000 balanced domains, comprising 10,000 DGA domains from 52 families and 10,000 legitimate domains. A total of 21 features were extracted, including domain length, name length, character counts and ratios, unique character counts, domain structure, and Shannon entropy. Evaluation was conducted using Random Stratified Split and Unseen Family Split. Random Stratified Split achieved 91.83% accuracy, 92.14% precision, 91.45% recall, and 91.79% F1-score. Unseen Family Split achieved 85.55% accuracy, 91.47% precision, 78.41% recall, and 84.44% F1-score. The results indicate a change in classification performance when the model encounters unseen DGA families. Feature importance identified name length and the ratio of name length to domain length as the two most important features. The findings provide an evaluation basis for DGA classification under unseen-family conditions.
References
Biros, H., & Kantor, M. (2025). Enhancing DGA detection with machine learning algorithms. Journal of Telecommunications and Information Technology, 31–44. https://doi.org/10.26636/jtit.2025.FITCE2024.2033
Chen, S., Lang, B., Chen, Y., & Xie, C. (2023). Detection of algorithmically generated malicious domain names with feature fusion of meaningful word segmentation and N-gram sequences. Applied Sciences, 13(7). https://doi.org/10.3390/app13074406
Divya, T., Amritha, P. P., & Viswanathan, S. (2022). A model to detect domain names generated by DGA malware. Procedia Computer Science, 215, 403–412. https://doi.org/10.1016/j.procs.2022.12.042
Fan, B., Ma, H., Liu, Y., Yuan, X., & Ke, W. (2024). KDTM: Multi-stage knowledge distillation transfer model for long-tailed DGA detection. Mathematics, 12(5). https://doi.org/10.3390/math12050626
Jeremiah, D., Rafiq, H., Ta, V. T., Usman, M., Raza, M., & Awais, M. (2025). NIOM-DGA: Nature-inspired optimised ML-based model for DGA detection. Computers & Security, 157. https://doi.org/10.1016/j.cose.2025.104561
Jiang, K., Wu, S., Huang, R., & Deng, Z. (2025). DGA domain name detection model based on gated convolution and LSTM. KSII Transactions on Internet and Information Systems, 19(3), 987–1006. https://doi.org/10.3837/tiis.2025.03.015
Lee, H., Do Yoo, J., Jeong, S., & Kim, H. K. (2024). Detecting domain names generated by DGAs with low false positives in Chinese domain names. IEEE Access, 12, 123716–123730. https://doi.org/10.1109/ACCESS.2024.3454242
Liew, S. R. C., & Law, N. F. (2023). Use of subword tokenization for domain generation algorithm classification. Cybersecurity, 6(1). https://doi.org/10.1186/s42400-023-00183-8
Nadagoudar, R. B., & Ramakrishna, M. (2024). DGA domain name detection and classification using deep learning models. International Journal of Advanced Computer Science and Applications, 15. https://doi.org/10.14569/IJACSA.2024.0150730
Nie, Y., Liu, S., Qian, C., Deng, C., Li, X., Wang, Z., & Kuang, X. (2023). Multimodel collaboration to combat malicious domain fluxing. Electronics, 12(19). https://doi.org/10.3390/electronics12194121
Pelayo-Benedet, T., Rodríguez, R. J., & Gañán, C. H. (2025). The machines are watching: Exploring the potential of large language models for detecting algorithmically generated domains. Journal of Information Security and Applications, 93. https://doi.org/10.1016/j.jisa.2025.104176
Peng, X., He, J., Ni, L., & Yang, G. (2026). Beyond pattern matching: A cognitive-driven framework for DGA detection via dual-perspective anomaly perception. Electronics, 15(9). https://doi.org/10.3390/electronics15091934
Selvaraj, S., & Panjanathan, R. (2024). WordDGA: Hybrid knowledge-based word-level domain names against DGA classifiers and adversarial DGAs. Informatics, 11(4). https://doi.org/10.3390/informatics11040092
Sun, X., & Liu, Z. (2023). Domain generation algorithms detection with feature extraction and Domain Center construction. PLOS ONE, 18(1). https://doi.org/10.1371/journal.pone.0279866
Suryotrisongko, H., Musashi, Y., Tsuneda, A., & Sugitani, K. (2022). Robust botnet DGA detection: Blending XAI and OSINT for cyber threat intelligence sharing. IEEE Access, 10, 34613–34624. https://doi.org/10.1109/ACCESS.2022.3162588
Tang, J., Guan, Y., Zhao, S., Wang, H., & Chen, Y. (2024). DGA domain detection based on Transformer and rapid selective kernel network. Electronics, 13(24). https://doi.org/10.3390/electronics13244982
Vranken, H., & Alizadeh, H. (2022). Detection of DGA-generated domain names with TF-IDF. Electronics, 11(3). https://doi.org/10.3390/electronics11030414
Wang, Z., Guo, Y., & Montgomery, D. (2022). Machine learning-based algorithmically generated domain detection. Computers and Electrical Engineering, 100. https://doi.org/10.1016/j.compeleceng.2022.107841
Xie, M., He, R., & He, A. (2024). Deep learning DGA malicious domain name detection based on multi-stage feature fusion. Applied and Computational Engineering, 64(1), 1–8. https://doi.org/10.54254/2755-2721/64/20241334
Yang, C., Lu, T., Yan, S., Zhang, J., & Yu, X. (2022). N-Trans: Parallel detection algorithm for DGA domain names. Future Internet, 14(7). https://doi.org/10.3390/fi14070209
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Jurnal Publikasi Teknik Informatika

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.




