COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE AND SYSTEMS ANALYSIS

Reward Model in Reinforcement Learning for Web Application Security Testing

Authors

Vinnytsia National Technical University ROR
L. M. Kupershtein ORCID 0000-0001-6737-7134
Vinnytsia National Technical University ROR

Keywords

markov decision process cybersecurity threat attack graph reinforcement learning neural network reward function machine learning artificial intelligence modeling

Abstract

The paper addresses the problem of reward function design for reinforcement learning in cyberattack scenario analysis, where traditional approaches within Markov decision processes, focused primarily on reaching a terminal state, insufficiently reflect the temporal dynamics of compromise and the varying importance of intermediate objectives. A reward model is proposed that combines a temporal component time-to-compromise (TTC) with semantic weights of intermediate goals, particularly actions related to privilege escalation and access to critical resources. Unlike conventional reward schemes, this approach produces a locally informative reinforcement signal that guides learning not only toward achieving compromise but also toward constructing time-efficient and structurally meaningful attack trajectories. The experimental study was conducted on 12 stochastically generated MAL-scenarios (7200 runs per method) based on open attack models. The results showed that the proposed model yields a substantially lower mean TTC compared to three baseline approaches (improvement ranging from 5 % to 21 %) with high statistical significance, while simultaneously increasing the rate of successful target state achievement and reducing the number of inadmissible actions. The resulting policies are also more interpretable: the share of targeted actions in the constructed trajectories increases, while auxiliary steps are displaced in favor of strategically significant privilege escalation actions. Sensitivity analysis confirmed the robustness of the results to the choice of model parameters. The practical significance of the work lies in the potential application of the proposed approach for identifying, ranking, and prioritizing the most dangerous web application compromise scenarios with respect to their temporal and structural criticality.

2 0

How to Cite

[1]
“Reward Model in Reinforcement Learning for Web Application Security Testing”, Вісник ВПІ, no. 4, pp. 119–131, Sep. 2026, doi: 10.31649/.

Author Biographies

A. V. Prytula, Vinnytsia National Technical University

Post-Graduate Student of the Chair of Information Protection

L. M. Kupershtein, Vinnytsia National Technical University

Cand. Sc. (Eng.), Associate Professor of the Chair of Information Protection,

References

[1] M. C. Ghanem, and T. M. Chen, “Reinforcement Learning for Efficient Network Penetration Testing,” Information, vol. 11, no. 1, p. 6, 2019. https://doi.org/10.3390/info11010006 .
[2] M. C. Ghanem, T. M. Chen, and E. G. Nepomuceno, “Hierarchical reinforcement learning for efficient and effective automated penetration testing of large networks,” Journal of Intelligent Information Systems, vol. 60, no. 2, pp. 281-303, 2022. https://doi.org/10.1007/s10844-022-00738-0 .
[3] F. M. Zennaro, and L. Erdődi, “Modelling penetration testing with reinforcement learning using capture-the-flag challenges: Trade-offs between model-free learning and a priori knowledge,” IET Information Security, vol. 17, no. 3, pp. 441-457, 2023. https://doi.org/10.1049/ise2.12107 .
[4] J. Yi, and X. Liu, “Deep Reinforcement Learning for Intelligent Penetration Testing Path Design,” Applied Sciences, vol. 13, no. 16, p. 9467, 2023. https://doi.org/10.3390/app13169467 .
[5] А. Толкачова, та М.-М. Посувайло, «Тестування на проникнення з використанням глибокого навчання з підкріпленням,» Кібербезпека: освіта, наука, техніка, т. 3, № 23, 2024. https://doi.org/10.28925/2663-4023.2024.23.1730 .
[6] В. Вікулов, та І. Пишнограєв, «Автоматизоване виявлення вразливостей SQL-ін’єкцій у чат-ботах за допомогою навчання з підкріпленням,» Кібербезпека: освіта, наука, техніка, т. 1, № 29, 2025. https://doi.org/10.28925/2663-4023.2025.29.873.
[7] D. J. Leversage, and E. J. Byres, “Estimating a System’s Mean Time-to-Compromise,” IEEE Security & Privacy, vol. 6, no. 1, pp. 52-60, 2008. https://doi.org/10.1109/msp.2008.9 .
[8] W. Nzoukou, L. Wang, S. Jajodia, and A. Singhal, “A Unified Framework for Measuring a Network’s Mean Time-to-Compromise," in 2013 IEEE 32nd International Symposium on Reliable Distributed Systems, 2013, pp. 215-224. https://doi.org/10.1109/srds.2013.30 .
[9] U. Garg, G. Sikka, and L. K. Awasthi, “Empirical risk assessment of attack graphs using time to compromise framework,” International Journal of Information and Computer Security, vol. 16, no. 1/2, p. 33, 2021. https://doi.org/10.1504/ijics.2021.117393 .
[10] E. Rencelj Ling, and M. Ekstedt, “Estimating Time-To-Compromise for Industrial Control System Attack Techniques Through Vulnerability Data,” SN Computer Science, vol. 4, no. 3, 2023. https://doi.org/10.1007/s42979-023-01750-z .
[11] B. E. Strom, A. Applebaum, D. P. Miller, K. C. Nickels, A. G. Pennington, and C. B. Thomas, MITRE ATT&CK: Design and philosophy (Technical Report MP180360R1). The MITRE Corporation, 2020. [Online]. Available: https://www.mitre.org/sites/default/files/2021-11/prs-19-01075-28-mitre-attack-design-and-philosophy.pdf .
[12] P. Mell, K. Scarfone, and S. Romanosky, “Common Vulnerability Scoring System,” IEEE Security and Privacy Magazine, vol. 4, no. 6, pp. 85-89, 2006. https://doi.org/10.1109/msp.2006.145 .
[13] M. Ibrahim, and R. Elhafiz, “Security Analysis of Cyber-Physical Systems Using Reinforcement Learning,” Sensors, vol. 23, no. 3, p. 1634, 2023. https://doi.org/10.3390/s23031634 .
[14] M. Ibrahim, and R. Elhafiz, “Security Assessment of Industrial Control System Applying Reinforcement Learning,” Processes, vol. 12, no. 4, p. 801, 2024. https://doi.org/10.3390/pr12040801 .
[15] Y. Wang, Y. Li, X. Xiong, J. Zhang, Q. Yao, and C. Shen, “DQfD-AIPT: An Intelligent Penetration Testing Framework Incorporating Expert Demonstration Data,” Security and Communication Networks, vol. 2023, pp. 1-15, 2023. https://doi.org/10.1155/2023/5834434 .
[16] Л. Д. Ганенко, «Метод адаптивного формування винагороди за умов невизначеності динамічних об’єктів,» Телекомунікаційні та інформаційні технології, № 1, с. 23-30, 2026. https://doi.org/10.31673/2412-4338.2026.019003 .
[17] B.-S. Kim, H.-W. Suk, Y.-H. Choi, D.-S. Moon, and M.-S. Kim, “Optimal Cyber Attack Strategy Using Reinforcement Learning Based on Common Vulnerability Scoring System,” Computer Modeling in Engineering & Sciences, vol. 141, no. 2, pp. 1551-1574, 2024. https://doi.org/10.32604/cmes.2024.052375 .
[18] А. Притула, та Л. Куперштейн, «Моделювання сценаріїв кібератак як марковського процесу із семантично обмеженим простором дій,» Кібербезпека: освіта, наука, техніка, № 33, с. 555-569, 2026. https://doi.org/10.28925/2663-4023.2026.33.1232 .
[19] P. Johnson, R. Lagerström, and M. Ekstedt, “A Meta Language for Threat Modeling and Attack Simulations,” in Proceedings of the 13th International Conference on Availability, Reliability and Security, pp. 1-8, 2013. https://doi.org/10.1145/3230833.3232799 .
[20] V. Mnih, et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529-537, 2015. https://doi.org/10.1038/nature14236 .
[21] EnterpriseLang: Enterprise Language for the Meta Attack Language framework. [Online]. Available: https://github.com/mal-lang/enterpriseLang .
[22] WebGoat: A deliberately insecure web application [Online]. Available: https://github.com/WebGoat/WebGoat .

Most read articles by the same author(s)