Reward Model in Reinforcement Learning for Web Application Security Testing
Keywords
Abstract
The paper addresses the problem of reward function design for reinforcement learning in cyberattack scenario analysis, where traditional approaches within Markov decision processes, focused primarily on reaching a terminal state, insufficiently reflect the temporal dynamics of compromise and the varying importance of intermediate objectives. A reward model is proposed that combines a temporal component time-to-compromise (TTC) with semantic weights of intermediate goals, particularly actions related to privilege escalation and access to critical resources. Unlike conventional reward schemes, this approach produces a locally informative reinforcement signal that guides learning not only toward achieving compromise but also toward constructing time-efficient and structurally meaningful attack trajectories. The experimental study was conducted on 12 stochastically generated MAL-scenarios (7200 runs per method) based on open attack models. The results showed that the proposed model yields a substantially lower mean TTC compared to three baseline approaches (improvement ranging from 5 % to 21 %) with high statistical significance, while simultaneously increasing the rate of successful target state achievement and reducing the number of inadmissible actions. The resulting policies are also more interpretable: the share of targeted actions in the constructed trajectories increases, while auxiliary steps are displaced in favor of strategically significant privilege escalation actions. Sensitivity analysis confirmed the robustness of the results to the choice of model parameters. The practical significance of the work lies in the potential application of the proposed approach for identifying, ranking, and prioritizing the most dangerous web application compromise scenarios with respect to their temporal and structural criticality.
