Specifics of Detecting Training Data Compromise in Artificial Intelligence Systems
DOI: 10.21293/1818-0442-2026-29-1-124-134
DOI: 10.21293/1818-0442-2026-29-1-124-134
Abstract: Relevance. Artificial intelligence systems are currently widely deployed, but many of them are vulnerable to training data poisoning attacks. When attacking, an attacker purposefully introduces malicious instances into the dataset, which leads to a decrease in model accuracy or the emergence of hidden backdoors. Purpose of the study. Development of an approach to analyzing ways of negative impact on training sets. This approach will not only determine the most optimal impacts, but also to quantify their severity (the study proposes developing a system of metrics for quantitatively assessing poisoning attacks and building a model of optimal actions of an attacker based on Markov decision processes). Methods. The following tests were applied: the two-sample Kolmogorov–Smirnov test and the chi-square for comparing feature distributions, the Jensen–Shannon distance to estimate data drift, the Adversarial Validation method, and the value iteration algorithm to find the optimal policy in the Markov decision process (MDP). Novelty. A formal model of the interaction between an attacker and an analyst is proposed as MDP with a reward function that simultaneously accounts for the decrease in model accuracy, the indistinguishability of poisoned data and the degree of their statistical drift relative to a control set. Results. The dependencies of the probability of a successful attack on its scale for various states of analyst trust are obtained. Optimal and suboptimal sequences of attack actions are identified. It is shown that the proposed metrics quantitatively characterize both the effectiveness and stealth of the attack. Practical significance. The proposed approach allows for assessing the susceptibility of datasets to poisoning attacks during the development of artificial intelligence systems and substantiating preventative measures: regular monitoring of JS divergence, checking the distribution of class labels and applying several statistical tests to detect anomalies.
Keywords: network attack, neural network, dataset, feature matrix, activation function, Python programming language, Markov decision processes
For citation:
Vetrov I. A., Satsuta A. I., Podtopelnyy V. V. Specifics of Detecting Training Data Compromise in Artificial Intelligence Systems. Doklady Tomskogo gosudarstvennogo universiteta sistem upravleniya i radioelektroniki, 2026, vol. 29, no. 1, pp. 124–134. DOI: 10.21293/1818-0442-2026-29-1-124-134
Authors and copyright holders:
Executive Secretary of the Editor’s Office
Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia
Phone / Fax: + 7 (3822) 701-582
Viktor N. Maslennikov
Executive Secretary of the Editor’s Office
Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia
Phone / Fax: + 7 (3822) 51-21-21 / 51-43-02