Dataset for Detecting AI-Generated Source Code
DOI: 10.21293/1818-0442-2025-28-2-106-110
DOI: 10.21293/1818-0442-2025-28-2-106-110
Abstract: Modern generative language models are actively used for automated source code generation, necessitating the development of detection methods. However, the creation of datasets for identifying machine-generated code remains a challenging task. This paper analyzes existing datasets, identifying their limitations. An original dataset is developed, comprising Python programming language solutions to problems written by humans and generated by state-of-the-art language models. Experimental evaluation is conducted using machine learning methods. The results demonstrate the promise of the proposed dataset while indicating the need for its further expansion or conducting new experiments to identify the optimal model.
Keywords: source code, machine learning, language models, dataset, code classification
For citation:
Bukina S. G., Harchenko S. S. Dataset for Detecting AI-Generated Source Code. Doklady Tomskogo gosudarstvennogo universiteta sistem upravleniya i radioelektroniki, 2025, vol. 28, no. 2, pp. 106–110. DOI: 10.21293/1818-0442-2025-28-2-106-110
Authors and copyright holders:
Executive Secretary of the Editor’s Office
Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia
Phone / Fax: + 7 (3822) 701-582
Viktor N. Maslennikov
Executive Secretary of the Editor’s Office
Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia
Phone / Fax: + 7 (3822) 51-21-21 / 51-43-02