Dataset for Detecting AI-Generated Source Code

DOI: 10.21293/1818-0442-2025-28-2-106-110

Download article in PDF format

JATS xml

Abstract: Modern generative language models are actively used for automated source code generation, necessitating the development of detection methods. However, the creation of datasets for identifying machine-generated code remains a challenging task. This paper analyzes existing datasets, identifying their limitations. An original dataset is developed, comprising Python programming language solutions to problems written by humans and generated by state-of-the-art language models. Experimental evaluation is conducted using machine learning methods. The results demonstrate the promise of the proposed dataset while indicating the need for its further expansion or conducting new experiments to identify the optimal model.

Keywords: source code, machine learning, language models, dataset, code classification

For citation:
Bukina S. G., Harchenko S. S. Dataset for Detecting AI-Generated Source Code. Doklady Tomskogo gosudarstvennogo universiteta sistem upravleniya i radioelektroniki, 2025, vol. 28, no. 2, pp. 106–110. DOI: 10.21293/1818-0442-2025-28-2-106-110

Authors and copyright holders:

  • Bukina S. G. , Tomsk State University of Control Systems and Radioelectronics (Tomsk, Russia)
  • Harchenko S. S. , Tomsk State University of Control Systems and Radioelectronics (Tomsk, Russia)

  • 1. Hype or not? AI’s benefits for developers explored in the 2023 Developer Survey. Available at: https://stackoverflow.blog/2023/06/14/hype-or-not-developers-have-something-to-say-about-ai/ (Аccessed: 16 September 2024).
  • 2. 2024 Developer Survey. Available at: https://survey.stackoverflow.co/2024/ai/ (Аccessed: 16 September 2024).
  • 3. Ma W., Song Y., Xue M., Wen S., Xiang Y. The «Code» of Ethics: A Holistic Audit of AI Code Generators. IEEE Transactions on Dependable and Secure Computing, 2024, vol. 21, no. 5, pp. 4997–5013.
  • 4. Oedingen M., Denz R., Engelhardt R., Hammer M. ChatGPT Code Detection: Techniques for Uncovering the Source of Code. AI Journal, 2024, vol. 5, no. 3, pp. 1066– 1094.
  • 5. Hoq M., Shi Y., Leinonen J., Babalola D. Detecting ChatGPT-Generated Code Submissions in a CS1 Course Using Machine Learning Models. Proceedings of the 55th ACM Technical Symposium on Computer Science Education, Portland, Oregon, United States, 2024, vol. 1, pp. 526–532.
  • 6. Idialu O.J., Mathews N.S., Maipradit R., Atlee J.M. Whodunit: Classifying Code as Human Authored or GPT-4 Generated - A case study on CodeChef problems. Proceedings of the 21st International Conference on Mining Software Repositories, Lisbon, Portugal, 2024, pp. 394–406.
  • 7. Li K., Hong S., Fu C., Zhang Y., Liu M. Discriminating Human-authored from ChatGPT-Generated Code Via Discernable Feature Analysis. IEEE 34th International Symposium on Software Reliability Engineering Workshops, Florence, Italy, 2023, pp. 120–127.
  • 8. Nguyen P., Rocco J., Sipio C., Rubei R. GPTSniffer: A CodeBERT-based classifier to detect source code written by ChatGPT. Journal of Systems and Software, 2024, vol. 214, 112059.
  • 9. Bukhari S.A. Issues in Detection of AI-Generated Source Code: The Requirements for the degree of Masters of Science, Calgary, 2024, 102 p.
  • 10. Sjoerd S. The Detection of AI Generated Coding Content: The Requirements for the degree of Masters of Science, Utrecht, 2024, 75 p.
  • 11. Xu Z., Sheng V. Detecting AI-Generated Code Assignments Using Perplexity of Large Language Models. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, Canada, 2024, pp. 23155–23162.
  • 12. Dataset APPS on Hugging Face. Available at: https://huggingface.co/datasets/codeparrot/apps (Аccessed: 24 October 2024).
  • 13. Chen M., Tworek J., Jun H. Evaluating Large Language Models Trained on Code. Available at: https://arxiv.org/abs/2107.03374 (Accessed: 16 February 2025).
  • 14. Allamanis M., Barr E., Devanbu P. A Survey of Machine Learning for Big Code and Naturalness. ACM Computing Surveys (CSUR), 2018, vol. 51, no. 81, pp. 1–37.
  • 15. Vaithilingam P., Wu T., Glassman E. Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. Proceedings of the CHI Conference on Human Factors in Computing Systems Extended Abstracts, 2022, no. 332, pp. 1–7.
Editorial office address

Executive Secretary of the Editor’s Office

 Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia

  Phone / Fax: + 7 (3822) 701-582

  journal@tusur.ru

 

Viktor N. Maslennikov

Executive Secretary of the Editor’s Office

 Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia

  Phone / Fax: + 7 (3822) 51-21-21 / 51-43-02

Subscription for updates