Speech Synthesis Quality Assessment Using Pairwise Comparison and Siamese Neural Networks

DOI: 10.21293/1818-0442-2026-29-1-144-151

Download article in PDF format

JATS xml

Abstract: Relevance. The development of speech synthesis technologies necessitates the need to establish reliable methods for assessing their realism to ensure information security. Purpose. A comparative study of the realism of speech synthesis models (RVC, Coqui TTS, and Fish Speech) based on an integrated approach combining subjective pairwise comparison by auditors and objective classification of a Siamese neural network. Methods. Pairwise comparison of audio recordings by a group of independent auditors was used, as well as a Siamese neural network featuring a trainable classifier based on MFCC spectrograms, trained on authors' original Russian-language dataset. Novelty. For the problem of identifying synthesized speech, a modified Siamese neural network architecture is proposed, in which a concatenation of embeddings followed by processing via a multilayer perceptron is used instead of a fixed distance metric. Results. The results of the study showed that the Fish Speech model exhibits the highest acoustic similarity to the original, causing the maximum number of false identifications. The trained Siamese network detects synthesized speech with an overall accuracy of 84%. Practical significance. The findings demonstrate the applicability of the proposed architecture for independent automated evaluation of synthesis quality and enable further research into protecting speech systems from spoofing attacks, with the possibility of comparison with human perceptual abilities.

Keywords: Retrieval-based Voice Conversion, speech synthesis, neural networks, pairwise comparison, expert evaluation, Text-to-Speech, Fish Speech, Siamese neural network

Funding: This work was supported by the Ministry of Science and High-er Education of Russia: FEWM-2026-0009 (TUSUR).

For citation:
Shelupanov A. A., Kostyuchenko E. Yu., Pogorelov G. L., Nikitin N. D., Maksimov A. V. Speech Synthesis Quality Assessment Using Pairwise Comparison and Siamese Neural Networks. Doklady Tomskogo gosudarstvennogo universiteta sistem upravleniya i radioelektroniki, 2026, vol. 29, no. 1, pp. 144–151. DOI: 10.21293/1818-0442-2026-29-1-144-151

Authors and copyright holders:

  • Shelupanov A. A. , Tomsk State University of Control Systems and Radioelectronics (Tomsk, Russia)
  • Kostyuchenko E. Yu. , Tomsk State University of Control Systems and Radioelectronics (Tomsk, Russia)
  • Pogorelov G. L. , Tomsk State University of Control Systems and Radioelectronics (Tomsk, Russia)
  • Nikitin N. D. , Tomsk State University of Control Systems and Radioelectronics (Tomsk, Russia)
  • Maksimov A. V. , Tomsk State University of Control Systems and Radioelectronics (Tomsk, Russia)

  • 1. Repyuk N.S., Konev A.A. [Software package for the study of speech signals]. Doklady Tomskogo gosudarstvennogo universiteta sistem upravleniya i radioelektroniki, 2025, vol. 28, no. 1, pp. 93–99 (in Russ.).
  • 2. GOST R 50840–95. Peredacha rechi po traktam svyazi. Metody otsenki kachestva, razborchivosti i uznavayemosti [Speech transmission over varies communication channels. Techniques for measurements of speech quality, intelligibility and voice identification]. Moscow, Standartinform Publ., 1996, 18 p. (in Russ.).
  • 3. Kostuchenko E. et al. The evaluation process automation of phrase and word intelligibility using speech recognition systems. International Conference on Speech and Computer, Istanbul, Turkey, Cham, Springer Publ., 2019, pp. 237–246.
  • 4. Semenchuk K.V., Minkevich P.E., Puzyrevskaya A.A. [The essence of the pairwise comparison method in functional-cost analysis]. Science Time, 2016, no. 12 (36), pp. 107–109 (in Russ.).
  • 5. Solomennik A.I., Talanov A.O., Solomennik M.V., Khomitsevich O.G., Chistikov P.G. [Evaluating the quality of synthesized speech: problems and solutions]. Journal of Instrument Engineering, 2013, vol. 56, no. 2, pp. 38–41 (in Russ.).
  • 6. Novokhrestova D.I., Kostyuchenko E.Y., Khodashinsky I.A. [Algorithm and methodology for quantitative assessment of speech signal similarity]. Doklady Tomskogo gosudarstvennogo universiteta sistem upravleniya i radioelektroniki, 2022, vol. 25, no. 3, pp. 45–51 (in Russ.).
  • 7. Sorokin V.N., Vyugin V.V., Tananykin A.A. [Personal voice recognition: an analytical review]. Information Processes, 2012, vol. 12, no. 1, pp. 1–30 (in Russ.).
  • 8. Zaman K., Sah M., Direkoglu C., Unoki M. A Survey of Audio Classification Using Deep Learning. IEEE Access, 2023, vol. 11, pp. 106620–106649.
  • 9. Kataev M.Yu. [Methods of speech command recognition in school information systems]. Speech Technologies, 2024, no. 1, pp. 18–34 (in Russ.).
  • 10. Oganesyan O., Sargsyan D., Malajyan A. [Comparison of voice cloning algorithms in zero-shot and few-shot scenarios]. Proceedings of ISP RAS, 2024, vol. 36, no. 4, pp. 7–16 (in Russ.).
  • 11. Coqui TTS Documentation: XTTS. Available at: https://coqui-tts.readthedocs.io/en/latest/models/xtts.html (accessed: 15 January 2025).
  • 12. Liao S., Wang Y., Li T. et al. Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis. arXiv e-prints, 2024, Art. no. arXiv: 2411.01156.
  • 13. McFee B., Raffel C., Liang D. et al. Librosa: Audio and music signal analysis in Python. Proc. of the 14th Python in Science Conf. (SciPy), Austin, TX, USA, 2015, vol. 8, pp. 18–25.
  • 14. Yang X.K., He L., Dan Q., Zhang W.Q. Voice activity detection algorithm based on long-term pitch information. EURASIP Journal on Audio, Speech, and Music Processing, 2016, no. 1, pp. 1–9.
  • 15. Gourisaria M.K. et al. Comparative analysis of audio classification with MFCC and STFT features using machine learning techniques. Discover Internet of Things, 2024, vol. 4, no. 1, pp. 1–23.
Editorial office address

Executive Secretary of the Editor’s Office

 Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia

  Phone / Fax: + 7 (3822) 701-582

  journal@tusur.ru

 

Viktor N. Maslennikov

Executive Secretary of the Editor’s Office

 Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia

  Phone / Fax: + 7 (3822) 51-21-21 / 51-43-02

Subscription for updates