Speech Synthesis Quality Assessment Using Pairwise Comparison and Siamese Neural Networks
DOI: 10.21293/1818-0442-2026-29-1-144-151
DOI: 10.21293/1818-0442-2026-29-1-144-151
Abstract: Relevance. The development of speech synthesis technologies necessitates the need to establish reliable methods for assessing their realism to ensure information security. Purpose. A comparative study of the realism of speech synthesis models (RVC, Coqui TTS, and Fish Speech) based on an integrated approach combining subjective pairwise comparison by auditors and objective classification of a Siamese neural network. Methods. Pairwise comparison of audio recordings by a group of independent auditors was used, as well as a Siamese neural network featuring a trainable classifier based on MFCC spectrograms, trained on authors' original Russian-language dataset. Novelty. For the problem of identifying synthesized speech, a modified Siamese neural network architecture is proposed, in which a concatenation of embeddings followed by processing via a multilayer perceptron is used instead of a fixed distance metric. Results. The results of the study showed that the Fish Speech model exhibits the highest acoustic similarity to the original, causing the maximum number of false identifications. The trained Siamese network detects synthesized speech with an overall accuracy of 84%. Practical significance. The findings demonstrate the applicability of the proposed architecture for independent automated evaluation of synthesis quality and enable further research into protecting speech systems from spoofing attacks, with the possibility of comparison with human perceptual abilities.
Keywords: Retrieval-based Voice Conversion, speech synthesis, neural networks, pairwise comparison, expert evaluation, Text-to-Speech, Fish Speech, Siamese neural network
Funding: This work was supported by the Ministry of Science and High-er Education of Russia: FEWM-2026-0009 (TUSUR).
For citation:
Shelupanov A. A., Kostyuchenko E. Yu., Pogorelov G. L., Nikitin N. D., Maksimov A. V. Speech Synthesis Quality Assessment Using Pairwise Comparison and Siamese Neural Networks. Doklady Tomskogo gosudarstvennogo universiteta sistem upravleniya i radioelektroniki, 2026, vol. 29, no. 1, pp. 144–151. DOI: 10.21293/1818-0442-2026-29-1-144-151
Authors and copyright holders:
Executive Secretary of the Editor’s Office
Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia
Phone / Fax: + 7 (3822) 701-582
Viktor N. Maslennikov
Executive Secretary of the Editor’s Office
Editor’s Office: 40 Lenina Prospect, Tomsk, 634050, Russia
Phone / Fax: + 7 (3822) 51-21-21 / 51-43-02