| ||||
| ||||
![]() Title:Adapting a Multilingual Zero-Shot Text-to-Speech Model for High-Fidelity Synthesis in Polish Despite Limited Training Data Conference:ISD2026 Tags:Cross-Lingual Fine-Tuning, Low-Resource, Polish TTS, Speech synthesis and Zero-Shot Voice Cloning Abstract: Modern Text-to-Speech (TTS) systems demand massive training corpora, creating a severe bottleneck for low-resource languages like Polish. Furthermore, relying on commercial cloud-based TTS platforms introduces critical risks to data privacy and latency. This paper presents an effective cross-lingual fine-tuning strategy that adapts a pre-trained, English-centric non-autoregressive architecture (F5-TTS) for high-fidelity, locally deployable Polish speech synthesis. To overcome the open-source data deficit, we utilize a hybrid dataset combining crowdsourced amateur recordings with a proprietary professional studio corpus, which notably includes the deep and distinct voice of the prominent Polish actor Piotr Fronczewski. Objective metrics (UTMOS, WER) and subjective MUSHRA evaluations demonstrate that our fine-tuned model significantly outperforms the current local state-of-the-art baseline (XTTS-v2), achieving highly natural, artifact-free zero-shot voice cloning. Coupled with a lightweight on-premise GUI, this work provides a robust solution for Polish speech synthesis, serving as a blueprint for adapting a foundational TTS architecture to other underrepresented languages. Adapting a Multilingual Zero-Shot Text-to-Speech Model for High-Fidelity Synthesis in Polish Despite Limited Training Data ![]() Adapting a Multilingual Zero-Shot Text-to-Speech Model for High-Fidelity Synthesis in Polish Despite Limited Training Data | ||||
| Copyright © 2002 – 2026 EasyChair |
