GPT-SoVITS
Train a good text-to-speech model from 1 minute of voice data. Few-shot voice cloning.
TL;DR · 30-second scan
GPT-SoVITS (Python) — Train a good text-to-speech model from 1 minute of voice data. Few-shot voice cloning.
You want your own voice cloned once and reused forever with zero API cost.
Voice & Audio Generation
The most-starred voice cloning repo on GitHub for a reason. One minute of clean voice data trains a usable TTS clone. Output quality rivals paid tools like ElevenLabs at zero recurring cost. The UI is functional-but-ugly and setup takes 30-60 minutes if you have never touched Python audio libs. Best for creators building content in a single consistent voice (podcasts, faceless YouTube narration, personal AI assistants).
You want your own voice cloned once and reused forever with zero API cost.
You need instant cloning inside a hosted app — ElevenLabs API is still the fastest to integrate.
Add this badge to your README to show your project is curated on StackPicks. Free, lightweight (180×28 SVG), and gives your visitors a one-click way to see honest take + alternatives.
[](https://stackpicks.dev/repo/gpt-sovits)
<a href="https://stackpicks.dev/repo/gpt-sovits"><img src="https://stackpicks.dev/api/badge/gpt-sovits" alt="Featured on StackPicks" width="180" height="28" /></a>
Are you the maintainer of RVC-Boss/GPT-SoVITS? Add the badge and we'll feature your project in the weekly curator newsletter.