
VibeVoice
Open source
FreeFree forever
Overview
VibeVoice is a research framework for advanced text to speech generation capable of producing long-form multi-speaker conversational audio such as podcasts. It uses continuous acoustic and semantic speech tokenizers with a next-token diffusion approach guided by a language model to maintain speaker consistency natural turn taking and high audio fidelity for very long sequences.
Key features
- Multi speaker synthesis
- Long form audio up to 90 minutes
- Continuous speech tokenizers
- Next token diffusion model
- High speaker consistency
- Natural turn taking
- Efficient low frame rate processing
- High fidelity output

