MARS5 an open-source TTS model to replicate performances (from 2-3s of audio reference) in 140+ languages, even for extremely tough prosodic scenarios like sports commentary, movies, anime & more. Join our Discord https://discord.com/invite/ZzsKTAKM today!
MARS5 TTS team, this open-source text-to-speech model looks incredible! The ability to handle tough prosodic scenarios is super impressive. How does MARS5 ensure naturalness and emotional expression across such a wide range of languages and contexts, especially for niche areas like anime and sports commentary? 🎙️
@dmytro_semonov Hey Dmytro, MARS follows a AR<>NAR pipeline with a distinctive NAR model. Our architecture and datasets are differentiated and have allowed us to reach this level of prosody.
Report
Whoa, launching something like MARS5 TTS is like, next level! The way it handles 140+ languages and those tricky prosodic scenarios is mind-blowing.
Quick question tho: how does MARS5 manage the balance between maintaining natural intonation in such a diverse set of languages and ensuring consistency across different dialects and accents? Curious to know how you guys tackled that challenge! 🚀🧠
@p_val We kept a low-resource language-first approach. Helps us capture nuances in each one. Our technology generalizes well even on some unseen languages :)
The ability to replicate performances from just 2-3 seconds of audio reference is fascinating. Could you share more about the technical challenges you faced in achieving this and how you overcame them?
Great launch!
@ramya_pk it's built upon years of research experience that our team at CAMB.AI brings. Definitely not something that was built overnight; lots of blood, sweat and tears haha
MARS5 TTS
MARS5 TTS
MARS5 TTS
MARS5 TTS
MARS5 TTS
MARS5 TTS
Swatle
MARS5 TTS