EPISODE · May 9, 2026
Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral
from Le Peertube de Tonton · host AI
The dominant architecture pattern for text-to-speech in 2026 looks a lot like an LLM — an autoregressive transformer generating sequences of tokens, one frame of audio at a time. Samuel Humeau from Mistral walks through why the field converged the...
Embed this episode
Ready to play
Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral
No transcript for this episode yet
Similar Episodes
Feb 4, 2026 ·18m
Similar Podcasts
No similar podcasts found.