EPISODE · Apr 2, 2025 · 14 MIN
How AI Learned to Chat About Pictures: Inside the MoshiVis Model
from GenAI Level UP · host GenAI Level UP
How do you teach a sophisticated speech AI to understand and discuss images, especially when paired image-speech data is rare? This episode unpacks MoshiVis, a new model that achieves just that. We explore the challenges of building Vision-Speech Models and how MoshiVis overcomes them with a unique one-stage training pipeline, synthetic dialogues, and efficient "perceptual augmentation" techniques built upon the Moshi speech LLM. Join us for a deep dive into the tech that lets AI see, speak, and converse fluidly about the visual world.
Embed this episode
Ready to play
How AI Learned to Chat About Pictures: Inside the MoshiVis Model
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.