How AI Learned to Chat About Pictures: Inside the MoshiVis Model episode artwork

EPISODE · Apr 2, 2025 · 14 MIN

How AI Learned to Chat About Pictures: Inside the MoshiVis Model

from GenAI Level UP · host GenAI Level UP

How do you teach a sophisticated speech AI to understand and discuss images, especially when paired image-speech data is rare? This episode unpacks MoshiVis, a new model that achieves just that. We explore the challenges of building Vision-Speech Models and how MoshiVis overcomes them with a unique one-stage training pipeline, synthetic dialogues, and efficient "perceptual augmentation" techniques built upon the Moshi speech LLM. Join us for a deep dive into the tech that lets AI see, speak, and converse fluidly about the visual world.

Episode metadata supplied by the publisher feed · Published Apr 2, 2025

Embed this episode

Ready to play

How AI Learned to Chat About Pictures: Inside the MoshiVis Model

0:00 14:30

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of GenAI Level UP?

This episode is 14 minutes long.

When was this GenAI Level UP episode published?

This episode was published on April 2, 2025.

Can I download this GenAI Level UP episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!