Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models episode artwork

EPISODE · Oct 31, 2024 · 18 MIN

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models

from Artificial Discourse · host Kenpachi

"Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models," details the development of a new family of multimodal language models (VLMs) called Molmo. Molmo is notable for its open-weight and open-data approach, meaning the model's weights, training data, and code are publicly available. This contrasts with the current trend of proprietary VLMs which keep their models closed. Molmo achieves state-of-the-art performance by utilizing a novel image captioning dataset called PixMo, collected from human annotators using speech-based descriptions. This approach avoids reliance on synthetic data generated by proprietary systems, enabling the creation of performant VLMs without the need for distilling closed models. The authors highlight Molmo's potential for various tasks, including question answering and image-based navigation.

Episode metadata supplied by the publisher feed · Published Oct 31, 2024

Embed this episode

NOW PLAYING

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models

0:00 18:40

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Artificial Discourse?

This episode is 18 minutes long.

When was this Artificial Discourse episode published?

This episode was published on October 31, 2024.

Can I download this Artificial Discourse episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!