HumanOmni:以人为中心的视频理解大型视觉语音语言模型 episode artwork

EPISODE · Feb 7, 2025 · 8 MIN

HumanOmni:以人为中心的视频理解大型视觉语音语言模型

from AI Podcast · host weedge

深入探讨HumanOmni,一个为理解以人为中心的场景而设计的多模态大型语言模型。我们讨论了其数据集构建、模型架构以及在情感识别、面部表情理解和动作理解等任务上的表现。

Episode metadata supplied by the publisher feed · Published Feb 7, 2025

Embed this episode

Ready to play

HumanOmni:以人为中心的视频理解大型视觉语音语言模型

0:00 8:05

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of AI Podcast?

This episode is 8 minutes long.

When was this AI Podcast episode published?

This episode was published on February 7, 2025.

Can I download this AI Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!