X-LLM: Treating Multimodalities as Foreign Languages episode artwork

EPISODE · Jun 23, 2026

X-LLM: Treating Multimodalities as Foreign Languages

from AI Post Transformers

This episode explores X-LLM, a 2023 system that treats images, video, and speech as foreign languages a frozen ChatGLM can learn to read through learned modality-to-language bridges. It breaks down the paper’s architecture, including Q-Former-based visual adapters and a separate speech pipeline with continuous integrate-and-fire modules, to show how three sensory routes feed a single dialogue model instead of one end-to-end multimodal transformer. The discussion argues that X-LLM mattered less as proof of a universal multimodal theory than as a practical open-model recipe shaped by 2023 compute limits, with its Chinese-language backbone playing a real methodological role rather than serving as background context. Listeners get a sharp comparison between this bridge-based approach and later end-to-end systems such as GPT-4o and Gemini 1.5, making the episode useful for understanding how modern multimodal assistants actually evolved. Sources: 1. X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages — Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang, Jing Shi, Shuang Xu, Bo Xu, 2023 http://arxiv.org/abs/2305.04160 2. Multimodal Few-Shot Learning with Frozen Language Models — Maria Tsimpoukelli, Jacob Menick, Oriol Vinyals, Felix Hill, 2021 https://scholar.google.com/scholar?q=Multimodal+Few-Shot+Learning+with+Frozen+Language+Models 3. Flamingo: a Visual Language Model for Few-Shot Learning — Jean-Baptiste Alayrac, Jeff Donahue, Karen Simonyan, Oriol Vinyals, Andrew Zisserman, 2022 https://scholar.google.com/scholar?q=Flamingo%3A+a+Visual+Language+Model+for+Few-Shot+Learning 4. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models — Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi, 2023 https://scholar.google.com/scholar?q=BLIP-2%3A+Bootstrapping+Language-Image+Pre-training+with+Frozen+Image+Encoders+and+Large+Language+Models 5. SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities — Dong Zhang, Shimin Li, Xin Zhang, Xipeng Qiu, 2023 https://scholar.google.com/scholar?q=SpeechGPT%3A+Empowering+Large+Language+Models+with+Intrinsic+Cross-Modal+Conversational+Abilities 6. Visual Instruction Tuning — Haotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae Lee, 2023 https://scholar.google.com/scholar?q=Visual+Instruction+Tuning 7. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models — Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, Mohamed Elhoseiny, 2023 https://scholar.google.com/scholar?q=MiniGPT-4%3A+Enhancing+Vision-Language+Understanding+with+Advanced+Large+Language+Models 8. PaLM-E: An Embodied Multimodal Language Model — Danny Driess et al., 2023 https://scholar.google.com/scholar?q=PaLM-E%3A+An+Embodied+Multimodal+Language+Model 9. CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition — Linhao Dong, Bo Xu, 2019 https://scholar.google.com/scholar?q=CIF%3A+Continuous+Integrate-and-Fire+for+End-to-End+Speech+Recognition 10. VL-JEPA: Joint Embedding Predictive Architecture for Vision-language — Delong Chen et al., 2025 https://scholar.google.com/scholar?q=VL-JEPA%3A+Joint+Embedding+Predictive+Architecture+for+Vision-language 11. TokenPacker: Efficient Visual Projector for Multimodal LLM — Wentong Li et al., 2024 https://scholar.google.com/scholar?q=TokenPacker%3A+Efficient+Visual+Projector+for+Multimodal+LLM 12. Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM — Donghwan Chi et al., 2025 https://scholar.google.com/scholar?q=Slot-MLLM%3A+Object-Centric+Visual+Tokenization+for+Multimodal+LLM 13. Auto-Encoding Morph-Tokens for Multimodal LLM — Kaihang Pan et al., 2024 https://scholar.google.com/scholar?q=Auto-Encoding+Morph-Tokens+for+Multimodal+LLM 14. ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention — Wenjie Liu et al., 2026 https://scholar.google.com/scholar?q=ViCA%3A+Efficient+Multimodal+LLMs+with+Vision-Only+Cross-Attention 15. F-LMM: Grounding Frozen Large Multimodal Models — Size Wu et al., 2024 https://scholar.google.com/scholar?q=F-LMM%3A+Grounding+Frozen+Large+Multimodal+Models 16. MultiModal-GPT: A Vision and Language Model for Dialogue with Humans — Tao Gong et al., 2023 https://scholar.google.com/scholar?q=MultiModal-GPT%3A+A+Vision+and+Language+Model+for+Dialogue+with+Humans 17. AI Post Transformers: UniVideo: Unified Video Understanding, Generation, and Editing — Hal Turing & Dr. Ada Shannon, Sat, https://podcast.do-not-panic.com/episodes/univideo-unified-video-understanding-generation-and-editing/ 18. AI Post Transformers: DeepSeek-OCR: Contexts Optical Compression — Hal Turing & Dr. Ada Shannon, Sat, https://podcast.do-not-panic.com/episodes/deepseek-ocr-contexts-optical-compression/ Interactive Visualization: X-LLM: Treating Multimodalities as Foreign Languages

Episode metadata supplied by the publisher feed · Published Jun 23, 2026

Embed this episode

NOW PLAYING

X-LLM: Treating Multimodalities as Foreign Languages

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 23, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!