Embodied AI 101 podcast artwork

PODCAST · technology

Embodied AI 101

Stay in the loop on research in AI and physical intelligence.

Publisher-supplied feed metadata · PodParley refreshed Jun 13, 2026 · Source feed

  1. 105

    Imagining Locomotion: Learning a Neural World Model for Legged Robots

    A neural dynamics world model paired with model-free RL policies for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction and imagined rollouts that outperform pure model-based RL in prediction accuracy, policy learning, and sim-to-real transfer.

  2. 104

    One Model to See, Plan, and Act: Introducing EO-1 for Embodied AI

    A 3B parameter unified decoder-only transformer that interleaves vision, text, and action tokens for perception, planning, reasoning, and control in a single model. Trained on the 1.5M-sample EO-1.5M dataset with strong results across manipulation tasks and benchmarks.

  3. 103

    Video-Action Models for Robot Learning

    Introduces Video-Action Models (VAMs) that leverage pretrained internet-scale video models such as Cosmos-Predict2 as backbones instead of VLMs, paired with a flow-matching action decoder. Claims approximately 10x sample efficiency gains over standard vision-language-action models on real-world pick-and-place tasks.

  4. 102

    FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control

    Introduces FastDSAC, a method that scales maximum entropy reinforcement learning to high-dimensional humanoid control tasks, improving sample efficiency and policy performance.

  5. 101

    A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

    Proposes a minimalist pipeline using retargeting-guided RL for dexterous manipulation in humanoid robots. Focuses on simplifying the training recipe while maintaining strong manipulation performance.

  6. 100

    DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation

    DexNDM bridges the sim-to-real gap for stable in-hand rotation of complex objects, enabling learning from biased real-world data without requiring any successful demonstrations.

  7. 99

    Learning Multi-Modal Trajectory Policies for Data-Efficient Robotic Manipulation

    Proposes learning multi-modal trajectory policies to improve data efficiency in robotic manipulation tasks.

  8. 98

    ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning

    A zero-shot workflow reasoning approach for agentic control of embodied manipulation, enabling robots to perform complex tasks without task-specific training.

  9. 97

    Let the Body Follow: Coupled Egocentric Control for Whole-Body Robot Teleoperation

    A coupled egocentric control approach for whole-body robot teleoperation that coordinates body and limb movements from an egocentric perspective, submitted to Humanoids 2026.

  10. 96

    GigaWorld-Policy-0.5: An Efficient Mixture-of-Transformers Policy for Real-Time Robot Control

    A Mixture-of-Transformers robot policy that achieves 85 ms inference on RTX 4090 by separating visual dynamics from action generation. Demonstrates efficient real-time robot policy execution through architectural decomposition.

  11. 95

    Xiaomi-Robotics-1: A Scalable Vision-Language-Action Foundation Model for Mobile Manipulation

    A scalable vision-language-action (VLA) foundation model pretrained on over 100k hours of real-world manipulation trajectories, enabling out-of-the-box mobile manipulation capabilities.

  12. 94

    Local Policies Enable Zero-shot Long-Horizon Manipulation

    Proposes training local policies in simulation that transfer zero-shot to real-world long-horizon robotic manipulation tasks. Addresses the challenge of generalizing learned behaviors across extended task sequences.

  13. 93

    Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids

    Trains perception-driven dexterous manipulation policies in simulation for zero-shot real-world transfer on humanoid robots. Focuses on bridging the sim-to-real gap for vision-based dexterous tasks.

  14. 92

    UniTracker: A Universal Motion Tracking Framework for Humanoid Robots

    A framework enabling humanoid robots to execute diverse motions within physical limits, with video demonstrations and open-source code. Targets generalizable motion tracking across varied humanoid morphologies.

  15. 91

    ViTacFormer: Dexterous Manipulation via Active Vision and High-Resolution Touch Sensing

    A pipeline combining active vision and high-resolution tactile sensing for dexterous manipulation on high-DoF robot hands, achieving approximately 2.5 minutes of continuous autonomous control on real-world tasks such as food preparation.

  16. 90

    Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language

    Uses language and demonstrations to learn which aspects of behavior matter for reward in imitation learning, enabling 5× faster learning by avoiding copying irrelevant motion details. Accepted at ICRA 2026.

  17. 89

    FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

    Steers frozen robot foundation models including VLAs, diffusion policies, and world models using human corrections via action inversion in latent space without fine-tuning the base policy. Outperforms supervised fine-tuning and latent RL with only 5–20 human interventions.

  18. 88

    RoboTTT: End-to-End Robotic Assembly with Test-Time Training

    Presents a single end-to-end policy that performs precise, unscripted assembly of complex objects, measuring every grasp and alignment without speed-ups or human intervention.

  19. 87

    Xiaomi-Robotics-U0: A 38B World-Foundation Model for Unified Embodied Perception and Synthesis

    38B autoregressive world foundation model unifying text-to-image, multi-view scene generation, embodied transfer, and video generation; ranks #1 on World Arena.

  20. 86

    SpatialPoint: Spatial-Aware Point Prediction for Embodied Localization

    A spatial-aware point prediction method for embodied localization tasks. Addresses grounding and spatial reasoning in embodied AI settings.

  21. 85

    Multi-AUV Scene-Adaptive Embodied Intelligence for Multi-Target Tracking

    An embodied AI framework for multi-AUV (Autonomous Underwater Vehicle) multi-target tracking in ad-hoc underwater networks. Incorporates scene-adaptive policies for dynamic underwater environments.

  22. 84

    Speedup Patch: Learning a Plug-and-Play Policy to Accelerate Embodied Manipulation

    Proposes a plug-and-play policy module designed to accelerate existing embodied manipulation policies without retraining them from scratch.

  23. 83

    R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation

    Introduces a real-time 3D-aware policy framework for embodied manipulation tasks, enabling spatially grounded action generation.

  24. 82

    VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation

    Presents a vision-language-action model grounded in 3D Gaussian representations that integrates geometric and semantic awareness for improved robotic manipulation.

  25. 81

    SKooP: Symmetric Koopman Predictions for Legged Robot Reinforcement Learning

    Leverages symmetric Koopman operator predictions within a reinforcement learning framework to achieve faster training and better generalization for legged robot locomotion.

  26. 80

    RobotTT: A Tactile Transformer for Dexterous Manipulation

    A tactile foundation model and sim-to-real pipeline designed for dexterous manipulation tasks. Combines transformer-based tactile sensing with a sim-to-real transfer approach.

  27. 79

    TAC-LOCO: Unified Whole-Body Control for Contact-Aware Locomotion and Manipulation

    A tactile-informed whole-body loco-manipulation controller for quadrupedal robots that unifies locomotion and manipulation using tactile sensing. Enables compliant and contact-aware whole-body control.

  28. 78

    TACTIC: Contact-Centric Control for Whole-Arm Manipulation

    A contact-centric whole-arm manipulation framework that conditions control on both tactile and visual inputs. Targets dexterous manipulation tasks requiring rich contact feedback.

  29. 77

    SIEVE: Structure-Aware Data Selection for Imitation Learning with VLAs

    SIEVE introduces structure-aware data selection specifically for imitation learning with Vision-Language-Action models, aiming to improve training efficiency and policy quality.

  30. 76

    DexMachina: RL for Long-Horizon Bimanual Dexterous Policies

    An RL algorithm that learns long-horizon bimanual dexterous policies for any robot hand from a single human demonstration, emphasizing generalization across hands, objects, and complex motions.

  31. 75

    Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

    A unified platform comprising GE-Base (video diffusion model trained on 1M+ manipulation episodes), GE-Act (flow-matching action model), and GE-Sim (neural world simulator for closed-loop control). All code, models, and benchmarks are open-sourced.

  32. 74

    RoboTTT: Test-Time Training for Visuomotor Policies

    Introduces test-time training (TTT) inside the policy to natively scale visuomotor context to 8K timesteps at constant inference cost, enabling one-shot imitation from human video demos, online self-recovery from errors, and long-horizon assembly tasks.

  33. 73

    HY-Embodied-VLM-1.0: Efficient Physical-World Agents

    An embodied vision-language-action model with released weights, code, and paper targeting robotic manipulation and embodied AI tasks.

  34. 72

    Learning Unified Force and Position Control for Legged Loco-Manipulation

    Introduces a unified RL policy that jointly handles force and position control on quadrupedal and humanoid robots without force sensors, enabling position+force tracking, compliant behaviors, and force-aware imitation learning for contact-rich tasks.

  35. 71

    HapticVLA: Extending Vision-Language-Action Models to Contact-Rich Tasks Without Touch Sensors

    Enables contact-rich robotic manipulation using a VLA model trained with tactile sensing data but requiring no tactile input at inference time. Distills haptic knowledge into the vision-language-action policy.

  36. 70

    AnoleVLA: Lightweight VLA with Deep State Space Models for Mobile Manipulation

    Proposes a lightweight VLA model leveraging deep state space models (SSMs) instead of transformers for efficient mobile manipulation. Targets resource-constrained deployment scenarios with competitive performance.

  37. 69

    ABot-N1: Visual Language Navigation Foundation Model

    Decouples cognition from control in a VLM-based navigation policy, delivering 35% POI arrival gains and over 92% success rates in complex indoor and outdoor scenes.

  38. 68

    ABot-AgentOS: A General-Purpose Robotic Agent Operating System

    Provides scene-conditioned planning, context-isolated skill execution, multi-modal memory, and self-evolution capabilities for long-horizon embodied tasks.

  39. 67

    Trust Region Policy Distillation (TOP-D)

    Transforms unstable on-policy distillation into a stable training paradigm via dynamic proximal teacher construction, improving sample efficiency without additional computational overhead.

  40. 66

    HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

    A new humanoid whole-body control method that distills complementary teachers for agentic task-space control on humanoids. The approach enables robust whole-body coordination for complex manipulation tasks.

  41. 65

    EVA-Client: A Unified Framework for Real-Robot Policy Iteration

    A unified open framework for real-robot policy iteration that integrates teleoperation data collection, model training, deployment, and evaluation into a single closed-loop pipeline.

  42. 64

    UniVR-34B: A Vision-Only Foundation Model for Physical Tasks

    First large-scale model to learn complex physical dynamics, visual reasoning, and long-horizon planning directly from visual demonstrations without text chains; released with 310k SFT and 3k RL samples across 16 sources alongside the VR-X benchmark.

  43. 63

    CLAP: Converting Vision-Language Models into Vision-Language-Action Models via Language-Prompted Actions

    Converts any pretrained vision-language model into a vision-language-action model with zero architectural changes by prepending natural-language action descriptions; reaches 90.8% success on LIBERO with a 2B model after less than 6 hours of post-training.

  44. 62

    InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action

    InternVLA-A15 is a new Vision-Language-Action (VLA) model released by InternRobotics, targeting robotic manipulation and control tasks. The model is accompanied by a Hugging Face collection and an associated technical paper.

  45. 61

    MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

    Proposes foundation world-action models targeting real-time humanoid loco-manipulation, integrating world modeling with action generation for unified locomotion and manipulation control.

  46. 60

    On-Device Diffusion Transformer Policy for Efficient Robot Manipulation

    LightDP introduces a framework to accelerate Diffusion Transformer-based policies for real-time, on-device deployment in robot manipulation tasks, achieving efficiency for visuomotor control without sacrificing performance. Accepted to ICCV 2025.

  47. 59

    Robotic World Model: Neural Dynamics for Locomotion

    A neural dynamics world model paired with model-free RL for quadruped and humanoid locomotion in IsaacLab, enabling long-horizon autoregressive prediction with policies trained in imagined rollouts that outperform model-based baselines on sim-to-real transfer.

  48. 58

    LingBot-Vision: Self-Supervised ViT with Masked Boundary Modeling

    A self-supervised ViT backbone pretrained with masked boundary modeling for dense spatial perception, achieving strong zero-shot results on depth estimation, segmentation, and embodied tasks.

  49. 57

    LaMem-VLA: Latent-Memory-Native VLA Framework

    A VLA framework that curates experience into short- and long-term memory vaults, condenses them to latent tokens, and integrates directly into VLA reasoning, evaluated on SimplerEnv and LIBERO benchmarks.

  50. 56

    LingBot-VA 2.0: Native Video-Action Foundation Model for Robot Control

    A native video-action foundation model pretrained from scratch with a semantic visual-action tokenizer and foresight reasoning, enabling real-time robot control at ≤150 Hz on consumer GPUs without relying on retrofitted VLMs or video generators.

Type above to search every episode's transcript for a word or phrase. Matches are scoped to this podcast.

Searching…

We're indexing this podcast's transcripts for the first time — this can take a minute or two. We'll show results as soon as they're ready.

No matches for "" in this podcast's transcripts.

Showing of matches

No topics indexed yet for this podcast.

Loading reviews...

ABOUT THIS SHOW

Stay in the loop on research in AI and physical intelligence.

HOSTED BY

Shaoqing Tan

CATEGORIES

Frequently Asked Questions

How many episodes does Embodied AI 101 have?

Embodied AI 101 currently has 50 episodes available on PodParley. New episodes are automatically indexed when they're published to the podcast feed.

What is Embodied AI 101 about?

Stay in the loop on research in AI and physical intelligence.

How often does Embodied AI 101 release new episodes?

Embodied AI 101 has 50 episodes. Check the episode list to see recent publication dates and frequency.

Where can I listen to Embodied AI 101?

You can listen to Embodied AI 101 on PodParley by clicking any episode. We provide an embedded audio player for direct listening, and you can also subscribe via your preferred podcast app using the RSS feed.

Who hosts Embodied AI 101?

Embodied AI 101 is created and hosted by Shaoqing Tan.
URL copied to clipboard!