All Episodes
Embodied AI 101 — 230 episodes
Imagining Locomotion: Learning a Neural World Model for Legged Robots
One Model to See, Plan, and Act: Introducing EO-1 for Embodied AI
Video-Action Models for Robot Learning
FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation
DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation
Learning Multi-Modal Trajectory Policies for Data-Efficient Robotic Manipulation
ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning
Let the Body Follow: Coupled Egocentric Control for Whole-Body Robot Teleoperation
GigaWorld-Policy-0.5: An Efficient Mixture-of-Transformers Policy for Real-Time Robot Control
Xiaomi-Robotics-1: A Scalable Vision-Language-Action Foundation Model for Mobile Manipulation
Local Policies Enable Zero-shot Long-Horizon Manipulation
Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids
UniTracker: A Universal Motion Tracking Framework for Humanoid Robots
ViTacFormer: Dexterous Manipulation via Active Vision and High-Resolution Touch Sensing
Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language
FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space
RoboTTT: End-to-End Robotic Assembly with Test-Time Training
Xiaomi-Robotics-U0: A 38B World-Foundation Model for Unified Embodied Perception and Synthesis
SpatialPoint: Spatial-Aware Point Prediction for Embodied Localization
Multi-AUV Scene-Adaptive Embodied Intelligence for Multi-Target Tracking
Speedup Patch: Learning a Plug-and-Play Policy to Accelerate Embodied Manipulation
R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation
SKooP: Symmetric Koopman Predictions for Legged Robot Reinforcement Learning
RobotTT: A Tactile Transformer for Dexterous Manipulation
TAC-LOCO: Unified Whole-Body Control for Contact-Aware Locomotion and Manipulation
TACTIC: Contact-Centric Control for Whole-Arm Manipulation
SIEVE: Structure-Aware Data Selection for Imitation Learning with VLAs
DexMachina: RL for Long-Horizon Bimanual Dexterous Policies
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
RoboTTT: Test-Time Training for Visuomotor Policies
HY-Embodied-VLM-1.0: Efficient Physical-World Agents
Learning Unified Force and Position Control for Legged Loco-Manipulation
HapticVLA: Extending Vision-Language-Action Models to Contact-Rich Tasks Without Touch Sensors
AnoleVLA: Lightweight VLA with Deep State Space Models for Mobile Manipulation
ABot-N1: Visual Language Navigation Foundation Model
ABot-AgentOS: A General-Purpose Robotic Agent Operating System
Trust Region Policy Distillation (TOP-D)
HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers
EVA-Client: A Unified Framework for Real-Robot Policy Iteration
UniVR-34B: A Vision-Only Foundation Model for Physical Tasks
CLAP: Converting Vision-Language Models into Vision-Language-Action Models via Language-Prompted Actions
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action
MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation
On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
Robotic World Model: Neural Dynamics for Locomotion
LingBot-Vision: Self-Supervised ViT with Masked Boundary Modeling
LaMem-VLA: Latent-Memory-Native VLA Framework
LingBot-VA 2.0: Native Video-Action Foundation Model for Robot Control
GigaWorld-1: Large-Scale World Model for Robot Policy Evaluation
Learning Unified Force and Position Control for Legged Loco-Manipulation
Robix: Unified Vision-Language Model for Robotic Reasoning and Planning
EmbodiedOneVision (EO-1): A Unified Decoder-Only Transformer for General Robot Control
mimic-video: Video-Action Models for Robot Learning
Lowering the Barrier: The GEM Open-Source Arm
RynnWorld-4D: A 4D Embodied World Model for Bimanual Robot Control
RynnWorld-Teleop: Digital Teleoperation via World Model Rendering
VLA-Corrector: Adaptive Action Horizons through Latent Visual Monitoring
Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization
Current as Touch: Proprioceptive Contact Feedback for Compliant Dexterous Manipulation
DART: One-Shot VLA Policy Adaptation via Weight-Space Arithmetic
Does VLA Even Know the Basics? Act2Answer Benchmark
Contact-Grounded Policy: Dexterous Visuotactile Policy with Generative Contact Grounding
Freeform Preference Learning (FPL) for Robotic Manipulation
Orca: The World is in Your Mind
Qwen-RobotNav: A Scalable Unified Navigation Model for Agentic Robotics
Scaling Robot Skills from Cheap Human Videos
ABC: An Open Behavior Cloning Stack for Bimanual Manipulation
Qwen-RobotManip: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
ASPIRE: Automated Skill Discovery for Robotics
SERF: 4D Latent Mapping for Long-Horizon Mobile Manipulation
ViserDex: Visual Sim-to-Real for Robust Dexterous In-Hand Reorientation
DexSkin: A High-Coverage, Conformable "Electronic Skin" for Robot Fingers
EBench: A Diagnostic Benchmark for Generalist Manipulation Policies
VITRA: A Foundation for Dexterous VLA via Human Video Pretraining
DexWM: A Dexterous Manipulation World Model from Human Videos
PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning
Continual Robot Policy Learning via Variational Neural Dynamics
PhysisForcing: Physics-Reinforced World Models for Robotic Manipulation
Translation as a Bridging Action
Play2Perfect: Dexterous Play Pretraining for Precise Assembly
Dexora: Open-Source VLA for High-DoF Bimanual Dexterity
WorldVLA: Towards Autoregressive Action World Model
HumDex: Humanoid Dexterous Manipulation Made Easy
ForceBand: Learning Forceful Manipulation with sEMG
In-Context World Modeling for Robotic Control
WOLF-VLA: Vision-Language-Action for Humanoid Walking
Motion-Focused Latent Action for Cross-Embodiment VLA from Human Videos
ManiFlow: Manipulation via Rectified Flow
RL-100: Toward Highly Reliable Real-World Robot Reinforcement Learning
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
Bi-HIL: Bilateral Control-Based Multimodal Hierarchical Imitation Learning for Long-Horizon Contact-Rich Manipulation
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning
ConstrainedMimic: Safe Humanoid Robot Motion Tracking
REAL: Robust Extreme Agility via Spatio-Temporal Policy Learning and Physics-Guided Filtering
HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Learning for Contact-Rich Manipulation
SARM2 + SPIRAL: Multi-Task Reward Models and RL Refinement for Long-Horizon Dexterous Manipulation
Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm VLA Systems
ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation
Playful Agentic Robot Learning
Learning Unified Force and Position Control for Legged Loco-Manipulation
Robots that Collaborate: Sequential Asymmetric Imitation for Learning Coupled Robot Policies
AstraBrain-WBC 0.5: A Humanoid Robot Cerebellum Foundation Model
SRL: Combining SLIP Model and Reinforcement Learning for Agile Robotic Jumping
DataClaw0: Agentic Tailoring for Raw Multimodal Streams
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
VERA: Video-to-Action World Model Policy
GEN-1: Scaled Dexterous Manipulation Foundation Model
Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot Manipulation
Cortical Policy: A Dual-Stream View Transformer for Robotic Manipulation
VisualClaw: A Self-Evolving Wearable Vision Agent
Kairos: A Native World Model Stack for Physical AI
DragMesh-2: A Contact-Driven Framework for Dexterous Hand–Object Interaction
Guava: A Universal Harness for Robot Manipulation
Geometric Action Model for Robot Policies
ENPIRE: Physical AutoResearch with a Fleet of 8 Robots
MolmoAct2: An Open Foundation Model for Real-World Robotics
Hy-Embodied-0.5-VLA: A Massive Bimanual Teleoperation Dataset for Vision-Language-Action
Q-Guided Flow: Test-Time Gradient Guidance of Flow Policies
Flow Reversal Steering: Guiding Diffusion-Based Robot Policies with High-Level Reasoning
Test-Time Compute Scaling for Robot Policies (DIRECT)
LabVLA: Bringing Vision-Language-Action to the Chemistry Lab
Humanoid-GPT: A Foundation Model for Zero-Shot Humanoid Control
CHORUS: Decentralized Multi-Robot Collaboration with a Single Shared VLA Model
RISE: Self-Improving Robot Policy with Compositional World Model
EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot Control
Robix: A Unified Model for Robot Interaction, Reasoning and Planning
Robotic World Model: Learning to Simulate for Robust Robot Control
AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization
ArtiFixer: Few-Step Diffusion for 3D Scene Reconstruction
Deployment-Time Memorization in Foundation-Model Agents
Adversarial Machine Learning: Taxonomy, Threat Models, and Mitigation Strategies in Deep Neural Networks
SoCRATES: Evaluating LLM Mediators in Conflict Scenarios
Unembedding Matrix as a Feature Lens: Unlocking Better Text Embeddings
LeanMarathon: Autonomous Formalization of Math Proofs on Erdős Problems
Deep Research Agents: Survey and Roadmap for Autonomous AI Research
Cosmos 3: Omnimodal World Models for Physical AI
Humanoid-GPT: GPT-Style Transformer for Zero-Shot Dynamic Humanoid Control
Bending Paper, Shaping Dexterity: The Robotic Origami Challenge
GraspGen-X: A Foundation Model for Zero-Shot 6-DoF Grasping
When Does Deep RL Beat Calibrated Baselines?
Training Deep Networks as Random Effects: An Optimization–Inference Duality
Generative Depth Supervision for Embodied Vision-Language Models
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
LT2: Linear-Time Looped Transformers
One Learning Rate Doesn't Fit All: Layerwise Spectral Scheduling for Transformers
SimToolReal: Procedural Tool Generation and a Universal Objective for Zero-Shot Tool Manipulation
Robometer and the Future of Robotic Reward Modeling
Qwen-VLA: A Generalist Vision–Language–Action Robot Model
EXPO-FT: Sample-Efficient Reinforcement Learning Fine-Tuning for Vision-Language-Action Models
RoboMeter: Learning Dense Rewards from Successes and Failures
MobileGym: A Controllable, Parallel Sandbox for Mobile GUI Agents
ANY2ANY: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking
TriSplat: Feed-Forward 3D Reconstruction with Triangulated Meshes
MIKASA-Robo-VLA: A Memory-Intensive Benchmark for Vision-Language-Action Robotics
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
Bimanual Pegboard Manipulation: A Benchmark for Vision-Language-Action Models
FutureSim: Replaying Real-World Events to Evaluate AI Forecasting Agents
AgentFloor: A Benchmark for Long-Horizon Agent Planning
AlexNet: The Deep Convolutional Network That Transformed Vision
A Few Useful Things to Know About Machine Learning
SimToolReal: A Universal Dexterous Tool-Use Policy
Mimic-Video: Learning Physics Priors from Web-Scale Video for Robot Dexterity
Deep Residual Learning for Image Recognition (ResNet)
Attention Is All You Need – The Transformer Revolution
NVIDIA Cosmos: World Foundation Models for Physical AI
LATENT: Teaching a Humanoid to Play Tennis from Imperfect Data
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
World Action Models: The Next Frontier in Embodied AI
Training a Whole-Body Control Foundation Model
DexJoCo: A Unified Benchmark for Task-Oriented Dexterous Manipulation
MMSkills: Building Multimodal Skill Libraries for Visual Agents
PhysBrain 1.0 VLA (TwinBrainVLA): Dual-Brain Vision-Language-Action with Physics-Grounded Learning
MolmoAct2-LIBERO: An Open Vision-Language-Action Model for Robotics
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Diffusion Transformers
WildClawBench: A Real-World, Long-Horizon Benchmark for AI Agents
MCP-Cosmos: Bring Your Own World Model
OpenAI o1: Teaching LLMs to Think Slow and Deep
The Llama 3 Herd of Models
LATENT: Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data
AnyFlow: Any-Step Video Diffusion for Predictive World Modeling
# Robotics: The Endgame
Claw-Eval: Toward Trustworthy and Transparent Evaluation of Autonomous Agents
LIBERO-Para: Paraphrase Robustness in Robotic Manipulation
YOR: Your Own Mobile Manipulator for Generalizable Robotics
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
Accelerating Video World Models: From Generative Videos to Real-Time Simulators
From Tokens to Thoughts: Continuous Latent Reasoning in Large Models and Robot Control
CaP-X: Coding Agents for Physical eXecution
DoRA: Weight-Decomposed Low-Rank Adaptation
AI Model Collapse: What Happens When AI Trains on Its Own Outputs
PhAIL: Benchmarking Vision-Language-Action Models on Real-World Bin-Picking
Co-training Large Behavior Models: Data Modalities and Training Strategies for Robot Manipulation
HyDRA: Hybrid Memory for Dynamic Video World Models
# WildWorld: Dynamic World Modeling with Actions and Explicit State
Omni-WorldBench: Evaluating Interactive 4D World Models
SIMART: From Static Meshes to Sim-Ready Articulated Models
EgoSim: An Egocentric World Simulator for Embodied Interaction
Digit's New Motor Cortex: Sim-to-Real RL for Whole-Body Control
EgoNav: Diffusion-Based Humanoid Navigation from Human Egocentric Video
CaP-X: A Code-as-Policy Framework for Robot Manipulation
Embodied Intelligence Breakthrough: Generalist AI’s GEN-1 Robots
CaP-X: LMs' First Physical Exam
AI Model Collapse: The Danger of Training on AI-Generated Data
High-Level Automated Reasoning with Qwen2.5-7B
Co-Training Large Behavior Models: Multimodal Data for Robot Manipulation
HyDRA: Hybrid Memory for Dynamic Video World Models
DexWM: Leveraging Human Videos for Dexterous Robot World Models
World Models in Robotics
SIMART: Decomposing Monolithic Meshes into Sim-Ready Articulated Assets
LeWorldModel: A Stable JEPA World Model from Pixels
World Models for Robots: The Next Big Leap?
Harnessing Long-Running AI in Embodied Systems
HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations
TurboQuant: Redefining AI Efficiency with Extreme Compression
DexWM: Learning Dexterous Object Manipulation from Human Videos
FlashAttention-3: Fast & Accurate Attention with Asynchrony & Low-Precision
When AI Trains on Its Own Output: The Model Collapse Problem
MolmoBot: A Vision-Language Model for Zero-Shot Robot Manipulation
LeWorldModel: Stable End-to-End JEPA from Pixels
EgoVerse: An Egocentric Data Ecosystem for Scaling Robot Learning
HSImul3R: Physics-Driven Reconstruction of Human–Scene Interactions
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
DreamZero: World Action Models Are Zero-Shot Policies
Kinema4D: A 4D Generative Simulator for Embodied AI
VEGA-3D: Teaching multimodal LLMs spatial reasoning through video generation