PODCAST · technology
HuggingFace 每日AI论文速递
by duan
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
-
624
2026.09.03 | 代码库蒸馏成AI技能;长视频世界模型可扩展
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:31] 🤖 Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills(从仓库到技能:将GitHub代码库蒸馏为AI4AI技能)[01:26] 🌍 SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models(SolarWM:面向长时程视频世界模型的开放数据与可扩展训练)[02:19] ⏱ EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction(EarlyEval:通过早期结果预测实现更廉价的智能体评估)[03:21] 🤝 It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning(两相匹配:基于强化学习的生成式检索器协同进化)[04:18] 🎯 Language Models Can Control Their Own Attention(语言模型能够控制自身的注意力)[05:10] 🖼 On the Design Fundamentals of Pixel Text Representation Learning(论像素文本表示学习的设计基础)[06:12] 🔍 Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering(超越视觉相似性:面向知识型视觉问答的实体对齐检索)[07:14] 🤖 HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?(HarnessDev:大语言模型能否创建并演化自己的智能体执行框架?)[08:05] 🗜 ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes(ZipTok3D:利用紧凑令牌前缀的高保真三维令牌化)[09:04] 🎯 Cliff: Learning Process Rewards from the First Mistake(Cliff:从第一个错误中学习过程奖励)[10:05] 🎯 Aspire: Can Models Self-Evolve from Vague Goals?(Aspire:模型能否从模糊目标中自我进化?)[10:46] 🎭 Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation(影响导向蒸馏:解决采样词元在线策略蒸馏中的多样性瓶颈)[11:42] 🎮 S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?(S3Gym:大语言模型能否将自我测试与自我评判转化为自我改进?)[12:36] 👀 A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss(一瞥足矣:基于SimLoss的单遍细粒度图像描述)[13:35] 🎙 VibeVoice-ASR-Streaming Technical Report(VibeVoice-ASR-Streaming技术报告)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
623
2026.09.02 | 学生模拟器实现因材施教;自动驾驶模型融合感知规划。
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:28] 🎓 StudentSim: Training LLM-based Student Simulators(StudentSim:训练基于大语言模型的学生模拟器)[01:34] 🚗 Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving(Qwen-Drive-1.0:迈向自动驾驶视觉语言基础模型的第一步)[02:46] 🔁 SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers(SMELT:计算匹配的MoE循环Transformer的缩放定律)[03:44] 🤖 UI-Venus-2 Technical Report(UI-Venus-2技术报告)[04:41] 🤖 ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training(ZimaBlue:通过可扩展视频预训练演化可泛化的世界动作模型)[05:40] 🌍 H3-World: Turning Language Understanding into World Control(H3-World:将语言理解转化为世界控制)[06:41] 🔍 Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering(Hi-Q:层次化证据引导的多跳问答查询细化)[07:36] 🚁 Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching(评估多模态大语言模型作为无人机控制的通用视觉-语言-行动智能体:指挥、接近、跟踪与搜索)[08:33] 🔗 Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System(揭示原生统一多模态模型中理解与生成的协同效应:从表征、任务到系统)[09:27] 🛡 Safin-1: Safety from Within through Memory-Native State Evolution(Safin-1:通过记忆原生状态演化实现由内而外的安全)[10:24] 🏢 From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix(从生产流量到后训练:构建覆盖企业请求组合的自托管大语言模型)[11:15] 🩺 DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory(DiagEvo:通过分层错误记忆进行诊断引导的自我进化)[12:07] 🧠 EM^2Mem: Event-Centric Multimodal Memory for Large Language Models(EM²Mem:面向大语言模型的事件中心多模态记忆)[13:09] 🤖 Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement(Harness-of-Harness:多日持续改进的自主软件开发)[14:07] 🛡 Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs(控制流与数据流分离:多智能体大语言模型中的稳定提示优化)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
622
2026.09.01 | 在线策略蒸馏实为自我改进;原生2K音视频联合生成
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:30] 🔄 Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement(在线策略蒸馏真的在蒸馏吗?从噪声教师到自我改进)[01:24] 🎬 DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution(DreamX-Creator:在2K分辨率下实现原生音视频生成的民主化)[02:13] 🧠 GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling(GenFirst:先生成后重建的稳定端到端潜在生成建模)[03:06] 🧩 Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling(Lucida:面向可组合真实到仿真场景建模的解析、生成与放置)[03:57] ⚖ Normalized Low-Rank Adaptation(归一化低秩自适应)[04:54] 🎯 PaperGym: Rubric-Centered Evolution for Research-Plan Generation(PaperGym:以评分标准为中心的研究计划生成进化)[05:53] ⚡ On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability(论Qwen3.8-Next架构的设计:评估、效率与训练稳定性)[06:51] 🎓 CogEvol: Towards Efficient and Reliable Learning Environment Generation(CogEvol:迈向高效可靠的学习环境生成)[07:50] 🧭 LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation(LightNav-0:激发视觉语言模型空间智能,实现通用具身导航)[08:43] 🧩 SHAPE of Chain-of-Thought in Math Reasoning(数学推理中思维链的SHAPE框架)[09:48] 🧠 Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence(超越人类监督的大规模推理模型扩展:通往超级智能之路)[10:51] 🧩 Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase(超级库智能体:超越单一代码库的多应用联合生成与维护)[11:48] 🤖 Evaluating the Hidden Costs of Personalization in Large Language Models(评估大型语言模型中个性化的隐藏成本)[12:42] 📋 Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents(先评估后改进:面向自动科研智能体的自动评分标准归纳)[13:32] 🎭 Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions(看得见的谎言:具身社会交互中VLM智能体的言语与非言语联合欺骗)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
621
2026.08.31 | LoopArena揭示循环控制难题;DART-SD拓扑感知优化工具调用。
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:29] 🔁 LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering(LoopArena:将模型作为循环工程的运行时控制器进行基准测试)[01:27] 💎 DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents(DART-SD:多轮工具调用代理的菱形拓扑感知检索与自蒸馏调优)[02:25] 🛠 Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities(智能体工件创建:系统、评估、原则与机遇)[03:25] 🤖 Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models(超越数据规模:面向视觉-语言-动作模型的以表征为中心的持续预训练)[04:42] 🌍 Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning(代码即世界:面向物理推理的可执行世界表示的智能体发现)[05:45] 🔄 J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data(J-Zero:零数据下挑战者—求解者—评判者的统一协同进化)[06:37] 🎥 Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction(重新审视面向长时程流式三维重建的局部上下文)[07:35] 🧠 ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL(ContextPilot:通过细粒度强化学习训练智能体进行主动上下文管理)[08:33] 🧠 LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation(LayerRecall:用于视频生成长时程一致性的状态条件记忆路由器)[09:38] 💰 Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090(Puro-2B:穷实验室在RTX 5090上以不到5090美元训练出的Qwen2-1.5B)[10:27] 🐘 Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge(盲人摸象:探究长尾分歧知识下大语言模型的认识论短视)[11:18] 🛡 StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing(StepGuard:通过可扩展监督与安全效用平衡学习步骤级护栏)[12:12] 🎨 Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents(画你所见:多模态智能体中灵巧视觉工具使用的基准测试)[13:08] ⚡ Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding(在视频中定位一切:重新思考高效生成式时空视频定位)[14:04] 🧠 PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control(PonderPounce:预训练多模态大语言模型作为机器人控制的回合上下文引擎)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
620
【周末特辑】8月第5周最火AI论文 | 视觉轨迹才是推理核心;智能体需以可验证进展为目标
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 5 篇论文如下:[00:42] TOP1(🔥253) | 🧠 VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning(VBVR-Pro:可扩展且可验证的原生视觉推理套件)[03:58] TOP2(🔥201) | 🤖 Apodex 1.1: Scaling Agentic Intelligence for Complex Work(Apodex 1.1:扩展智能体智能以应对复杂工作)[07:26] TOP3(🔥170) | 🎬 VGI-Bench: Probing Visual Intelligence in Video Generation Models(VGI-Bench:探究视频生成模型的视觉智能)[10:06] TOP4(🔥167) | 🧠 VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction(VoiceMem:面向实时交互的流式双脑记忆)[13:16] TOP5(🔥140) | 🧪 FrontierChallenge: Evaluating Scientific Workflow Completion(前沿挑战:评估科学工作流完成度)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
619
2026.08.28 | 视频模型概率校准不足;智能体数据需ACE平衡
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:30] 🎲 PAWBench: How Far Are We from Probabilistically Aligned World Modeling?(PAWBench:我们距离概率对齐的世界建模还有多远?)[01:09] 🤖 What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents(什么造就了好的智能体数据?面向LLM智能体的数据生成的ACE视角)[02:01] 🎯 TTPO: Test-Time Policy Optimization(TTPO:测试时策略优化)[02:49] 🤖 Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report(训练智能体与其运行框架协同进化:TaoLive数字人智能体技术报告)[03:41] 🏙 UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City(UrbanGround:从局部感知到真实尺度城市中的空间能动性)[04:34] 🎮 GameWAM: A World Action Model for Video Games(GameWAM:视频游戏的世界动作模型)[05:24] 🔄 PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents(PILOT在环:面向长时程智能体的实时自我改进)[06:03] 🤖 Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization(Zero-WAM:基于人类视频的上下文世界-动作建模实现开放式任务泛化)[06:55] ⚗ Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher(Self-OPD:无需教师模型的流匹配模型在线策略蒸馏)[07:51] 🧬 Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO(理解面向LLM推理的进化策略:比GRPO更广泛的推理覆盖)[08:44] 🎮 Magpie: Real-Time World Renderer for Interactive Games(Magpie:用于交互式游戏的实时世界渲染器)[09:41] 🧩 Procedura: Agentic 3D Modeling with Procedural Control(Procedura:程序化控制下的智能体3D建模)[10:39] 🎬 Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning(思考镜头:基于智能体推理的一致性多镜头视频编辑)[11:40] 🧠 WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution(WikiSkill:将智能体经验编译为持久知识以促进技能进化)[12:46] 🔗 CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval(CaSKG:面向可扩展智能体技能检索的反事实因果技能图)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
618
2026.08.27 | 双脑记忆让语音助手实时又准确;科学工作流多数模型难善终
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:31] 🧠 VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction(VoiceMem:面向实时交互的流式双脑记忆)[01:30] 🧪 FrontierChallenge: Evaluating Scientific Workflow Completion(前沿挑战:评估科学工作流完成度)[02:29] 🎬 VGI-Bench: Probing Visual Intelligence in Video Generation Models(VGI-Bench:探究视频生成模型的视觉智能)[03:32] 🚀 WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation(WarpSAC:重思探索与利用,迈向可扩展离策略强化学习的巅峰)[04:27] ⚙ JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution(JIT-Agent:通过即时框架演化实现框架智能的规模化扩展)[05:20] 🧠 VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning(VBVR-Pro:可扩展且可验证的原生视觉推理套件)[06:15] 🎛 D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation(D³-MOPD:面向高效多教师蒸馏的自适应动态领域调度)[07:22] 🧠 Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data(下一块推理强化学习真的比SFT更好吗?重新审视无CoT数据下的训练策略)[08:26] 🎯 Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning(Agent-G²:面向智能体强化学习的高斯引导)[09:13] 🎬 Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds(面向持久故事与交互世界的长时程音视频生成)[10:19] ⚖ Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation(Open-MOPD:多教师同策略蒸馏中能力失衡的诊断与修复)[11:26] 🤖 StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models(StreamPI:面向视觉-语言-动作模型的流式多模态时序建模)[12:29] 🎬 Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios(视频-IFBench:评估多模态大语言模型在视频理解场景中的指令跟随能力)[13:30] 🧠 Code World Model: Coding Agent as World Brain(代码世界模型:代码智能体作为世界大脑)[14:25] 🎯 V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning(V-Rubrics:基于评分标准的强化学习实现视觉忠实性)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
617
2026.08.26 | 将标注视为轨迹以加速视频强化学习;微信多模态嵌入模型刷新纪录
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:32] 🎯 Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs(将标注视为轨迹:高效可扩展的视频多模态大语言模型强化学习)[01:25] 🧲 WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report(WeMM-Embedding:微信多模态嵌入技术报告)[02:17] 🔧 AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces(AutoSaddler:基于智能体执行轨迹的自动框架优化与持久更新)[03:17] 🎯 On-Policy Self-Distillation in Diffusion Models(扩散模型中的同策略自蒸馏)[04:02] 🛡 CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild(CyberFactory:利用现实世界实例规模化网络安全能力)[05:00] 🧠 Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses(面向长时程智能体框架的递归经验-工作记忆演化)[05:53] 🎯 Best Practice Critic Optimization(最佳实践评论家优化)[06:45] 🎯 On-policy Distillation with Verifiable Reward(基于可验证奖励的在线策略蒸馏)[07:42] 🕶 From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms(从看见到行动:智能眼镜作为第一人称智能平台)[08:41] 🎮 Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training(Game2World引擎:解锁真实场景游戏视频以训练世界模型)[09:34] 📏 Length-Adaptive Decoding for Masked Diffusion Machine Translation(掩码扩散机器翻译的长度自适应解码)[10:27] 🎬 LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training(LAION-BVD:面向多模态预训练的千万小时开放视频数据集)[11:21] 🔁 Meta$^n$: Recursive Self-Improvement through Emergent Depth(Meta^n:通过涌现深度实现递归自我改进)[12:13] 🤖 CAFE: Self-Improving Search Agents Need Co-Evolving Feedback(CAFE:自我改进的搜索智能体需要协同演化的反馈)[13:14] 🔮 Latent Action as Intention Enables Efficient Future Imagination for World Action Models(以潜在动作作为意图,实现世界动作模型的高效未来想象)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
616
2026.08.25 | 智能体从会答到会交付;全模态世界开放可进入
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:30] 🤖 Apodex 1.1: Scaling Agentic Intelligence for Complex Work(Apodex 1.1:扩展智能体智能以应对复杂工作)[01:36] 🎮 EchoWM: Open and Enterable Omnimodal World Models(EchoWM:开放且可进入的全模态世界模型)[02:40] 🛒 TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming(TLive-Omni:面向电商直播的全模态理解模型)[03:22] 🎨 Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision(通过概念缩放与密集监督解锁图像编辑的潜力)[04:20] 📱 MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks(MobilePA-Bench:面向复杂真实世界任务的移动规划智能体基准评测)[05:16] 🤖 Prime Agent: A Self-Improving RLM Harness(Prime Agent:一种自我改进的递归语言模型工具框架)[06:22] 🧊 Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion(Block3D:通过块状扩散的高效文本到三维生成)[07:09] 🚗 RISE: Adaptive Imagination for World Action Models(RISE:面向世界行动模型的自适应想象)[08:19] 📈 Towards a Densing Law for User Representation Learning at Billion-Scale Capacity(面向十亿级容量用户表示学习的稠密化定律)[09:26] ⚖ ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction(ARC:开放式真实世界交互中的公平相对优势比较)[10:27] 🧠 ReWorld: An Interactive World Model with Long-Horizon Memory(ReWorld:一个具有长时记忆的交互式世界模型)[11:29] ⚖ Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization(超越稳定性-探索困境:面向大语言模型策略优化的环境正则化)[12:23] 🎮 GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?(GameXpert-Bench:编码智能体距离专家级游戏开发还有多远?)[13:17] 🧪 One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows(一次成功不等于可靠:Thinkingbox——面向有状态业务工作流的智能体沙盒与基准测试)[14:26] 🎯 Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection(Task-CoEvolve:通过自适应验证任务选择实现高效的工具链优化)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
615
2026.08.24 | 大模型超参迁移降本;图工程引领系统协作
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:32] 🚀 Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts(让我们一步一步扩展规模:面向大规模混合专家模型的高效超参数迁移)[01:24] 🕸 Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence(大语言模型智能体时代的图工程:从个体智能到系统智能)[02:21] 🤖 OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs(OmniAssistBench:全模态大语言模型的助手式交互基准)[03:17] ⚡ ParaTempo: Efficient Parallel Reasoning via Temporal Confidence(ParaTempo:基于时间置信度的高效并行推理)[04:00] ♾ InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter(InfinityEdit:基于轻量级编辑触发适配器的无限视频编辑)[04:53] 🪙 Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models(每一枚硬币都有两面:论大型语言模型同策略蒸馏中泛化的双重性)[05:37] 🧩 EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking(EviRank:面向多模态图像重排序的结构化相关性证据)[06:46] 🧠 Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs(超越正确性:混合思考多模态大语言模型的响应行为基准测试与对齐)[07:45] 🎨 UniSpace: Unified Visual Representation and Scalable Multimodal Modeling(UniSpace:统一视觉表示与可扩展多模态建模)[08:39] 🛒 Towards Faithful Simulation of Human Shopping Behavior(面向人类购物行为的忠实模拟)[09:35] ⚙ AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale(AgentMercury:你的智能体能够大规模合成可验证的商业场景环境)[10:41] ⚡ Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference(Daedalus-150M:为CPU推理设计的卷积-注意力混合模型)[11:35] 🛡 CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment(CLEAR:面向保持实用性的大语言模型安全对齐的连续潜在适配器路由)[12:42] 📱 Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs(Llama-Mobile:高效2.7比特视觉语言模型量化)[13:32] 🍳 FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth(FlavourBench:用可执行的烹饪真值对前沿语言模型进行排名)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
614
【周末特辑】8月第4周最火AI论文 | 操控系统升级让AI更准;视频检测难敌生成攻击
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 5 篇论文如下:[00:43] TOP1(🔥435) | ⚙ StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling(StateM:通过 Harness 扩展在 Terminal-Bench 2.1 上达到 95.3% 原始准确率,或一次 15 美元的前沿运行)[03:30] TOP2(🔥277) | 🛡 Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination(我们能防御针对现实世界危机事件的AI生成视频攻击吗?对检测器、生成器与社会传播的系统评估)[06:28] TOP3(🔥245) | 🌍 EnvHarness: Awakening Static Worlds for Agent Learning(EnvHarness:为智能体学习唤醒静态世界)[09:45] TOP4(🔥167) | 👁 Self-Supervised Visual On-Policy Distillation(自监督视觉同策略蒸馏)[12:21] TOP5(🔥162) | 🧩 Demystifying Agent Skills: Why They Work-Until They Don't(揭秘智能体技能:它们为何有效——直到失效)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
613
2026.08.21 | 动态环境适配弱点;任务合成保留源意图
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:31] 🌍 EnvHarness: Awakening Static Worlds for Agent Learning(EnvHarness:为智能体学习唤醒静态世界)[01:28] 🖥 FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis(FACET:在终端任务合成中保留源意图与可执行状态)[02:29] 🧪 SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?(SWE-bench Science:编码智能体能否解决科学领域的工程任务?)[03:15] 🕺 4DAnyone: Create Anyone in 4D from a Casual Monocular Video(4DAnyone:从随意单目视频创建任意人物的4D形象)[04:11] 👥 WithEveryone: Unified Planning and Identity Grounding for Group Image Generation(WithEveryone:面向群像生成的统一规划与身份锚定)[05:01] 🧠 MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use(MemTrapBench:大语言模型记忆使用中的认知陷阱基准测试)[05:52] 🔄 SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback(SkillEvo:来自多轮交互反馈的自我更新进化梯度)[06:50] 🎮 ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models(ForgeWM:面向少步动作条件视频世界模型的渐进式因果训练)[07:48] 🧩 Repo0: Design-Driven Zero-to-All Code Generation(Repo0:设计驱动的从零到全代码生成)[08:46] ⚡ FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving(FlashPrefill V2:面向长上下文大语言模型服务的块稀疏预填充注意力)[09:41] 🧠 Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization(注入、对齐、恢复:面向无检索文档知识内化的分阶段后训练)[10:42] 🤖 EXIMO: VLM Guided Exploration of VLA Policies(EXIMO:视觉语言模型引导的VLA策略探索)[11:36] 🧠 Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See(用低资源语言思考:SFT构建了什么,RL修复了什么,准确率无法看到什么)[12:18] 🎯 Towards Quantifying Benchmark Optimization in ASR Models(面向ASR模型中基准优化的量化研究)[13:20] 🛡 PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents(PolicyGuide:从守护单一动作到引导策略合规型LLM智能体的整个工作流)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
612
2026.08.20 | 闭环进化提升具身智能;验证门控保障工业代码
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:34] 🤖 Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence(Zetta ζ:面向自进化物理智能的高效闭环具身智能体框架)[01:29] ✅ SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation(SemaPLC:一种基于项目、以验证为门控的PLC代码生成智能体框架)[02:31] 🎯 SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation(SemComp-Bench:视频生成中语义任务完成的基准测试)[03:28] 🔬 OmniScientist: An Omni-Modal Omni-Discipline AI Scientist(全能科学家:一个全模态、全学科的人工智能科学家)[04:26] 🧠 Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL(Co-RL:多智能体强化学习中多样化群体催生无监督推理)[05:13] 🎮 SPADE: Self-Play in Adaptive Synthetic Executable Environments(SPADE:自适应合成可执行环境中的自博弈)[06:08] 🧪 Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis(训练面向单步逆合成的化学合理性感知大语言模型)[07:04] 🎯 Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning(潜在世界模型中的决策度量对齐:诊断与面向MPC规划的动作条件目标)[07:58] 🧬 Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification(训练留下痕迹:面向语言模型谱系验证的中心化残差签名)[08:50] 🔁 Looped Language Models Improve Compositional Tool Calling(循环语言模型提升组合式工具调用能力)[09:36] 🖐 SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation(SoftVTBench:面向可变形物体操作的变形感知视触觉数据集与基准)[10:33] ⚽ FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents(FM-Bench:面向竞争智能体的长时程管理基准)[11:25] ✍ Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion(借助属性引导的体裁扩展,将创意写作扩展到故事中心数据之外)[12:15] 🔥 The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning(越热门越难遗忘:大语言模型遗忘的自适应流行度方法)[13:01] 🔍 Temporal Multi-Signal Fusion for Token-Level Hallucination Detection(面向Token级幻觉检测的时序多信号融合)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
611
2026.08.19 | 进化策略微调省显存;技能应用有双刃剑
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:34] 🧬 Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements(Agentic ESOpt:以极低GPU需求微调长程LLM智能体)[01:46] 🧩 Demystifying Agent Skills: Why They Work-Until They Don't(揭秘智能体技能:它们为何有效——直到失效)[02:35] 🔬 ASI-Bench: At the Dawn of Artificial Superintelligence(ASI-Bench:人工超级智能的黎明)[03:26] 💻 FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution(FreeToken:高效的边缘原生MoE服务与带宽自适应执行)[04:20] 🧭 Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation(具身导航器:指向、思考、记忆与对齐实现高效导航)[05:22] 🎬 AVA-Encoder: Towards Agent-Native Video Representation Learning(AVA-编码器:迈向智能体原生的视频表示学习)[06:12] 🖼 EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing(EDITBRIDGE:迈向忠实且高效的超高分辨率图像编辑)[07:13] ⚡ Agent Lightning v1.0: Towards Harnessed Agentic RL(Agent Lightning v1.0:迈向框架化的智能体强化学习)[08:12] 🎥 V-RAE: Rethinking Video Latent Spaces for Generation(V-RAE:重新思考用于生成的视频潜空间)[09:17] 🎬 CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing(CoinVE-200K:面向组合式指令引导视频编辑的大规模高质量数据集)[10:17] 🛡 DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization(DiSCO:通过分布引导的对比提示优化防御文本到图像生成)[11:22] 🧠 Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents(驾驭记忆:对记忆智能体中记忆底层介质的整体评估)[12:32] 🎨 From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation(从语料到协同演进的能力:以能力为中心的通用图像生成数据设计)[13:31] 📊 StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows(StartupBench:对通用智能体在市场验证的端到端工作流上的基准测评)[14:34] ⚡ Energy-Guided Flow Matching(能量引导的流匹配)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
610
2026.08.18 | 智能体化评测让世界模型可诊断;多模态三维生成仍有瓶颈
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:32] 🕵 HarnessEval-W: Agentifying the Evaluation of Visual Worlds(HarnessEval-W:使视觉世界的评估智能体化)[01:32] 🌍 VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?(VibeWorlding:多模态智能体能否端到端构建3D开放世界?)[02:10] ⚡ MOSS-VL Technical Report(MOSS-VL 技术报告)[03:05] 🤖 ClawGym II: Exploring Black-Box RL on Agent Harness(ClawGym II:在智能体框架上探索黑盒强化学习)[03:56] 🎯 Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization(学习尚未掌握的,而非已经精通的:面向多奖励策略优化的饱和感知优势重加权)[04:51] 🤖 UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations(UI-Mate:利用上下文演示推进开放权重的基础图形用户界面智能体)[05:46] 🔬 Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search(大型发现模型:基于经验建模的开放式搜索)[06:37] 🎨 An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models(训练像素空间文本到图像扩散模型的实证研究)[07:31] 🤖 Agentic Transaction: Towards ACID-Compliant Agent Systems(智能体事务:迈向ACID合规的智能体系统)[08:29] 🔬 How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks(智能体如何在自动研究中失败:基于100个真实前沿研究任务的端到端诊断评估)[09:24] ⚡ GenRouter: Unified Workflow Routing for Agentic Image Generation(GenRouter:用于智能体图像生成的统一工作流路由)[10:29] 🧩 MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling(MegaParts:通过词元高效的自回归建模将部件感知的3D物体生成扩展到300个部件)[11:22] 🧠 Understanding Cognition-Induced Risks in Agentic AI Systems(理解智能体AI系统中认知引发的风险)[12:21] 🔗 Advancing Open and Reproducible Relational Learning: RelArena-$α$, TabPFN-Rel and RPI(推进开放可复现的关系学习:RelArena-α、TabPFN-Rel与RPI)[13:24] 🛡 Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs(Ventor-QTest:威胁模型驱动的供应商托管LLM API验证)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
609
2026.08.17 | 视频检测难防伪;自监督蒸馏促提升
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:30] 🛡 Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination(我们能防御针对现实世界危机事件的AI生成视频攻击吗?对检测器、生成器与社会传播的系统评估)[01:27] 👁 Self-Supervised Visual On-Policy Distillation(自监督视觉同策略蒸馏)[02:21] 🤖 Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development(超越最终得分:对长周期AI研究与开发智能体的系统评估)[03:15] 🧠 Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning(Intern-S2-Mobius:知识与推理解耦的基础模型)[04:09] 🎮 Marionette: Predicting World States, Rendering Geometry, Painting Appearance(Marionette:预测世界状态,渲染几何,绘制外观)[05:00] 🧠 SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning(SimpleOPD:面向长上下文推理的简单分词器无关在线策略蒸馏)[06:03] 🧠 MobileMem: Learning from a Year of Mobile Experiences(移动记忆:从一年的移动体验中学习)[06:54] 🤖 DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data(DFM Mimir v1:仅使用合规后训练数据、以1B参数实现前沿性能的开放HRM)[07:50] 🤸 HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark(HumanTracker:迈向全面且与人类感知一致的运动追踪基准)[08:50] 🎨 CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing(CPI-Bench:一个面向真实世界图像编辑的全面、实用且智能的基准)[09:48] 🧠 Latent On-Policy Self-Distillation(潜在同策略自蒸馏)[10:53] 🔍 Claim-Level Reliability Assessment for Efficient Test-Time Reasoning(面向高效测试时推理的声明级可靠性评估)[11:43] 🤖 PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment(PRM-as-a-Judge 1.5:机器人过程评估工具包)[12:42] 📉 Forecast Collapse in Time-Series Foundation Models(时间序列基础模型中的预测崩溃)[13:41] 🤔 Second Thought: Reasoning in Parallel as LLM Agents Act and Observe(第二思考:LLM智能体在行动与观察时并行推理)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
608
【周末特辑】8月第3周最火AI论文 | 潜空间推理降本;持续学习促进化
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 5 篇论文如下:[00:46] TOP1(🔥639) | 🧠 BDH-CQ: In-Context Learning with Recurrent Latent Reasoning(BDH-CQ:基于循环潜在推理的上下文学习)[03:27] TOP2(🔥332) | 🔄 Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA(Macaron-V1:迈向具备自我改进和LoRA混合的开放持续学习)[06:39] TOP3(🔥278) | 📄 Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill(从火花到论文:作为可组合技能的端到端研究论文生成)[10:15] TOP4(🔥255) | 🧬 OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution(OpenART:通过开放式环境演化扩展智能体红队测试)[13:01] TOP5(🔥208) | 🧠 On-Policy Self-Distillation without Any Supervision(无需任何监督的在线策略自蒸馏)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
607
2026.08.14 | 动作条件视频世界模型引入几何感知;长时记忆外部化实现无尽世界
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:32] 🤖 DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation(DreamX-Phi 1.0:面向机器人操作的动作条件视频世界模型)[01:24] 🌍 Alaya-EVOKE: From Linear-Scaling Supervision to Endless World(Alaya-EVOKE:从线性扩展监督到无尽世界)[02:30] 🔀 LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers(LLMRouter:开发、评估和部署LLM路由器的统一基础设施)[03:25] 🧬 DarwinX: Evolving Agent Harnesses Through Natural Selection(DarwinX:通过自然选择进化智能体框架)[04:26] 🔬 Intern-S2-Preview: Scientific Agentic Foundation Model(Intern-S2-Preview:科学智能体基础模型)[05:20] 🎮 PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives(PlayWorld:使用智能体玩家在长程目标上对世界模型进行基准测试)[06:20] 🤖 AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design(AutoDesign:面向长时程智能体设计的元框架优化)[07:23] 🧠 Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence(空间记忆智能体:基于经验的程序记忆实现空间智能)[08:12] ⚡ Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus(混合线性注意力大语言模型中的大规模激活:注意力前尖峰与尖峰间平台)[09:03] 🎭 UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos(UniSwap:面向说话视频的流式音频-视觉身份交换)[10:12] ⚡ LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time(LiveAnimate:实时稳定长格式流式人体动画生成)[11:08] 🤖 How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review(修辞何以能对AI审稿人进行奖励黑客?解析基于AI的同行评审中的修辞敏感性)[12:00] ✂ An AI4AI Framework for Visual Token Pruning(面向视觉Token剪枝的AI4AI框架)[12:56] 🤖 H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models(H2R-Bench:在世界模型中评估人类到机器人的操作视频生成)[14:08] 🔄 Full-bandwidth transformer(全带宽Transformer)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
606
2026.08.13 | 演化环境揭示智能体风险;组合技能实现论文生成
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:26] 🧬 OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution(OpenART:通过开放式环境演化扩展智能体红队测试)[01:35] 📄 Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill(从火花到论文:作为可组合技能的端到端研究论文生成)[02:43] 🧩 AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses(测试时AI4AI:通过推理支架实现强到弱能力迁移)[03:37] 🔬 Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence(Mechanist:将人工智能作为揭示智能机制的科学仪器)[04:33] 🎭 Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives(大语言模型智能体能否坚守剧本?交互式叙事中长程一致性的基准测试)[05:33] 🌍 StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization(StateFlow:为预可视化构建、演化与访问3D世界状态)[06:28] 📐 Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models(自几何:面向几何一致的3D视觉基础模型的无真值即插即用测试时自适应)[07:25] 🪞 From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection(从合成到去除:物理驱动的反射模拟与基于扩散模型的视频去反射)[08:22] 🛡 ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents(ToolHazard:扩展对抗性环境,用于基于大语言模型的智能体的安全评估与对齐)[09:15] 🔍 The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images(视觉工具使用的幻象:图像思维的因果审计)[10:16] 🤖 Self-Evolving Embodied Agents via Skill-Harness Evolution(基于技能与执行框架演化的自进化具身智能体)[11:19] 🧩 SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries(SkillZip:面向可扩展智能体技能库的合约保持图压缩)[12:14] 🛡 Agent Safety Should Be a Runtime Contract(智能体安全应当是一种运行时契约)[13:01] 🧬 Persistent Recursive Worlds Enable Autonomous Software Evolution(持久递归世界赋能自主软件演化)[13:46] 💡 MBA: Multimodal Benchmark and Agents for Real-World Business Ideation(MBA:面向真实世界商业构思的多模态基准与智能体)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
605
2026.08.12 | 共体智能体以人为中心助人成长;智能体与环境共演化迈向自我导向
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:35] 🤖 ComBodied Agents: a New Paradigm of Human-Centric Agentic AI(共体智能体:以人为中心的智能体人工智能新范式)[01:22] 🧬 Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design(智能体系统中的共同演化:迈向超越人类设计的自我导向演化)[02:22] 🌍 Beyond Pixels: From Video Priors to 4D Worlds(超越像素:从视频先验到4D世界)[03:12] 🧩 Articulated Object Reconstruction from Rest-State Observation(基于静止状态观测的铰接物体重建)[04:09] ⚔ AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss(AdvFD:通过对抗性弗雷歇距离损失提升视觉生成)[05:05] 🧬 Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution(孟德尔·哥德尔机:通过比较进化实现递归自我改进的编码智能体)[06:00] 🎭 Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence(Ex-Omni-2D:具备原生视觉临场感的表现性全模态对话模型)[06:48] 🌍 VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?(VibeLifeBench:你的生活智能体能否在动态世界中主动且持久地行动?)[07:46] 🚫 Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness(解码级禁忌:大语言模型鲁棒性的诊断性压力测试)[08:33] 📦 SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure(SkillZip:通过发现可复用结构实现自我进化智能体的免评估技能压缩)[09:33] 📱 SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information(SPIEval:评估大型语言模型作为移动助手处理分散个人信息的能力)[10:36] ✂ Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents(并不值得再投入一个词元:高效深度研究智能体的边际价值估计)[11:34] 🌐 Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation(开放大语言模型用于多语言机器翻译的无参考后训练)[12:28] 🔍 InSight-doc: Agentic Visual Perception for Long-Document Understanding(InSight-doc:面向长文档理解的智能体视觉感知)[13:27] 🔀 UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models(UniMoMo:基于专家合并的大型推荐模型MoE加速方法)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
604
2026.08.11 | 自进化混合专家赋能持续学习;代码重构基准揭示智能体局限
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:30] 🔄 Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA(Macaron-V1:迈向具备自我改进和LoRA混合的开放持续学习)[01:24] 🔧 SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring(SWE-Bench ProMax:面向大规模多语言代码重构的智能体基准评测)[02:21] 🐍 Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution(Ouroboros:通过核心评审进化实现自我发展的前沿编程智能体)[03:31] 🧠 BDH-CQ: In-Context Learning with Recurrent Latent Reasoning(BDH-CQ:基于循环潜在推理的上下文学习)[04:13] 🧠 Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory(智能体记忆蒸馏:利用分层教师记忆赋能小型大语言模型智能体)[05:11] 🧠 Motif 3: Technical Report(Motif 3:技术报告)[05:59] 🔬 Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains(Sci-VBench:评估科学领域中知识与推理密集型视频生成)[06:49] 🖼 What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems(下一步编辑什么:对话系统中的视觉对齐图像编辑后续建议)[07:53] 🎯 SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation(SPOT:面向同策略蒸馏的稀疏探测与结果校准)[08:58] ⚡ OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching(OasisKV:通过前瞻稀疏预取将解码期KV缓存扩展到HBM之外)[09:49] 🧠 RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States(RoMeRL:通过降阶效用状态平衡自进化智能体记忆中的反馈覆盖与记忆-奖励陷阱)[10:43] 🔍 Evidence-RL: Towards Evidence-intensive Visual Reasoning(证据强化学习:迈向证据密集型视觉推理)[11:42] 🧠 Scaling Inherently Interpretable Language Models(扩展内在可解释的语言模型)[12:40] 🧬 Evo-Bench: Can Language Models Improve Agent Harness?(Evo-Bench:语言模型能否改进智能体运行框架?)[13:40] 🔓 Stealing Reasoning Traces from Proprietary LLM APIs(从专有大语言模型API中窃取推理轨迹)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
603
2026.08.10 | 多模态智能体环境设计重质轻量;强化学习利于多任务共存
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:30] 🌍 Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning(超越单纯的环境规模扩展:为多模态智能体学习设计有效的环境分布)[01:36] ⚔ SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs(SFT冲突、RL共存:大语言模型多任务学习的理论与实证分析)[02:42] 🚗 SimWAM: A Simple World Action Model for End-to-End Autonomous Driving(SimWAM:面向端到端自动驾驶的简单世界动作模型)[03:40] 🎯 YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family(YOLO-PEFT:YOLO系列上的参数高效微调)[04:43] 🎥 StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding(StreamArena:迈向连续、交互与长时程的智能体流式视频理解)[05:36] 🎧 Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning(以演化评分标准作为奖励的音频推理强化学习)[06:19] 🙈 When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles(当激活预言机学会不读取:微调预言机中的概念特异性盲区)[07:09] ⚡ Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss(面向大语言模型的高效知识蒸馏:离线Top-K Logits与融合分块KL损失)[08:16] 🚁 Uncertainty-Aware World Model for Aerial Image-Goal Navigation(面向航拍图像目标导航的不确定性感知世界模型)[09:19] 🎯 Douyin Multimodal Embedding Model Technical Report(抖音多模态嵌入模型技术报告)[10:19] 📈 Skaling: Chinchilla's Exponents Meet Kaplan's Coupling(Skaling:Chinchilla指数与Kaplan耦合的融合)[11:11] 🧠 Addressable Memory for Video World Models(视频世界模型的可寻址记忆)[12:05] 🧩 Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression(相关但不完整:硬提示压缩中作为范式级失败模式的指称悬空)[13:01] 🔄 Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors(往返一致性:双向扩散模型可预测自身的展开误差)[14:01] 🧠 The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows(优化器即智能体:跨提示、程序与机器学习工作流的推理驱动搜索)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
602
【周末特辑】8月第2周最火AI论文 | 递归合成扩数据;状态管理提长程
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 5 篇论文如下:[00:41] TOP1(🔥221) | 🔁 Recursive Synthesis for Long-Horizon Terminal Tasks(面向长时程终端任务的递归式合成)[03:39] TOP2(🔥162) | 🧭 LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks(LongHorizon-Harness:推动面向真实世界任务的长时程智能体)[06:29] TOP3(🔥154) | 🎙 SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks(SwanTale:面向指令与零样本任务的统一多说话人语音与音频生成)[09:28] TOP4(🔥140) | 🚗 Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs(自动驾驶视觉语言模型中用于可验证推理的未来轨迹延迟暴露)[12:30] TOP5(🔥108) | ⚓ DAPD: Dual-Anchored Policy Distillation(双锚定策略蒸馏)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
601
2026.08.07 | 递归自蒸馏重塑智能体信用;开源裁判低成本评估操作
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:32] 🎯 AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning(AgentOPSD:面向智能体强化学习的递归自蒸馏)[01:35] 🤖 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models(OSReward:为跨平台计算机使用奖励模型制定标准化评估)[02:24] 🌍 WorldClaw: Agentic 3D Open-World Generation at Scale(WorldClaw:大规模智能体式3D开放世界生成)[03:12] 🗺 GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?(GST-Bench:视觉语言模型能否从视频中形成全局空间意识?)[04:17] 💭 EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning(EnvACE:通过世界预演将环境动态内化于智能体强化学习)[05:10] 🔍 Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval(从失败中学习:基于硬负样本的检索中心思维链用于统一多模态检索)[06:08] ⏳ ChronoVision: Temporal Reasoning via Latent State Reconstruction(ChronoVision:通过潜在状态重建实现时序推理)[07:08] 🌐 From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models(从经济主体到主体经济:经济世界模型的系统蓝图)[08:10] 🧮 On-Policy Delta Distillation for Multilingual Math Reasoning(面向多语言数学推理的同策略差值蒸馏)[08:59] ⚙ HarnessOpt-Bench: Evaluating LLMs at Harness Optimization(HarnessOpt-Bench:评估大语言模型在智能体运行框架优化上的表现)[09:43] 🇬 Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains(教Nemotron希腊语:面向专业领域的现代希腊语语料挖掘、检索适配与有据生成)[10:43] 🤖 DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation(DyPES-VLA:学习共享动力学先验与具身特定控制以实现跨具身操作)[11:38] 🤖 World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation(世界到手腕:面向精细机器人操作的任务条件化未来手腕建模)[12:44] 📊 DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces(DataSpace:面向异构工作空间的可验证分析数据代理基准测试)[13:44] 🪄 EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal(EffectLearner:面向真实世界视频目标移除的世界感知对象-效应推理)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
600
2026.08.06 | 回溯答案训练搜索智能体;统一工具实现图像生成
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:32] 🔍 ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment(ABSeeker:通过答案回溯的信用分配训练长程搜索智能体)[01:25] 🎨 ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation(ToolArtist:面向智能体图像生成的工具使用统一多模态模型)[02:16] 🪞 The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads(个性化幻象:大语言模型如何捏造用户画像,以及自我监控为何具有误导性)[03:17] ⚛ Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes(迈向多模态预训练的物理机制:知识流动、模态协同、早期统一与配方)[04:20] 🤖 OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents(OneDayAgent:面向自主智能体的长周期任务执行框架)[05:14] 🧬 GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks(GDPevo:在真实业务任务中评估智能体的自我进化)[06:05] 🎯 When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation(当教师误导:虚假信号感知的同策略蒸馏)[06:56] 🧩 Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning(迈向技能原生的大语言模型:用于长程推理评测与训练的技能熵)[08:02] 🧩 NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap(NOLLI:用于诊断英韩性能差距的难度校准谜题基准)[08:57] 🤖 Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data(Ego2Robot:从第一人称人类数据中可扩展合成机器人数据)[09:54] 🎬 AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities(AVE-Compass:面向音视频编辑能力的全面评估)[10:55] 🧠 When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents(当记忆说谎:VLM智能体中空间记忆陈旧性的实证研究)[11:53] 🎯 Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance(在失败处蒸馏:利用自适应教师指导恢复负强化学习组的学习信号)[12:49] 🧠 FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory(FocusMem:分解潜在GUI记忆中的内容、读取与信任)[13:52] 👋 HelloWorld: Enabling Socially Interactive Characters in Video World Models(HelloWorld:在视频世界模型中实现社交互动角色)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
599
2026.08.05 | 智能体长期运营后净资产不足人类三成;实时视频编辑达高清流畅
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:27] 🛒 MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations(MerchantBench:电商运营中LLM智能体长期一致性的基准评测)[01:19] 🎬 JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion(JoyAI-Video-Edit:利用自回归扩散的实时开放式视频编辑)[02:19] 🧊 Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing(Hunyuan3D-Buffalo 1.0:一种可扩展的三维生成、理解与编辑统一多模态模型)[03:15] 🌌 AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling(AURORA-LM:自编码统一表示用于连续潜变量扩散语言建模)[04:18] 🎥 Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent(视频深度研究:迈向下一代多模态深度研究智能体)[05:15] 🔄 Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation(知识-几何解耦:面向流式推荐的可刷新预训练迁移)[06:20] 🤖 PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning(PCSD:智能体强化学习中自蒸馏的持久一致性)[07:08] 🌍 Quo Vadis, World Modeling?(世界建模,何去何从?)[07:59] 🔁 PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents(PAST-Bench:对个人智能体中递归自我改进的基础进行基准测试)[08:56] 🌉 Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging(Any-OPD:通过表示空间桥接实现异构流匹配模型的在策略蒸馏)[09:57] 🗜 OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models(OmniPack:用于高效全模态大语言模型的统一令牌压缩)[11:02] 📈 LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models(LLaDA MoE v2:扩展混合专家扩散语言模型)[12:00] 🎯 CAPEval: A Decoupled Caption Evaluation across Understanding and Generation(CAPEval:面向理解与生成的解耦式图像描述评估)[12:55] 🧩 UniWorld-Design: From Pixel Generation to Layer-Native Design(UniWorld-Design:从像素生成到图层原生设计)[13:56] 🧠 SkillJack: Persistent Skill Backdoors in Self-Evolving Agents(SkillJack:自进化智能体中的持久性技能后门)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
598
2026.08.04 | 长时程智能体成功率提升至八成;多说话人语音音频统一生成
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:32] 🧭 LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks(LongHorizon-Harness:推动面向真实世界任务的长时程智能体)[01:30] 🎙 SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks(SwanTale:面向指令与零样本任务的统一多说话人语音与音频生成)[02:31] 🎯 VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation(VAD:在多模态在策略蒸馏中为目标重建归因视觉证据)[03:37] 🤖 Progressive Agent Skill Generation via Reinforcement Learning(基于强化学习的渐进式智能体技能生成)[04:39] ⚓ DAPD: Dual-Anchored Policy Distillation(双锚定策略蒸馏)[05:37] 🧲 UEmbed: Unified Sparse and Dense Multimodal Embeddings(UEmbed:统一稀疏与稠密多模态嵌入)[06:44] 🌍 WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity(WorldExam:从表象外观到内在反应性的世界模型基准评测)[07:54] 🔗 CADENA: Stepwise CAD Reverse Engineering(CADENA:逐步式CAD逆向工程)[08:55] 🛠 SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation(SKT:通过经验证的合成数据生成实现规模化技能使用训练)[09:58] 🤖 SWE-Touch: Benchmarking Coding Agents When Users Touch the Code(SWE-Touch:在用户改动代码时对编码代理的基准测试)[11:08] 🚗 Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs(自动驾驶视觉语言模型中用于可验证推理的未来轨迹延迟暴露)[12:09] 🧠 WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning(WCM:面向视觉-语言-动作强化学习的世界评论家模型)[13:05] 🔄 Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations(超越形态的运动:从抽象运动表征引导跨类别运动迁移)[14:04] 🧠 GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning(GradCuit:信用分配的梯度流实现稳健且可解释的测试时潜在推理)[14:59] 🛋 Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis(Roomer:面向三维室内布局合成的反思式对象级模型编辑与修复)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
597
【月末特辑】7月最火AI论文 | 虎鲸模型预测世界状态;Kimi K3开源逼近顶尖
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 10 篇论文如下:[00:42] TOP1(🔥473) | 🌍 Orca: The World is in Your Mind(虎鲸:世界在你心中)[03:50] TOP2(🔥438) | 🧠 Kimi K3: Open Frontier Intelligence(Kimi K3:开放前沿智能)[06:31] TOP3(🔥309) | 🧩 Program-as-Weights: A Programming Paradigm for Fuzzy Functions(程序即权重:面向模糊函数的编程范式)[09:09] TOP4(🔥307) | 🌍 ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU(ABot-World-0:在单个桌面GPU上实现无限交互式世界展开)[12:34] TOP5(🔥293) | 🧪 AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis(AskChem:以论断为中心的化学文献综合基础设施)[15:21] TOP6(🔥289) | 🤖 Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents(Qwen-UI-Agent技术报告:迈向下一代以真实世界为中心的基础GUI智能体)[18:00] TOP7(🔥259) | 🧠 Metis: Memory Foundation Model(Metis:记忆基础模型)[22:07] TOP8(🔥232) | 🧭 Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable(驾驭手册:使不断演化的智能体驾驭系统可读、可导航且可编辑)[25:09] TOP9(🔥205) | 🧠 LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget(长稻草:在固定GPU预算下实现超过200万Token的长上下文强化学习)[28:24] TOP10(🔥198) | 🤖 RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model(RynnBrain 1.1:迈向更强大和更通用的具身基础模型)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
596
【周末特辑】8月第1周最火AI论文 | Kimi K3开源前沿智能;AskChem论断级化学检索
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 5 篇论文如下:[00:44] TOP1(🔥431) | 🧠 Kimi K3: Open Frontier Intelligence(Kimi K3:开放前沿智能)[03:17] TOP2(🔥292) | 🧪 AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis(AskChem:以论断为中心的化学文献综合基础设施)[05:53] TOP3(🔥281) | 🤖 Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents(Qwen-UI-Agent技术报告:迈向下一代以真实世界为中心的基础GUI智能体)[08:46] TOP4(🔥256) | 🧠 Metis: Memory Foundation Model(Metis:记忆基础模型)[11:40] TOP5(🔥192) | 🤖 Progress Reward Modeling for Robotic Learning: A Comprehensive Survey(机器人学习的进度奖励建模:一项全面综述)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
595
2026.07.30 | TurboVLA实现消费级显卡实时操控;CoRT精细化信用分配提升指令遵循。
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:33] ⚡ TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM(TurboVLA:在RTX 4090上以32Hz频率运行且显存占用低于1GB的实时视觉-语言-动作模型)[01:36] 🎯 CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization(CoRT:用于令牌级准则引导策略优化的反事实重放)[02:31] 🤖 HumanCLAW: Can Vision-Language Models Act Through a Body?(HumanCLAW:视觉-语言模型能否通过身体行动?)[03:41] 🧬 DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space(DecoEvo:文本空间中求解器与评估生成器技能的解耦协同进化)[04:39] 🧠 CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition(CLBench-V:从基础定位到知识获取的多模态上下文学习评估)[05:32] 🧩 CAST: Game Solvers as Turn-Level Teachers for LLM Agents(CAST:游戏求解器作为LLM智能体的回合级教师)[06:28] 🧠 SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution(技能崛起:面向跨任务技能演化的智能体强化学习)[07:18] 🎮 StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation(StatePlay:状态感知的游戏世界模型用于机制一致的内容生成)[08:05] 📊 OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding(OmegaUse-OfficeVal:基于经济基准评估LLM智能体在长期办公套件任务中的表现)[09:08] 🤖 Can AI agents conduct open-ended AI research? Early evidence from two case studies(AI智能体能否进行开放式的AI研究?来自两个案例研究的早期证据)[10:04] 🛡 GPT-Red: Automated Red Teaming via Self-Play at Scale(GPT-Red:通过大规模自我对弈实现自动化红队测试)[11:04] 🎬 Explicit Layer Modeling for Video Object Insertion and Layer Decomposition(显式层建模用于视频对象插入与层分解)[12:01] 🕵 StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents(StealthBench:衡量自主攻击安全代理的操作隐蔽性)[12:54] 🧠 Memory for Large Language Models(大型语言模型的记忆机制)[13:48] 📜 Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems(为叙述者评级:面向多智能体知识系统中声明级溯源的一种伊斯纳德-里贾尔框架)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
594
2026.07.29 | 高保真数据训练策略,逼近真机效果;相关性动态引导搜索,精准高效检索
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:31] 🤖 HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone(HiFi-UMI:仅从高保真UMI数据学习可部署的操作策略)[01:31] 🔍 A New Role for Relevance: Guiding Corpus Interaction in Agentic Search(相关性的新角色:在智能体搜索中引导语料库交互)[02:20] 🎨 ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition(ReDesign:通过智能体分解从图像中恢复可编辑的设计结构)[03:07] 🧠 Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory(保持铭记:基准测试代理记忆中的内隐关联盲点)[03:56] 🏃 Pass the Baton: Trajectory-Relayed On-Policy Distillation(传递接力棒:轨迹中继的在线策略蒸馏)[04:47] ⚡ Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model(Mage-VL:一种高效的编解码器原生流式多模态基础模型)[05:49] 🧩 CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents(CodeNib:一种为编码智能体提供仓库上下文服务的多视图数据系统)[06:48] 🌍 Wonder: Video World Model Done Better(Wonder:更优的视频世界模型)[07:51] 👁 PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models(感知基准:评估多模态大语言模型中的原子视觉感知能力)[08:44] 🔍 Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking(新颖主张还是似曾相识?重新思考多模态自动事实核查的“无污染”动态评估)[09:40] 🛡 Shieldstral(盾星)[10:41] 🔀 MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities(MODUS:仅解码器的任意模态到任意模态多样化建模)[11:41] ⚡ Parallel Decoding Distillation for Fast Image and Video Generation(并行解码蒸馏:面向快速图像与视频生成的方法)[12:23] 🎬 Visual prompt engineering for video models(视频模型的视觉提示工程)[13:16] 🎯 OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs(OmniDelta:面向全模态大语言模型令牌压缩的技能驱动预算分配方法)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
593
2026.07.28 | Kimi K3开源模型性能领先;JarvisHub画布框架革新创意协作
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:33] 🧠 Kimi K3: Open Frontier Intelligence(Kimi K3:开放前沿智能)[01:22] 🎨 JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents(JarvisHub:一个面向画布原生多模态创意代理的开放框架)[02:20] 🤖 From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search(从专有到开源:通过多智能体协议蒸馏弥合智能搜索中的分布差距)[03:18] 🤖 Progress Reward Modeling for Robotic Learning: A Comprehensive Survey(机器人学习的进度奖励建模:一项全面综述)[04:17] 🤖 StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents(StateAct:面向长周期计算机使用代理,程序状态优先于像素)[05:23] 🧠 Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation(重新思考在线策略扩散蒸馏中的无分类器引导)[06:27] 🗼 Data Pyramid for Embodied Manipulation(具身操作的数据金字塔)[07:33] ⚡ Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification(Sol-Attn:通过即时注意力稀疏化加速视频生成推理)[08:37] 🎬 OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation(OmniVAE:一种具有跨模态对齐的音频-视频VAE,用于联合生成)[09:30] 🧠 The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation(多轮长程规划中的物理学:通过单教师与多教师在线策略智能体蒸馏从预训练到后训练)[10:26] 👗 Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On(氧气试穿:面向时尚的原生基础模型,实现任意物品虚拟试穿)[11:17] 🦎 Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling(Chamaileon:基于情境化建模与混合采样的跨情境结合剂设计)[12:11] 🔮 dRAE: Representation Autoencoder with Hyper-Spherical Codes(dRAE:超球面码的表示自编码器)[13:09] 🧩 DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes(解耦混合:用于可扩展VLM数据配方的解耦比率搜索与凸分配)[14:08] 🏥 ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding(ClinFusion:一种面向整体医学理解的以视觉为中心的多模态大语言模型系统)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
592
2026.07.27 | 技能自我对弈推动模型能力前沿;智能体上下文管理优化成本与推理。
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 14 篇论文如下:[00:32] 🔄 Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills(技能自我对弈:通过协同进化技能推动大语言模型能力前沿)[01:22] 🧠 Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems(智能体上下文管理:将代理记忆与成本视为生命周期与架构问题)[02:18] 🤖 Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning(Molt:一个可扩展的、原生PyTorch的智能体强化学习训练框架)[03:07] 🧪 DataPrep-Bench: Benchmarking LLMs as Training Data Preparators(数据准备基准:将大语言模型作为训练数据准备工具的基准测试)[04:04] 📊 Scaling Native Multimodal Pre-Training From Scratch(从头开始扩展原生多模态预训练)[04:54] 🎯 Three-Body Scattering for Generative Modeling(用于生成建模的三体散射)[05:48] 🌐 LAMAR: An Open Language-Aware Multilingual Alignment Reranker(LAMAR:一种开放的语言感知多语言对齐重排序器)[06:46] 🧠 Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making(多头潜在控制:大语言模型智能体决策的统一接口)[07:43] 🎛 Spectral Prior for Reducing Exposure Bias in Diffusion Models(用于减少扩散模型曝光偏差的频谱先验)[08:34] 🧠 IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation(IDEAgent:面向研究创意生成的主体性质量-多样性搜索)[09:27] 🎯 SceneActBench: Can Agents Act on the 3D Scenes They See?(场景动作基准:智能体能否对所见的三维场景采取行动?)[10:37] 🔄 Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering(闭环:无需训练的循环一致性自回归生成渲染)[11:31] 🔊 Multimodal Speaker Verification as a Threat to Speaker Anonymization(多模态说话人验证对说话人匿名化的威胁)[12:26] 🧠 VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression(VisCo:利用大语言模型作为视觉标记压缩的内在编码器)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
591
【周末特辑】7月第4周最火AI论文 | 单卡GPU实现无限交互世界;三维定位让机器人更智能
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 5 篇论文如下:[00:47] TOP1(🔥297) | 🌍 ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU(ABot-World-0:在单个桌面GPU上实现无限交互式世界展开)[02:39] TOP2(🔥197) | 🤖 RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model(RynnBrain 1.1:迈向更强大和更通用的具身基础模型)[04:48] TOP3(🔥165) | ⏱ TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs(TimeLens2:基于多模态大语言模型的通用视频时间定位)[06:35] TOP4(🔥144) | 🔍 RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM(RAGU:一种具有紧凑领域自适应大语言模型的多步图检索增强生成引擎)[09:10] TOP5(🔥141) | 🔍 AREX: Towards a Recursively Self-Improving Agent for Deep Research(AREX:迈向递归自我改进的深度研究智能体)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
590
2026.07.24 | AREX验证驱动递归改进;ReferTrack先指认后跟踪
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:32] 🔍 AREX: Towards a Recursively Self-Improving Agent for Deep Research(AREX:迈向递归自我改进的深度研究智能体)[01:34] 🤖 ReferTrack: Referring Then Tracking for Embodied Visual Tracking(ReferTrack:面向具身视觉追踪的“先指认后跟踪”范式)[02:20] 📚 K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs(K12-KGraph:一个面向课程对齐的知识图谱,用于基准测试和训练教育大语言模型)[03:13] 🖼 Visual Contrastive Self-Distillation(视觉对比自蒸馏)[04:02] 🗺 Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text(展示而非叙述:在生成像素而非LLM文本中评估空间认知)[04:51] 🎨 Color Pass-Through via Camera-Display Coupling(通过相机-显示耦合的色彩直通)[05:45] 🛠 Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction(腾讯工作伙伴基准:一个具有抗污染任务构建的多领域编码智能体基准)[06:43] 🧭 LLMs Get Lost in Evolving User Intent(大语言模型在用户意图演变中迷失方向)[07:33] 🎥 Self-Supervised Learning of Structured Dynamics from Videos(从视频中自监督学习结构化动力学)[08:28] 🧠 Sample-Efficient Learning from Agent Experience(从智能体经验中进行样本高效学习)[09:28] 🌀 Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation(用于高效高保真表示的递归正弦隐式神经表示)[10:27] 🌍 Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers(流式多智能体自回归扩散模型与世界状态寄存器)[11:28] 🤖 Robostral Navigate(罗博斯特拉导航)[12:18] 🎭 Predictive Divergence Masks for LLM RL(预测性散度掩码用于大语言模型强化学习)[13:04] 🎬 SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation(SANA-Video 2.0:混合线性注意力与注意力残差实现高效视频生成)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
589
2026.07.22 | 实时游戏渲染提速56倍;AI数据管道成本降七成。
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:34] 🎮 Generative World Renderer at the Speed of Play(以游戏速度运行的生成式世界渲染器)[01:27] 🔀 DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines(数据流驾驭平台:一种用于构建可编辑LLM数据管道的接地代码代理平台)[02:16] 🔍 Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers(文本模板标记是扩散Transformer中的隐式语义寄存器)[03:04] ⚡ Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing(Mage-Flow:一种用于图像生成与编辑的高效原生分辨率基础模型)[03:52] 🌍 AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report(AlayaWorld:交互式长时域世界建模——完整技术报告)[04:49] 🤖 Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning(陈旧但稳定:面向异步强化学习稳定化的陈旧性自适应信任区域方法)[05:52] 🌍 ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU(ABot-World-0:在单个桌面GPU上实现无限交互式世界展开)[06:49] 📊 SciForma: Structure-Faithful Generation of Scientific Diagrams(SciForma:科学图表的结构忠实生成)[07:45] 🔍 AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents(AgentDebugX:用于LLM Agent故障可观测性、归因与恢复的开源工具包)[08:41] ⚡ HPD-Parsing: Hierarchical Parallel Document Parsing(HPD-Parsing:层级并行文档解析)[09:38] 📊 Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness(用于评估开放生成的两级元评价标准:GAMUT,一个面向事实完整性的基准测试)[10:33] 🧠 ISO: An RLVR-Native Optimization Stack(ISO:一种原生于RLVR的优化栈)[11:41] 🎙 Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing(转录策略作为潜在变量:通过词级时序激活可控的逐字语音识别)[12:43] 🎓 EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration(EduPanel:一种用于教学视频的三智能体LLM评审器——可靠性、互补性与人类信任校准)[13:29] 🎬 Masked Visual Actions for Unified World Modeling(掩蔽视觉动作:用于统一世界建模)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
588
2026.07.21 | 时间定位与高效剪枝;两篇论文提升模型精准度与效率
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:31] ⏱ TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs(TimeLens2:基于多模态大语言模型的通用视频时间定位)[01:14] ✂ SWE-Pruner Pro: The Coder LLM Already Knows What to Prune(SWE-Pruner Pro:编码器大语言模型已知道该剪枝什么)[02:05] 🌍 EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World(演化世界:面向交互式文学世界中角色与世界观协同演化的开放模式框架)[03:08] 🔍 DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment(DeepSearch-World:可验证环境中深度搜索代理的自我蒸馏)[04:06] 🎬 HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement(HOMIE:通过多模态智能增强实现以人-物为中心的视频个性化)[04:59] 🤖 RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model(RynnBrain 1.1:迈向更强大和更通用的具身基础模型)[05:51] 🍎 Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence(苹果-π:面向法律约束的物理智能,以视频为基准测试思维过程)[06:49] 🎯 Group Entropy-Controlled Policy Optimization(群体熵控制策略优化)[07:52] 🧠 ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams(反射世界-多模态:面向开放视频流的实体导向多模态记忆系统)[09:03] 🌍 GigaAM Multilingual: Foundation Model for Underrepresented Languages(GigaAM多语言模型:面向低资源语言的基座模型)[09:51] ⏱ GigaChat Audio: Time-aware Large Audio Language Model(GigaChat音频:时间感知的大规模音频语言模型)[10:50] 🤖 Environment-free Synthetic Data Generation for API-Calling Agents(面向API调用智能体的无环境合成数据生成)[11:42] 🎬 FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry(FlowMimic: 基于像素对扭曲流场的无掩码视觉编辑与生成——用于在线视频编辑数据生成与模态模仿)[12:30] 🧊 DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation(DiffGI:面向高保真薄壳三维生成的可微几何图像)[13:20] 🎓 LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks(LLM作为教练:非可验证任务的体验式学习)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
587
2026.07.20 | RAGU引擎用轻量模型低成本构建知识图谱
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:32] 🔍 RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM(RAGU:一种具有紧凑领域自适应大语言模型的多步图检索增强生成引擎)[01:35] 🎯 RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources(RESOURCE2SKILL:从人类创建的多模态资源中提炼可执行的智能体技能)[02:34] 🤖 Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories(小米机器人-1:基于超过10万小时真实世界轨迹数据扩展视觉-语言-动作模型)[03:21] 🔗 xHC: Expanded Hyper-Connections(xHC:扩展超连接)[04:09] 🔄 Loop the Loopies!(循环循环!)[05:01] 🏥 Cura 1T: Specialized Model for Agentic Healthcare(Cura 1T:面向智能体医疗的专用模型)[05:47] 🧠 RecGPT-V3 Technical Report(RecGPT-V3技术报告)[06:34] 🧪 On-Policy Delta Distillation(在策略增量蒸馏)[07:20] 🤖 From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality(从以人为中心到智能体代码审查:不同代际生成式人工智能技术对审查质量的影响)[08:14] 🎵 Qwen-Music Technical Report(Qwen-Music技术报告)[09:08] ♟ Understanding Reasoning from Pretraining to Post-Training(理解从预训练到后训练中的推理能力)[10:00] 🤖 When Does Muon Help Agentic Reinforcement Learning?(缪子优化器何时有助于智能体强化学习?)[10:56] 🔄 Recursive Harness Self-Improvement(递归式框架自我改进)[11:50] 🎥 VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders(VideoRAE:通过表示自编码器驯服视频基础模型以进行生成建模)[12:49] 🦩 Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos(音频-视觉火烈鸟:面向长视频与复杂视频的开放音频-视觉智能)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
586
【周末特辑】7月第3周最火AI论文 | 行为定位新突破;长上下文训练显存优化
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 5 篇论文如下:[00:52] TOP1(🔥200) | 🧭 Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable(驾驭手册:使不断演化的智能体驾驭系统可读、可导航且可编辑)[03:22] TOP2(🔥179) | 🧠 LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget(长稻草:在固定GPU预算下实现超过200万Token的长上下文强化学习)[06:16] TOP3(🔥150) | 🎥 VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding(VideoChat3:面向高效与通用视频理解的全开放视频多模态大语言模型)[08:55] TOP4(🔥130) | 🧠 Weak-to-Strong Generalization via Direct On-Policy Distillation(弱到强泛化:通过直接在线策略蒸馏)[11:09] TOP5(🔥127) | 🖼 Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation(Boogu-Image-0.1:推动开源统一多模态理解与生成)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
585
2026.07.16 | 行为定位与高效多模态;开源模型追赶闭源系统
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:33] 🧭 Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable(驾驭手册:使不断演化的智能体驾驭系统可读、可导航且可编辑)[01:19] 🖼 Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation(Boogu-Image-0.1:推动开源统一多模态理解与生成)[02:07] 🧠 Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning(环零:将零强化学习扩展至万亿参数以实现涌现推理)[03:08] 📄 OvisOCR2 Technical Report(OvisOCR2 技术报告)[04:12] 🤖 KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill(知深行准-GUIClaw:具备自进化记忆与技能的深度认知、精准执行个人GUI助手)[05:10] 🛡 PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails(政策转变守卫:基准测试与改进策略自适应图像防护机制)[06:02] 🤖 GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch(GigaWorld-Policy-0.5:一种由自动研究驱动的更快更强的世界动作模型)[06:57] 🖼 MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors(MetaView:具有尺度感知隐式几何先验的单目新视角合成)[07:46] ✂ ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation(ShortOPD:通过短到长在线策略蒸馏恢复剪枝后的大语言模型)[08:40] 🎭 Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation(Hallo4D:多模态幻觉缓解实现一致的时空生成)[09:28] 🔬 Registers Matter for Pixel-Space Diffusion Transformers(寄存器对像素空间扩散Transformer至关重要)[10:22] 🤖 Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos(Vinci2:在连续自我中心视频中提供主动辅助)[11:21] 🔍 Tracing Agentic Failure from the Flow of Success(从成功流程中追溯智能体失败)[12:20] 🔍 From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization(从噪声轨迹到根因:面向智能体优化的结构化轨迹分析与因果提取)[13:23] 🧭 AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities(AgentCompass:一种统一的智能体能力评估基础设施)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
584
2026.07.15 | SpectraReward:零样本多模态奖励模型;盲点基准:揭示AI的认知盲区
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 10 篇论文如下:[00:31] 🔄 Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation(重新读它:预训练多模态大语言模型是文本到图像生成的零样本奖励模型)[01:35] 🔍 Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models(盲点基准:评估多模态模型中的盲点)[02:30] 📄 SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding(SynthDocBench:面向长上下文视觉文档理解的受控基准)[03:31] 🔍 Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution(先知晓再修复:面向软件问题修复的基于问答的仓库知识获取)[04:20] 🎵 MuScriptor: An Open Model for Multi-Instrument Music Transcription(MuScriptor:面向多乐器音乐转录的开放模型)[05:12] 🔍 Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation(超越可教授的知识边界:在智能体视觉生成中演化知识边界)[06:00] 🎨 Let RGB Be the Language of Vision(让RGB成为视觉的语言)[06:50] 📄 MonkeyOCRv2: A Visual-Text Foundation Model for Document AI(MonkeyOCRv2:面向文档AI的视觉-文本基础模型)[07:50] 🤖 Towards Autonomous and Auditable Medical Imaging Model Development(迈向自主且可审计的医学影像模型开发)[08:40] 🧠 Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms(深度强化学习评估与设计范式的原则性分析)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
583
2026.07.14 | 弱到强泛化提升大模型;双系统架构导航新突破
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 14 篇论文如下:[00:32] 🧠 Weak-to-Strong Generalization via Direct On-Policy Distillation(弱到强泛化:通过直接在线策略蒸馏)[01:22] 🤖 ABot-N1: Toward a General Visual Language Navigation Foundation Model(ABot-N1:迈向通用视觉语言导航基础模型)[02:20] 🤖 ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory(ABot-AgentOS:一种具有终身多模态记忆的通用机器人智能体操作系统)[03:21] 🕺 4D Human-Scene Reconstruction from Low-Overlap Captures(低重叠度捕获下的四维人体场景重建)[04:13] 🧠 LightMem-Ego: Your AI Memory for Everyday Life(轻量记忆自我:你的日常生活AI记忆)[05:15] 🧮 AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification(高级数学推理基准:面向高级数学证明生成与验证的基准套件)[06:11] 🧠 Metacognition in LLMs: Foundations, Progress, and Opportunities(大语言模型中的元认知:基础、进展与机遇)[07:03] 🤖 EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos(EgoSteer:面向从自我中心视频实现可操控灵巧操作的全栈系统)[07:55] 🧩 Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals(代理探索与可复用引导:一种通过代理引导更新信号实现的模块化大语言模型后训练范式)[08:55] 🧠 NeuroCogMap Reveals Cognitive Organization of Large Language Models(神经认知图谱揭示大型语言模型的认知组织)[09:52] 🎬 Motion4Motion: Motion Transfer Across Subjects at Inference(Motion4Motion:推理时跨主体的运动迁移)[10:58] 👗 CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation(CtrlVTON:通过视觉实例提示分割实现可控虚拟试穿)[11:54] 🎭 Latent-Identity Tuning in Text-to-Image Personalization Models(文本到图像个性化模型中的潜在身份调优)[12:52] 🧊 LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow(LATO.2:基于顶点与拓扑流的分解式3D网格生成)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
582
2026.07.13 | LHTB基准揭示AI长时任务瓶颈;视觉预训练超越文本预训练
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 13 篇论文如下:[00:33] 🚀 Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading(长时终端基准:在具有密集奖励评分的长期终端任务中测试智能体的极限)[01:25] 👁 Scalable Visual Pretraining for Language Intelligence(面向语言智能的可扩展视觉预训练)[02:22] 🎥 Video Generation Models are General-Purpose Vision Learners(视频生成模型是通用视觉学习器)[03:16] 🎯 Trust Region Policy Distillation(信任区域策略蒸馏)[04:14] 🧠 KronQ: LLM Quantization via Kronecker-Factored Hessian(KronQ:通过Kronecker分解黑塞矩阵实现的大语言模型量化)[05:00] 🌍 PanoWorld: Real-World Panoramic Generation(PanoWorld:真实世界全景生成)[05:56] 🎨 From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models(从RGB生成到密集场读取:利用文本到图像模型进行像素空间密集预测)[06:49] 🎯 Self-Guided Test-Time Training for Long-Context LLMs(自引导测试时训练用于长上下文大语言模型)[07:36] 🚗 Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation(Flow-ERD:面向多样化交通仿真的智能体类型感知流匹配与熵正则化蒸馏)[08:28] 🔊 Phone Segmentation and Recognition through Phonological Activation Mapping(通过音位激活映射进行音素分割与识别)[09:28] 🧠 Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning(探究大语言模型微调中记忆知识为何无法泛化的机制性理解)[10:18] 🤖 A Sovereign, Open-Source Foundation Model for German and English(一个面向德语和英语的主权开源基础模型)[11:13] 🏛 VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery(VaseMuseum:面向古希腊陶器的数字智能博物馆)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
581
【周末特辑】7月第2周最火AI论文 | 训练-推理不匹配问题被揭示;实时交互视频生成实现突破
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 5 篇论文如下:[00:51] TOP1(🔥161) | 🎯 The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning(优化训练策略的幻影:单调推理策略作为大语言模型强化学习的真正目标)[02:56] TOP2(🔥119) | 🎬 Vidu S1: A Real-Time Interactive Video Generation Model(Vidu S1:一种实时交互式视频生成模型)[04:56] TOP3(🔥89) | 🤖 RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation(RynnWorld-4D:面向机器人操作的4D具身世界模型)[07:15] TOP4(🔥84) | 🎮 AlayaWorld: Long-Horizon and Playable Video World Generation(AlayaWorld:长视界且可游玩的视频世界生成)[09:58] TOP5(🔥83) | 🔬 Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning(基于深度原生结构推理的精确、跨学科且透明的构效关系理解)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
580
2026.07.10 | Vidu S1实现实时视频对话交互;视频绿洲揭示理解评估缺陷
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:31] 🎬 Vidu S1: A Real-Time Interactive Video Generation Model(Vidu S1:一种实时交互式视频生成模型)[01:29] 🔍 Video-Oasis: Rethinking Evaluation of Video Understanding(视频绿洲:重新思考视频理解的评估)[02:27] 🔑 Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition(为什么我打不开抽屉?缓解零样本组合动作识别中的对象驱动捷径)[03:21] 🤖 UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks(UniClawBench:面向真实世界任务中主动智能体的通用基准)[04:18] 🎥 LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models(LongE2V:基于视频扩散模型的长时域事件驱动视频重建、预测与帧插值)[05:24] 🧬 Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation(思想具有基因组:科学谱系推理与基于谱系的创意生成基准测试)[06:13] 🌐 Enhancing In-context Panoramic Generation via Geometric-aware Pretraining(通过几何感知预训练增强上下文全景生成)[06:54] 🎬 CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation(CineMobile:面向电影级摄像机运动生成的设备端图像到视频扩散)[07:52] 🎬 OpenCoF: Learning to Reason Through Video Generation(OpenCoF:通过视频生成学习推理)[08:45] 🚀 Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE(Jet-Long:基于动态双焦点旋转位置编码的高效长上下文扩展)[09:39] ⚡ Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing(线性注意力架构:机制、权衡与跨层路由)[10:38] 💊 DrugGen 2: A disease-aware language model for enhancing drug discovery(DrugGen 2:一种疾病感知型语言模型,用于增强药物发现)[11:23] ⚡ Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models(Flash-BoN:扩散模型中推理时缩放的即时草稿生成)[12:08] 🎵 A Quantized Native Runtime for On-Device Semantic Audio Generation(面向设备端语义音频生成的量化原生运行时)[13:01] 🚀 UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma(UP:用于打破探索-稳定性困境的无界正向非对称优化)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
579
2026.07.09 | 结构可推理,记忆可长存;双突破让AI更懂科学
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 7 篇论文如下:[00:30] 🔬 Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning(基于深度原生结构推理的精确、跨学科且透明的构效关系理解)[01:28] 🤖 Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation(视觉-语言-动作模型中的双潜记忆用于机器人操作)[02:34] 🤖 Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence(面向具身智能的混合专家视频预训练规模化)[03:28] 🌍 Infinite Worlds with Versatile Interactions(无限世界与多样化交互)[04:32] 🤖 RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies(RoboDojo:统一仿真与真实世界的通用机器人操作策略综合评估基准)[05:34] 🏙 WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence(WildCity:一个用于渲染、仿真和空间智能的真实世界城市规模测试平台)[06:35] 🧩 Teaching LLMs a Low-Resource Language: Enhancing Code Completion in Pharo(教会大语言模型一种低资源语言:提升Pharo语言中的代码补全能力)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
578
2026.07.08 | AlayaWorld实时生成长视频;RynnWorld-4D提升机器人操作成功率
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 15 篇论文如下:[00:30] 🎮 AlayaWorld: Long-Horizon and Playable Video World Generation(AlayaWorld:长视界且可游玩的视频世界生成)[01:31] 🤖 RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation(RynnWorld-4D:面向机器人操作的4D具身世界模型)[02:30] 🔍 Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling(分层稀疏注意力机制的正确实现:迈向无限上下文建模)[03:36] 👁 Vision as Unified Multimodal Generation(视觉作为统一多模态生成)[04:25] 🔦 Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory(光之全能:基于长期记忆的智能体视频理解中反思优于推理)[05:25] 🎥 Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning(并行化自回归解码用于全模态密集视频字幕生成)[06:14] 🧠 Gemma 4 Technical Report(Gemma 4 技术报告)[07:11] ⚡ SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe(SkillOpt-Lite:通过一行“氛围”实现更优更快的智能体自我进化)[08:07] ⚡ DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation(DSpark:基于置信度调度的半自回归推测解码)[09:03] 🧠 MentalThink: Shaping Thoughts in Mental SVG World(MentalThink:在思维SVG世界中塑造思想)[09:57] 🤖 RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation(RynnWorld-Teleop:一种用于数字遥操作的动作条件世界模型)[10:58] 🎯 TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training(TurnOPD:使在线知识蒸馏具备回合感知能力以实现高效的长时程智能体训练)[11:58] 🎨 CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration(CanvasAgent:通过可视化工具编排实现复杂图像创建与编辑)[13:05] 🧪 TREK: Distill to Explore, Reinforce to Refine(TREK:通过蒸馏进行探索,通过强化进行精炼)[13:51] 🖼 PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation(PointDiT:面向单目几何估计的像素空间扩散模型)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
577
2026.07.07 | 跨平台智能体学习新范式;科研构思可复用技能提炼
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 14 篇论文如下:[00:32] 🤖 UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning(UI-MOPD:面向持续GUI智能体学习的多平台在线策略蒸馏)[01:30] 💡 ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes(ResearchStudio-Idea:基于证据的科研构思技能套件——来自机器学习会议成果)[02:21] 🎨 PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space(PixWorld:在像素空间中统一3D场景生成与重建)[03:11] 🧩 OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers(OmniOpt:现代优化器的分类、几何结构与基准测试)[04:04] 🤖 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation(GigaWorld-1:构建用于机器人策略评估的世界模型路线图)[04:55] 🧩 Vision Pretraining for Dense Spatial Perception(面向密集空间感知的视觉预训练)[05:54] 🤖 EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots(EVA-Client:面向实体机器人上的具身策略的统一数据收集、推理与部署框架)[06:47] 🎥 Wan-Streamer v0.2: Higher Resolution, Same Latency(Wan-Streamer v0.2:更高分辨率,相同延迟)[07:48] 🤖 InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization(InternVLA-A1.5:统一理解、潜在预知与动作以实现组合泛化)[08:46] 🔍 Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval(所有视觉标记都同等重要吗?面向视觉-语言检索的保留对象证据的标记合并方法)[09:59] 🧠 KVpop -- Key-Value Cache Compression with Predictive Online Pruning(KVpop——基于预测性在线剪枝的键值缓存压缩)[10:50] 🧠 dOPSD: On-Policy Self-Distillation for Diffusion Language Models(dOPSD:扩散语言模型的在线自蒸馏方法)[11:41] 🎨 Perceptual Flow Matching for Few-Step Generative Modeling(感知流匹配:用于少步生成建模)[12:37] 🔬 Multi-Turn Agentic Scientific Literature Search via Workflow Induction(多轮交互式科学文献搜索:基于工作流归纳的方法)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
576
2026.07.06 | 策略对齐破解推理盲区;轻量外挂实现实时修正
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 9 篇论文如下:[00:35] 🎯 The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning(优化训练策略的幻影:单调推理策略作为大语言模型强化学习的真正目标)[01:34] 🛠 VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon(VLA-Corrector:一种用于自适应动作视界的轻量级检测与校正推理框架)[02:35] 🤖 Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots(Embodied.cpp:面向异构机器人的具身AI模型便携推理运行时)[03:43] 🔢 OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers(OrbitQuant:面向图像和视频扩散Transformer的数据无关量化方法)[04:38] 📊 DataComp-VLM: Improved Open Datasets for Vision-Language Models(DataComp-VLM:改进的视觉语言模型开放数据集)[05:36] 🛡 Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming(保护AI智能体:一种多层次智能体红队测试的统一框架)[06:36] ☁ Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment(面向解译的云去除:基于观测锚定残差流与地理上下文对齐的方法)[07:38] 📄 MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering(多注意力归因:长文档问答中的无训练多模态归因)[08:25] 🔗 AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation(AGE:面向图检索增强生成的自适应掩码图嵌入方法)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
-
575
【周末特辑】7月第1周最火AI论文 | Orca:从视频中学习世界模型;智能体弃权:学会何时停止。
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 5 篇论文如下:[00:43] TOP1(🔥230) | 🌍 Orca: The World is in Your Mind(虎鲸:世界在你心中)[03:18] TOP2(🔥141) | 🤔 Agentic Abstention: Do Agents Know When to Stop Instead of Act?(智能体式弃权:智能体知道何时该停止而非行动吗?)[05:32] TOP3(🔥103) | 🧪 Dockerless: Environment-Free Program Verifier for Coding Agents(无Docker:面向编码智能体的无环境程序验证器)[07:58] TOP4(🔥93) | 🎭 DOPD: Dual On-policy Distillation(双在线策略蒸馏)[10:13] TOP5(🔥86) | 🧠 Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent(扩展智能体视野而非参数规模:以35B智能体达到万亿参数级性能)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
We're indexing this podcast's transcripts for the first time — this can take a minute or two. We'll show results as soon as they're ready.
No matches for "" in this podcast's transcripts.
No topics indexed yet for this podcast.
Loading reviews...
ABOUT THIS SHOW
每天10分钟,带您快速了解当日HuggingFace热门AI论文内容。每个工作日更新,欢迎订阅。📢播客节目在小宇宙、Apple Podcast平台搜索【HuggingFace 每日AI论文速递】🖼另外还有图文版,可在小红书搜索并关注【AI速递】
HOSTED BY
duan
CATEGORIES
Loading similar podcasts...