EPISODE · Jul 7, 2026 · 14 MIN
2026.07.07 | 跨平台智能体学习新范式;科研构思可复用技能提炼
from HuggingFace 每日AI论文速递
【赞助商】OpenClaw快报每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43【目录】本期的 14 篇论文如下:[00:32] 🤖 UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning(UI-MOPD:面向持续GUI智能体学习的多平台在线策略蒸馏)[01:30] 💡 ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes(ResearchStudio-Idea:基于证据的科研构思技能套件——来自机器学习会议成果)[02:21] 🎨 PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space(PixWorld:在像素空间中统一3D场景生成与重建)[03:11] 🧩 OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers(OmniOpt:现代优化器的分类、几何结构与基准测试)[04:04] 🤖 GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation(GigaWorld-1:构建用于机器人策略评估的世界模型路线图)[04:55] 🧩 Vision Pretraining for Dense Spatial Perception(面向密集空间感知的视觉预训练)[05:54] 🤖 EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots(EVA-Client:面向实体机器人上的具身策略的统一数据收集、推理与部署框架)[06:47] 🎥 Wan-Streamer v0.2: Higher Resolution, Same Latency(Wan-Streamer v0.2:更高分辨率,相同延迟)[07:48] 🤖 InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization(InternVLA-A1.5:统一理解、潜在预知与动作以实现组合泛化)[08:46] 🔍 Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval(所有视觉标记都同等重要吗?面向视觉-语言检索的保留对象证据的标记合并方法)[09:59] 🧠 KVpop -- Key-Value Cache Compression with Predictive Online Pruning(KVpop——基于预测性在线剪枝的键值缓存压缩)[10:50] 🧠 dOPSD: On-Policy Self-Distillation for Diffusion Language Models(dOPSD:扩散语言模型的在线自蒸馏方法)[11:41] 🎨 Perceptual Flow Matching for Few-Step Generative Modeling(感知流匹配:用于少步生成建模)[12:37] 🔬 Multi-Turn Agentic Scientific Literature Search via Workflow Induction(多轮交互式科学文献搜索:基于工作流归纳的方法)【关注我们】您还可以在以下平台找到我们,获得播客内容以外更多信息小红书: AI速递在小宇宙查看该单集文稿
Embed this episode
Ready to play
2026.07.07 | 跨平台智能体学习新范式;科研构思可复用技能提炼
No transcript for this episode yet
Similar Episodes
No similar episodes found.