具身智能・机器人・无人机前沿论文与项目索引

核验时间:2026-09-01。优先收录论文原文、作者项目页、官方实验室博客和官方文档,不把二手媒体作为关键技术事实来源。

建议:用 Google Scholar 搜标题,再按 Since 2025 / Since 2026Sort by date 追踪引用链和后续工作。


1. 2026:Robot Foundation Model / VLA / Embodied Reasoning

1.1 Gemini Robotics 2

  • Google DeepMind, Gemini Robotics 2 brings whole body intelligence to robots, 2026-07-30.
  • 重点:whole-body intelligence、dexterity、多机器人协作、跨机器人适配。
  • 官方:https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
  • 模型总览:https://deepmind.google/models/gemini-robotics/

1.2 Gemini Robotics ER 2

  • 2026-07-30。
  • 定位:高层 embodied reasoning,负责视频理解、复杂任务规划、机器人协作,再把底层执行交给 VLA。
  • Model card:https://deepmind.google/models/model-cards/gemini-robotics-er-2/

1.3 Gemini Robotics On-Device 2

  • 2026-07-30。
  • 重点:在本地机器人设备运行的 VLA,代表“低延迟、断网可用、端侧具身模型”方向。
  • Model card:https://deepmind.google/models/model-cards/gemini-robotics-on-device-2/

1.4 NVIDIA Isaac GR00T 1.7

  • NVIDIA, Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T, 2026-07-07.
  • 重点:人形 VLA、真实+仿真数据、ONNX/TensorRT、Isaac Lab-Arena、LeRobot 生态。
  • 官方技术博客:https://developer.nvidia.com/blog/develop-humanoid-robot-policies-end-to-end-with-nvidia-isaac-gr00t/
  • NVIDIA 物理 AI 开放模型新闻:https://nvidianews.nvidia.com/news/nvidia-expands-open-model-families-to-power-the-next-wave-of-agentic-physical-and-healthcare-ai

1.5 The Embodiment Gap in Robot Foundation Models

  • Domae et al., 2026-08.
  • 重点:跨本体不是“模型会泛化”就结束,还要考虑动作接口、传感器、控制阶段、对应关系与适配成本。
  • arXiv:https://arxiv.org/abs/2608.18433

1.6 LingBot-VLA 2.0 — From Foundation to Application

  • Wu et al., 2026-07-07。
  • 重点:约 60,000 小时预训练数据,其中约 50,000 小时机器人轨迹覆盖 20 种机器人配置,另有约 10,000 小时第一视角人类视频;扩展头部、腰部、移动底盘、灵巧手等 whole-body 动作空间,并加入预测动力学代理任务。
  • 为什么值得看:它非常直接地体现了 2026 年 VLA 从“实验室 benchmark”走向“多本体、全身动作、预测未来、工程部署”的趋势。
  • arXiv:https://arxiv.org/abs/2607.06403
  • 项目:https://technology.robbyant.com/lingbot-vla-v2
  • 代码:https://github.com/robbyant/lingbot-vla-v2

2. 2026:World Model / World-Action Model

2.1 World Model for Robot Learning: A Comprehensive Survey

  • Hou et al., 2026-04.
  • 重点:机器人世界模型的策略耦合、学习仿真器、视频世界模型、导航/驾驶、数据集与评测。
  • arXiv:https://arxiv.org/abs/2605.00080

2.2 DreamZero — World Action Models are Zero-shot Policies

  • Ye et al., 2026-02.
  • 重点:视频扩散世界模型 + 动作联合预测;强调 zero-shot 任务/环境泛化与跨本体迁移。
  • arXiv:https://arxiv.org/abs/2602.15922
  • 项目:https://dreamzero0.github.io/
  • 代码:https://github.com/dreamzero0/dreamzero

2.3 NVIDIA World Action Model 术语解释

  • 适合理解 WAM 与普通 VLA / World Model 的关系。
  • 官方:https://www.nvidia.com/en-gb/glossary/world-action-model/

3. 2026:VLA 数据、Benchmark 与 Data Engine

3.1 Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines

  • Wang et al., 2026-04.
  • 重点:VLA 的下一阶段瓶颈不仅是模型,而是数据基础设施、跨本体表示对齐、长时序评测与可扩展数据生成。
  • arXiv:https://arxiv.org/abs/2604.23001

3.2 AGIBOT WORLD 2026

  • 2026 数据主题覆盖真实世界 imitation learning、rich interaction、reinforcement learning 等。
  • 重点:100% 真实场景、多模态、力/力矩、错误修正、真机 RL 闭环。
  • 官方:https://agibot-world.com/
  • 2026 Theme 3 RL 介绍:https://www.agibot.com/article/231/detail/88.html

3.3 RoboCasa365

  • ICLR 2026。
  • 重点:365 个家居操作任务、海量场景资产、真人与自动生成演示、generalist policy benchmark。
  • 文档:https://robocasa.ai/docs/introduction/overview.html
  • Leaderboard:https://robocasa.ai/leaderboard.html

3.4 LeRobot v0.6.x

  • 2026-07 版本开始强化 world model policy、reward model、human-in-the-loop correction、统一仿真评测。
  • 官方博客:https://huggingface.co/blog/lerobot-release-v060
  • 文档:https://huggingface.co/docs/lerobot/main/index

4. 2026:触觉与接触智能

4.1 TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation

  • Zhou et al., 2026-07.
  • 重点:把慢速视觉语言规划、触觉世界模型、visuo-tactile action、快速触觉 residual correction 分层。
  • arXiv:https://arxiv.org/abs/2607.07287

5. 2026:无人机具身智能 / Aerial VLN / UAV VLA

5.1 Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

  • Xia et al., 2026-04.
  • 重点:Aerial VLN 综述,系统比较 seq2seq/attention、LLM/VLM end-to-end、hierarchical、multi-agent、dialog-based 方法。
  • 明确提出 long-horizon grounding、viewpoint robustness、6-DoF continuous execution、onboard deployment、swarm navigation 等开放问题。
  • arXiv:https://arxiv.org/abs/2604.07705

5.2 AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild

  • Sun et al., ICLR 2026。
  • 输入:RGB + language;输出可执行 UAV velocity action。
  • 重点:pseudo-depth、粗粒度目标、自主避障/规划、仿真+真实环境。
  • arXiv:https://arxiv.org/abs/2602.09657
  • 项目:https://xiaolousun.github.io/AutoFly/
  • 代码:https://github.com/xiaolousun/AutoFly-VLA

5.3 ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

  • Zhang et al., 2026-07.
  • 重点:Air-EQA → Observation Behavior Planning → Fine-grained Language-guided UAV Control,把主动观察与低层飞行控制串起来。
  • arXiv:https://arxiv.org/abs/2607.10180
  • 项目:https://lvmolvmo.github.io/ActiveFly/

5.4 Vision Language Action Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review

  • Sa et al., 2026-07.
  • 重点:把双臂 VLA 和无人机 VLA 的动作表示、协调、延迟、语言 grounding 等问题放在统一视角比较。
  • arXiv:https://arxiv.org/abs/2607.06706

6. 2024–2025:理解 VLA 演进必须看的代表工作

6.1 RT-2

  • Google DeepMind, 2023。
  • 贡献:把 VLM 的互联网视觉语言知识迁移到机器人动作;动作 token 化。
  • 官方:https://deepmind.google/blog/rt-2-new-model-translates-vision-and-language-into-action/

6.2 Open X-Embodiment / RT-X

  • 多机构、多机器人、统一机器人数据格式与跨本体训练的重要里程碑。
  • 官方仓库:https://github.com/google-deepmind/open_x_embodiment
  • DeepMind 博客:https://deepmind.google/blog/scaling-up-learning-across-many-different-robot-types

6.3 Octo

  • RSS 2024。
  • 重点:开源 generalist robot policy、800k trajectories、多机器人、多观测、多动作空间、扩散动作输出。
  • 项目:https://octo-models.github.io/

6.4 OpenVLA

  • 2024。
  • 重点:7B 级开源 VLA,推动 VLA 研究可复现。
  • 项目:https://openvla.github.io/

6.5 OpenVLA-OFT

  • 重点:parallel decoding、action chunking、continuous action representation、L1 objective,说明“动作解码方案”对实时性非常关键。
  • 项目:https://openvla-oft.github.io/

6.6 π₀

  • Physical Intelligence, 2024。
  • 重点:VLM + action expert + flow matching,跨 8 类机器人数据,面向灵巧操作。
  • 博客:https://www.physicalintelligence.company/blog/pi0
  • 论文:https://www.physicalintelligence.company/download/pi0.pdf

6.7 π₀.₅

  • 2025。
  • 重点:open-world generalization、heterogeneous co-training、长时序家务、新家庭环境。
  • 论文:https://www.physicalintelligence.company/download/pi05.pdf

6.8 SmolVLA

  • Hugging Face, 2025。
  • 450M,强调低成本、开源社区数据、消费级硬件、异步推理、flow-matching action expert。
  • 官方博客:https://huggingface.co/blog/smolvla
  • 文档:https://huggingface.co/docs/lerobot/smolvla

6.9 Diffusion Policy

  • Chi et al., 2023。
  • 贡献:将机器人 visuomotor policy 明确表述为条件扩散动作生成问题。
  • 项目:https://diffusion-policy.cs.columbia.edu/

6.10 ACT / ALOHA

  • Zhao et al., 2023。
  • 贡献:Action Chunking Transformer;低成本双臂遥操作与高精度模仿学习的重要基线。
  • 项目:https://tonyzhaozh.github.io/aloha/

7. 机器人数据集

Open X-Embodiment

  • 多机器人统一 RLDS 格式。
  • https://github.com/google-deepmind/open_x_embodiment

DROID

  • 76k demonstrations / 350h、564 scenes、86 tasks 的 in-the-wild 操作数据。
  • https://droid-dataset.github.io/

BridgeData V2

  • 53,896 trajectories,13 skills,24 environments。
  • https://bridgedata-v2.github.io/

AgiBot World

  • 百万级真实机器人轨迹生态。
  • https://github.com/OpenDriveLab/AgiBot-World

LeRobot Community Datasets

  • 与低成本硬件和统一训练代码紧密结合。
  • https://huggingface.co/lerobot

8. 仿真与 Benchmark

NVIDIA Isaac Sim

  • 高保真 GPU 仿真、OpenUSD、传感器与合成数据。
  • https://developer.nvidia.com/isaac/sim/

NVIDIA Isaac Lab

  • GPU 并行 RL/IL、humanoid、manipulator、AMR。
  • https://developer.nvidia.com/isaac/lab

Habitat / Habitat Lab

  • Embodied navigation、VLN、EQA 经典生态。
  • https://aihabitat.org/

ManiSkill

  • 操作学习与 benchmark。
  • https://maniskill.readthedocs.io/

RLBench

  • 多任务机械臂操作。
  • https://github.com/stepjam/RLBench

RoboCasa / RoboCasa365

  • 家居场景 generalist manipulation。
  • https://robocasa.ai/

AirSim

  • UAV / vehicle 仿真。
  • https://microsoft.github.io/AirSim/

9. ROS 2 / 机械臂 / 无人机工程官方资料

ROS 2

  • https://docs.ros.org/

MoveIt 2

  • https://moveit.picknik.ai/

Nav2

  • https://docs.nav2.org/

PX4

  • 主文档:https://docs.px4.io/
  • Offboard Mode:https://docs.px4.io/main/en/flight_modes/offboard
  • ROS 2 Offboard:https://docs.px4.io/main/en/ros2/offboard_control

MAVSDK

  • https://mavsdk.mavlink.io/

10. Google Scholar 建议搜索词

具身基础模型

"robot foundation model" vision language action 2026
"embodied reasoning" robotics 2026
"cross-embodiment" robot foundation model
"generalist robot policy" 2026

世界模型

"robot world model" 2026
"world action model" robotics
"video world model" robot manipulation
"latent world model" robot learning

动作生成

"flow matching" robot policy
"diffusion policy" robot manipulation
"action tokenizer" vision language action
"action chunking" robot transformer

触觉

"tactile foundation model" robotics
"visuotactile" dexterous manipulation
"tactile world model" robot

无人机

"aerial vision language navigation" 2026
"UAV vision language action" 2026
"aerial embodied intelligence" UAV
"active perception" UAV VLM
"multi-UAV" vision language model navigation

数据

"robot data engine" VLA
"robot dataset" cross embodiment 2026
"real-world reinforcement learning dataset" robot 2026

11. 跟踪前沿时的筛选原则

不要只按“新”筛论文。优先看:

  1. 是否真机;
  2. 是否报告动作空间与控制频率;
  3. 是否测新环境/新物体/新任务组合;
  4. 是否包含长时序任务;
  5. 是否给出数据规模;
  6. 是否报告 latency;
  7. 是否有 code/model/dataset;
  8. 是否与强基线公平比较;
  9. 是否分析失败模式;
  10. 是否说明安全与人工干预。

这比只看排行榜或宣传视频更可靠。

Logo

DAMO开发者矩阵,由阿里巴巴达摩院和中国互联网协会联合发起,致力于探讨最前沿的技术趋势与应用成果,搭建高质量的交流与分享平台,推动技术创新与产业应用链接,围绕“人工智能与新型计算”构建开放共享的开发者生态。

更多推荐