【论文速递】2026年第06周(Feb-01-07)(Robotics/Embodied AI/LLM)
中文使用 googletrans 翻译,翻译不对的地方以英文为准
目录
- Green-VLA:多面手机器人的分阶段视觉-语言-动作模型
- Kimi K2.5:视觉代理智能
- ERNIE 5.0技术报告
- PaperBanana:为人工智能科学家自动化学术插图
- Vision-DeepResearch:激励多模态大语言模型的深度研究能力
- FASA:频率感知稀疏注意力
- Vision-DeepResearch 基准:重新思考多模态大语言模型的视觉和文本搜索
- Golden Goose:从不可验证的互联网文本合成无限 RLVR 任务的简单技巧
- WideSeek-R1:通过多智能体强化学习探索宽度缩放以实现广泛的信息搜索
- CodeOCR:论视觉语言模型在代码理解中的有效性
- AOrchestra:自动创建子代理以进行代理编排
- CAR-bench:评估现实世界不确定性下 LLM 代理的一致性和限制意识
- 闭环:使用 RPG 编码器表示通用存储库
- DFlash:用于闪存推测解码的块扩散
- 多模式过程奖励模型中的训练数据效率
- UniReason 1.0:世界知识对齐图像生成和编辑的统一推理框架
- 思想链中没有全局计划:揭示法学硕士的潜在规划视野
- Spider-Sense:通过分层自适应筛选进行有效代理防御的内在风险感知
- MARS:用于自动化人工智能研究的具有反射搜索的模块化代理
- 用于生成视图自适应人类视频的 3D 感知隐式运动控制
- MemSkill:自我进化智能体的学习和进化记忆技能
- SWE-Universe:将真实世界的可验证环境扩展到数百万个
- 四方二:通过改进的无偏梯度估计在 NVFP4 中进行准确的 LLM 预训练
- ASTRA:代理轨迹和强化竞技场的自动合成
- 长度无偏序列策略优化:揭示和控制 RLVR 中的响应长度变化
- daVinci-Agency:有效解锁长期机构数据
Green-VLA:多面手机器人的分阶段视觉-语言-动作模型
- 标题: Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
- 作者: I. Apanasevich, M. Artemyev, R. Babakyan, P. Fedotova, D. Grankin, E. Kupryashin, A. Misailidi, D. Nerus, A. Nutalapati, G. Sidorov, I. Efremov, M. Gerasyov, D. Pikurov, Y. Senchenko, S. Davidenko, D. Kulikov, M. Sultankin, K. Askarbek, O. Shamanin, D. Statovoy, E. Zalyaev, I. Zorin, A. Letkin, E. Rusakov, A. Silchenko, V. Vorobyov, S. Sobolnikov, A. Postnikov
- 日期: 2026-01-31
- ArXiv主页: https://arxiv.org/abs/2602.00919
- 论文链接: https://arxiv.org/pdf/2602.00919
- 项目链接: https://greenvla.github.io/
- gitHub仓库: https://github.com/greenvla/GreenVLA
英文摘要
We introduce Green-VLA, a staged Vision-Language-Action (VLA) framework for real-world deployment on the Green humanoid robot while maintaining generalization across diverse embodiments. Green-VLA follows a five stage curriculum: (L0) foundational VLMs, (L1) multimodal grounding, (R0) multi-embodiment pretraining, (R1) embodiment-specific adaptation, and (R2) reinforcement-learning (RL) policy alignment. We couple a scalable data-processing pipeline (3,000 hours of demonstrations) with temporal alignment and quality filtering, and use a unified, embodiment-aware action interface enabling a single policy to control humanoids, mobile manipulators, and fixed-base arms. At inference, the VLA controller is enhanced with episode-progress prediction, out-of-distribution detection, and joint-prediction-based guidance to improve safety and precise target selection. Experiments on Simpler BRIDGE WidowX and CALVIN ABC-D, as well as real-robot evaluations, show strong generalization and performance gains from RL alignment in success rate, robustness, and long-horizon efficiency.
中文摘要
我们引入了 Green-VLA,这是一个分阶段的视觉-语言-动作 (VLA) 框架,用于在 Green 人形机器人上进行实际部署,同时保持跨不同实施例的通用性。Green-VLA 遵循五个阶段的课程:(L0) 基础 VLM、(L1) 多模式基础、(R0) 多实施例预训练、(R1) 特定实施例适应和 (R2) 强化学习 (RL) 策略调整。我们将可扩展的数据处理管道(3,000 小时的演示)与时间对齐和质量过滤结合起来,并使用统一的、可感知实施例的操作界面,支持单一策略来控制人形机器人、移动操纵器和固定底座手臂。在推理时,VLA 控制器通过事件进展预测、分布外检测和基于联合预测的引导得到增强,以提高安全性和精确的目标选择。Simpler BRIDGE WidowX 和 CALVIN ABC-D 上的实验以及真实机器人评估表明,RL 对齐在成功率、鲁棒性和长期效率方面具有很强的泛化性和性能增益。
Kimi K2.5:视觉代理智能
- 标题: Kimi K2.5: Visual Agentic Intelligence
- 作者: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, S. H. Cai, Yuan Cao, Y. Charles, H. S. Che, Cheng Chen, Guanduo Chen, Huarong Chen, Jia Chen, Jiahao Chen, Jianlong Chen, Jun Chen, Kefan Chen, Liang Chen, Ruijue Chen, Xinhao Chen, Yanru Chen, Yanxu Chen, Yicun Chen, Yimin Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Ziwei Chen, Dazhi Cheng, Minghan Chu, Jialei Cui, Jiaqi Deng, Muxi Diao, Hao Ding, Mengfan Dong, Mengnan Dong, Yuxin Dong, Yuhao Dong, Angang Du, Chenzhuang Du, Dikang Du, Lingxiao Du, Yulun Du, Yu Fan, Shengjun Fang, Qiulin Feng, Yichen Feng, Garimugai Fu, Kelin Fu, Hongcheng Gao, Tong Gao, Yuyao Ge, Shangyi Geng, Chengyang Gong, Xiaochen Gong, Zhuoma Gongque, Qizheng Gu, Xinran Gu, Yicheng Gu, Longyu Guan, Yuanying Guo, Xiaoru Hao, Weiran He, Wenyang He, Yunjia He, Chao Hong, Hao Hu, Jiaxi Hu, Yangyang Hu, Zhenxing Hu, Ke Huang, Ruiyuan Huang, Weixiao Huang, Zhiqi Huang, Tao Jiang, Zhejun Jiang, Xinyi Jin, Yu Jing, Guokun Lai, Aidi Li, C. Li, Cheng Li, Fang Li, Guanghe Li, Guanyu Li, Haitao Li, Haoyang Li, Jia Li, Jingwei Li, Junxiong Li, Lincan Li, Mo Li, Weihong Li, Wentao Li, Xinhang Li, Xinhao Li, Yang Li, Yanhao Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zheming Li, Weilong Liao, Jiawei Lin, Xiaohan Lin, Zhishan Lin, Zichao Lin, Cheng Liu, Chenyu Liu, Hongzhang Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Tianyu Liu, Weizhou Liu, Xiangyan Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yuanxin Liu, Yue Liu, Zhengying Liu, Zhongnuo Liu, Enzhe Lu, Haoyu Lu, Zhiyuan Lu, Junyu Luo, Tongxu Luo, Yashuo Luo, Long Ma, Yingwei Ma, Shaoguang Mao, Yuan Mei, Xin Men, Fanqing Meng, Zhiyong Meng, Yibo Miao, Minqing Ni, Kun Ouyang, Siyuan Pan, Bo Pang, Yuchao Qian, Ruoyu Qin, Zeyu Qin, Jiezhong Qiu, Bowen Qu, Zeyu Shang, Youbo Shao, Tianxiao Shen, Zhennan Shen, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Feifan Song, Pengwei Song, Tianhui Song, Xiaoxi Song, Hongjin Su, Jianlin Su, Zhaochen Su, Lin Sui, Jinsong Sun, Junyao Sun, Tongyu Sun, Flood Sung, Yunpeng Tai, Chuning Tang, Heyi Tang, Xiaojuan Tang, Zhengyang Tang, Jiawen Tao, Shiyuan Teng, Chaoran Tian, Pengfei Tian, Ao Wang, Bowen Wang, Chensi Wang, Chuang Wang, Congcong Wang, Dingkun Wang, Dinglu Wang, Dongliang Wang, Feng Wang, Hailong Wang, Haiming Wang, Hengzhi Wang, Huaqing Wang, Hui Wang, Jiahao Wang, Jinhong Wang, Jiuzheng Wang, Kaixin Wang, Linian Wang, Qibin Wang, Shengjie Wang, Shuyi Wang, Si Wang, Wei Wang, Xiaochen Wang, Xinyuan Wang, Yao Wang, Yejie Wang, Yipu Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhexu Wang, Zihan Wang, Zizhe Wang, Chu Wei, Ming Wei, Chuan Wen, Zichen Wen, Chengjie Wu, Haoning Wu, Junyan Wu, Rucong Wu, Wenhao Wu, Yuefeng Wu, Yuhao Wu, Yuxin Wu, Zijian Wu, Chenjun Xiao, Jin Xie, Xiaotong Xie, Yuchong Xie, Yifei Xin, Bowei Xing, Boyu Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L. H. Xu, Lin Xu, Suting Xu, Weixin Xu, Xinbo Xu, Xinran Xu, Yangchuan Xu, Yichang Xu, Yuemeng Xu, Zelai Xu, Ziyao Xu, Junjie Yan, Yuzi Yan, Guangyao Yang, Hao Yang, Junwei Yang, Kai Yang, Ningyuan Yang, Ruihan Yang, Xiaofei Yang, Xinlong Yang, Ying Yang, Yi Yang, Yi Yang, Zhen Yang, Zhilin Yang, Zonghan Yang, Haotian Yao, Dan Ye, Wenjie Ye, Zhuorui Ye, Bohong Yin, Chengzhen Yu, Longhui Yu, Tao Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Xiaokun Yuan, Yang Yue, Weihao Zeng, Dunyuan Zha, Haobing Zhan, Dehao Zhang, Hao Zhang, Jin Zhang, Puqi Zhang, Qiao Zhang, Rui Zhang, Xiaobin Zhang, Y. Zhang, Yadong Zhang, Yangkun Zhang, Yichi Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yushun Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Chenguang Zhao, Feifan Zhao, Jinxiang Zhao, Shuai Zhao, Xiangyu Zhao, Yikai Zhao, Zijia Zhao, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Junfeng Zhong, Longguang Zhong, Weiming Zhong, M. Zhou, Runjie Zhou, Xinyu Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yuxuan Zhu, Zhen Zhu, Jingze Zhuang, Weiyu Zhuang, Ying Zou, Xinxing Zu
- 日期: 2026-02-02
- ArXiv主页: https://arxiv.org/abs/2602.02276
- 论文链接: https://arxiv.org/pdf/2602.02276
- 项目链接: https://www.kimi.com/blog/kimi-k2-5.html
- gitHub仓库: https://github.com/MoonshotAI/Kimi-K2.5
英文摘要
We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5 introduces Agent Swarm, a self-directed parallel agent orchestration framework that dynamically decomposes complex tasks into heterogeneous sub-problems and executes them concurrently. Extensive evaluations show that Kimi K2.5 achieves state-of-the-art results across various domains including coding, vision, reasoning, and agentic tasks. Agent Swarm also reduces latency by up to 4.5times over single-agent baselines. We release the post-trained Kimi K2.5 model checkpoint to facilitate future research and real-world applications of agentic intelligence.
中文摘要
我们推出 Kimi K2.5,这是一种开源多模式代理模型,旨在推进通用代理智能。K2.5强调文本和视觉的联合优化,使两种方式相互促进。这包括联合文本视觉预训练、零视觉SFT、联合文本视觉强化学习等一系列技术。在此多模态基础上,K2.5 引入了 Agent Swarm,这是一种自主并行代理编排框架,可动态地将复杂任务分解为异构子问题并同时执行它们。广泛的评估表明,Kimi K2.5 在编码、视觉、推理和代理任务等各个领域都取得了最先进的结果。与单代理基线相比,Agent Swarm 的延迟时间最多可减少 4.5 倍。我们发布了训练后的 Kimi K2.5 模型检查点,以促进代理智能的未来研究和实际应用。
ERNIE 5.0技术报告
- 标题: ERNIE 5.0 Technical Report
- 作者: Haifeng Wang, Hua Wu, Tian Wu, Yu Sun, Jing Liu, Dianhai Yu, Yanjun Ma, Jingzhou He, Zhongjun He, Dou Hong, Qiwen Liu, Shuohuan Wang, Junyuan Shang, Zhenyu Zhang, Yuchen Ding, Jinle Zeng, Jiabin Yang, Liang Shen, Ruibiao Chen, Weichong Yin, Siyu Ding, Dai Dai, Shikun Feng, Siqi Bao, Bolei He, Yan Chen, Zhenyu Jiao, Ruiqing Zhang, Zeyu Chen, Qingqing Dang, Kaipeng Deng, Jiajun Jiang, Enlei Gong, Guoxia Wang, Yanlin Sha, Yi Liu, Yehan Zheng, Weijian Xu, Jiaxiang Liu, Zengfeng Zeng, Yingqi Qu, Zhongli Li, Zhengkun Zhang, Xiyang Wang, Zixiang Xu, Xinchao Xu, Zhengjie Huang, Dong Wang, Bingjin Chen, Yue Chang, Xing Yuan, Shiwei Huang, Qiao Zhao, Xinzhe Ding, Shuangshuang Qiao, Baoshan Yang, Bihong Tang, Bin Li, Bingquan Wang, Binhan Tang, Binxiong Zheng, Bo Cui, Bo Ke, Bo Zhang, Bowen Zhang, Boyan Zhang, Boyang Liu, Caiji Zhang, Can Li, Chang Xu, Chao Pang, Chao Zhang, Chaoyi Yuan, Chen Chen, Cheng Cui, Chenlin Yin, Chun Gan, Chunguang Chai, Chuyu Fang, Cuiyun Han, Dan Zhang, Danlei Feng, Danxiang Zhu, Dong Sun, Dongbo Li, Dongdong Li, Dongdong Liu, Dongxue Liu, Fan Ding, Fan Hu, Fan Li, Fan Mo, Feisheng Wu, Fengwei Liu, Gangqiang Hu, Gaofeng Lu, Gaopeng Yong, Gexiao Tian, Guan Wang, Guangchen Ni, Guangshuo Wu, Guanzhong Wang, Guihua Liu, Guishun Li, Haibin Li, Haijian Liang, Haipeng Ming, Haisu Wang, Haiyang Lu, Haiye Lin, Han Zhou, Hangting Lou, Hanwen Du, Hanzhi Zhang, Hao Chen, Hao Du, Hao Liu, Hao Zhou, Haochen Jiang, Haodong Tian, Haoshuang Wang, Haozhe Geng, Heju Yin, Hong Chen, Hongchen Xue, Hongen Liu, Honggeng Zhang, Hongji Xu, Hongwei Chen, Hongyang Zhang, Hongyuan Zhang, Hua Lu, Huan Chen, Huan Wang, Huang He, Hui Liu, Hui Zhong, Huibin Ruan, Jiafeng Lu, Jiage Liang, Jiahao Hu, Jiahao Hu, Jiajie Yang, Jialin Li, Jian Chen, Jian Wu, Jianfeng Yang, Jianguang Jiang, Jianhua Wang, Jianye Chen, Jiaodi Liu, Jiarui Zhou, Jiawei Lv, Jiaxin Zhou, Jiaxuan Liu, Jie Han, Jie Sun, Jiefan Fang, Jihan Liu, Jihua Liu, Jing Hu, Jing Qian, Jing Yan, Jingdong Du, Jingdong Wang, Jingjing Wu, Jingyong Li, Jinheng Wang, Jinjin Li, Jinliang Lu, Jinlin Yu, Jinnan Liu, Jixiang Feng, Jiyi Huang, Jiyuan Zhang, Jun Liang, Jun Xia, Jun Yu, Junda Chen, Junhao Feng, Junhong Xiang, Junliang Li, Kai Liu, Kailun Chen, Kairan Su, Kang Hu, Kangkang Zhou, Ke Chen, Ke Wei, Kui Huang, Kun Wu, Kunbin Chen, Lei Han, Lei Sun, Lei Wen, Linghui Meng, Linhao Yu, Liping Ouyang, Liwen Zhang, Longbin Ji, Longzhi Wang, Meng Sun, Meng Tian, Mengfei Li, Mengqi Zeng, Mengyu Zhang, Ming Hong, Mingcheng Zhou, Mingming Huang, Mingxin Chen, Mingzhu Cai, Naibin Gu, Nemin Qiu, Nian Wang, Peng Qiu, Peng Zhao, Pengyu Zou, Qi Wang, Qi Xin, Qian Wang, Qiang Zhu, Qianhui Luo, Qianwei Yang, Qianyue He, Qifei Wu, Qinrui Li, Qiwen Bao, Quan Zhang, Quanxiang Liu, Qunyi Xie, Rongrui Zhan, Rufeng Dai, Rui Peng, Ruian Liu, Ruihao Xu, Ruijie Wang, Ruixi Zhang, Ruixuan Liu, Runsheng Shi, Ruting Wang, Senbo Kang, Shan Lu, Shaofei Yu, Shaotian Gong, Shenwei Hu, Shifeng Zheng, Shihao Guo, Shilong Fan, Shiqin Liu, Shiwei Gu, Shixi Zhang, Shuai Yao, Shuang Zhang, Shuangqiao Liu, Shuhao Liang, Shuwei He, Shuwen Yang, Sijun He, Siming Dai, Siming Wu, Siyi Long, Songhe Deng, Suhui Dong, Suyin Liang, Teng Hu, Tianchan Xu, Tianliang Lv, Tianmeng Yang, Tianyi Wei, Tiezhu Gao, Ting Sun, Ting Zhang, Tingdan Luo, Wei He, Wei Luan, Wei Yin, Wei Zhang, Wei Zhou, Weibao Gong, Weibin Li, Weicheng Huang, Weichong Dang, Weiguo Zhu, Weilong Zhang, Weiqi Tan, Wen Huang, Wenbin Chang, Wenjing Du, Wenlong Miao, Wenpei Luo, Wenquan Wu, Xi Shi, Xi Zhao, Xiang Gao, Xiangguo Zhang, Xiangrui Yu, Xiangsen Wang, Xiangzhe Wang, Xianlong Luo, Xianying Ma, Xiao Tan, Xiaocong Lin, Xiaofei Wang, Xiaofeng Peng, Xiaofeng Wu, Xiaojian Xu, Xiaolan Yuan, Xiaopeng Cui, Xiaotian Han, Xiaoxiong Liu, Xiaoxu Fei, Xiaoxuan Wu, Xiaoyu Wang, Xiaoyu Zhang, Xin Sun, Xin Wang, Xinhui Huang, Xinming Zhu, Xintong Yu, Xinyi Xu, Xinyu Wang, Xiuxian Li, XuanShi Zhu, Xue Xu, Xueying Lv, Xuhong Li, Xulong Wei, Xuyi Chen, Yabing Shi, Yafeng Wang, Yamei Li, Yan Liu, Yanfu Cheng, Yang Gao, Yang Liang, Yang Wang, Yang Wang, Yang Yang, Yanlong Liu, Yannian Fu, Yanpeng Wang, Yanzheng Lin, Yao Chen, Yaozong Shen, Yaqian Han, Yehua Yang, Yekun Chai, Yesong Wang, Yi Song, Yichen Zhang, Yifei Wang, Yifeng Guo, Yifeng Kou, Yilong Chen, Yilong Guo, Yiming Wang, Ying Chen, Ying Wang, Yingsheng Wu, Yingzhan Lin, Yinqi Yang, Yiran Xing, Yishu Lei, Yixiang Tu, Yiyan Chen, Yong Zhang, Yonghua Li, Yongqiang Ma, Yongxing Dai, Yongyue Zhang, Yu Ran, Yu Sun, Yu-Wen Michael Zhang, Yuang Liu, Yuanle Liu, Yuanyuan Zhou, Yubo Zhang, Yuchen Han, Yucheng Wang, Yude Gao, Yuedong Luo, Yuehu Dong, Yufeng Hu, Yuhui Cao, Yuhui Yun, Yukun Chen, Yukun Gao, Yukun Li, Yumeng Zhang, Yun Fan, Yun Ma, Yunfei Zhang, Yunshen Xie, Yuping Xu, Yuqin Zhang, Yuqing Liu, Yurui Li, Yuwen Wang, Yuxiang Lu, Zefeng Cai, Zelin Zhao, Zelun Zhang, Zenan Lin, Zezhao Dong, Zhaowu Pan, Zhaoyu Liu, Zhe Dong, Zhe Zhang, Zhen Zhang, Zhengfan Wu, Zhengrui Wei, Zhengsheng Ning, Zhenxing Li, Zhenyu Li, Zhenyu Qian, Zhenyun Li, Zhi Li, Zhichao Chen, Zhicheng Dong, Zhida Feng, Zhifan Feng, Zhihao Deng, Zhijin Yu, Zhiyang Chen, Zhonghui Zheng, Zhuangzhuang Guo, Zhujun Zhang, Zhuo Sun, Zichang Liu, Zihan Lin, Zihao Huang, Zihe Zhu, Ziheng Zhao, Ziping Chen, Zixuan Zhu, Ziyang Xu, Ziyi Liang, Ziyuan Gao
- 日期: 2026-02-04
- ArXiv主页: https://arxiv.org/abs/2602.04705
- 论文链接: https://arxiv.org/pdf/2602.04705
英文摘要
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio. All modalities are trained from scratch under a unified next-group-of-tokens prediction objective, based on an ultra-sparse mixture-of-experts (MoE) architecture with modality-agnostic expert routing. To address practical challenges in large-scale deployment under diverse resource constraints, ERNIE 5.0 adopts a novel elastic training paradigm. Within a single pre-training run, the model learns a family of sub-models with varying depths, expert capacities, and routing sparsity, enabling flexible trade-offs among performance, model size, and inference latency in memory- or time-constrained scenarios. Moreover, we systematically address the challenges of scaling reinforcement learning to unified foundation models, thereby guaranteeing efficient and stable post-training under ultra-sparse MoE architectures and diverse multimodal settings. Extensive experiments demonstrate that ERNIE 5.0 achieves strong and balanced performance across multiple modalities. To the best of our knowledge, among publicly disclosed models, ERNIE 5.0 represents the first production-scale realization of a trillion-parameter unified autoregressive model that supports both multimodal understanding and generation. To facilitate further research, we present detailed visualizations of modality-agnostic expert routing in the unified model, alongside comprehensive empirical analysis of elastic training, aiming to offer profound insights to the community.
中文摘要
在本报告中,我们介绍了 ERNIE 5.0,这是一种原生自回归基础模型,旨在实现跨文本、图像、视频和音频的统一多模态理解和生成。所有模态都是在统一的下一组令牌预测目标下从头开始训练,基于具有模态不可知专家路由的超稀疏专家混合 (MoE) 架构。为了解决不同资源约束下大规模部署的实际挑战,ERNIE 5.0采用了一种新颖的弹性训练范式。在一次预训练运行中,该模型学习一系列具有不同深度、专家能力和路由稀疏性的子模型,从而在内存或时间受限的场景中实现性能、模型大小和推理延迟之间的灵活权衡。此外,我们系统地解决了将强化学习扩展到统一基础模型的挑战,从而保证了超稀疏 MoE 架构和多样化多模态设置下的高效稳定的后训练。大量实验表明,ERNIE 5.0 在多种模式下实现了强大且平衡的性能。据我们所知,在公开披露的模型中,ERNIE 5.0 代表了第一个支持多模态理解和生成的万亿参数统一自回归模型的生产规模实现。为了促进进一步的研究,我们在统一模型中提供了与模态无关的专家路由的详细可视化,以及弹性训练的全面实证分析,旨在为社区提供深刻的见解。
PaperBanana:为人工智能科学家自动化学术插图
- 标题: PaperBanana: Automating Academic Illustration for AI Scientists
- 作者: Dawei Zhu, Rui Meng, Yale Song, Xiyu Wei, Sujian Li, Tomas Pfister, Jinsung Yoon
- 日期: 2026-01-30
- ArXiv主页: https://arxiv.org/abs/2601.23265
- 论文链接: https://arxiv.org/pdf/2601.23265
- 项目链接: https://dwzhu-pku.github.io/PaperBanana/
- gitHub仓库: https://github.com/llmsresearch/paperbanana
英文摘要
Despite rapid advances in autonomous AI scientists powered by language models, generating publication-ready illustrations remains a labor-intensive bottleneck in the research workflow. To lift this burden, we introduce PaperBanana, an agentic framework for automated generation of publication-ready academic illustrations. Powered by state-of-the-art VLMs and image generation models, PaperBanana orchestrates specialized agents to retrieve references, plan content and style, render images, and iteratively refine via self-critique. To rigorously evaluate our framework, we introduce PaperBananaBench, comprising 292 test cases for methodology diagrams curated from NeurIPS 2025 publications, covering diverse research domains and illustration styles. Comprehensive experiments demonstrate that PaperBanana consistently outperforms leading baselines in faithfulness, conciseness, readability, and aesthetics. We further show that our method effectively extends to the generation of high-quality statistical plots. Collectively, PaperBanana paves the way for the automated generation of publication-ready illustrations.
中文摘要
尽管在语言模型的支持下,自主人工智能科学家取得了快速进步,但生成可发表的插图仍然是研究工作流程中的劳动密集型瓶颈。为了减轻这一负担,我们引入了 PaperBanana,这是一个用于自动生成可供出版的学术插图的代理框架。在最先进的 VLM 和图像生成模型的支持下,PaperBanana 协调专门的代理来检索参考、规划内容和风格、渲染图像,并通过自我批评进行迭代完善。为了严格评估我们的框架,我们引入了 PaperBananaBench,其中包含来自 NeurIPS 2025 出版物的 292 个方法图测试用例,涵盖不同的研究领域和插图风格。综合实验表明,PaperBanana 在忠实性、简洁性、可读性和美观性方面始终优于领先的基线。我们进一步表明,我们的方法可以有效地扩展到生成高质量的统计图。总的来说,PaperBanana 为自动生成可供出版的插图铺平了道路。
Vision-DeepResearch:激励多模态大语言模型的深度研究能力
- 标题: Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models
- 作者: Wenxuan Huang, Yu Zeng, Qiuchen Wang, Zhen Fang, Shaosheng Cao, Zheng Chu, Qingyu Yin, Shuang Chen, Zhenfei Yin, Lin Chen, Zehui Chen, Yao Hu, Philip Torr, Feng Zhao, Wanli Ouyang
- 日期: 2026-01-29
- ArXiv主页: https://arxiv.org/abs/2601.22060
- 论文链接: https://arxiv.org/pdf/2601.22060
- 项目链接: https://osilly.github.io/Vision-DeepResearch/
- gitHub仓库: https://github.com/Osilly/Vision-DeepResearch
英文摘要
Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity of their internal world knowledge, prior work has proposed augmenting MLLMs by ``reasoning-then-tool-call’’ for visual and textual search engines to obtain substantial gains on tasks requiring extensive factual information. However, these approaches typically define multimodal search in a naive setting, assuming that a single full-level or entity-level image query and few text query suffices to retrieve the key evidence needed to answer the question, which is unrealistic in real-world scenarios with substantial visual noise. Moreover, they are often limited in the reasoning depth and search breadth, making it difficult to solve complex questions that require aggregating evidence from diverse visual and textual sources. Building on this, we propose Vision-DeepResearch, which proposes one new multimodal deep-research paradigm, i.e., performs multi-turn, multi-entity and multi-scale visual and textual search to robustly hit real-world search engines under heavy noise. Our Vision-DeepResearch supports dozens of reasoning steps and hundreds of engine interactions, while internalizing deep-research capabilities into the MLLM via cold-start supervision and RL training, resulting in a strong end-to-end multimodal deep-research MLLM. It substantially outperforming existing multimodal deep-research MLLMs, and workflows built on strong closed-source foundation model such as GPT-5, Gemini-2.5-pro and Claude-4-Sonnet. The code will be released in https://github.com/Osilly/Vision-DeepResearch.
中文摘要
多模态大语言模型(MLLM)在广泛的视觉任务中取得了显着的成功。然而,受其内部世界知识能力的限制,先前的工作提出通过视觉和文本搜索引擎的“推理然后工具调用”来增强 MLLM,以便在需要大量事实信息的任务中获得实质性收益。然而,这些方法通常在幼稚的环境中定义多模态搜索,假设单个全级或实体级图像查询和少量文本查询足以检索回答问题所需的关键证据,这在具有大量视觉噪声的现实场景中是不现实的。此外,它们的推理深度和搜索广度通常受到限制,使得解决需要从不同视觉和文本来源聚合证据的复杂问题变得困难。在此基础上,我们提出了Vision-DeepResearch,它提出了一种新的多模态深度研究范式,即执行多回合、多实体、多尺度的视觉和文本搜索,以在强噪声下稳健地打击现实世界的搜索引擎。我们的 Vision-DeepResearch 支持数十个推理步骤和数百个引擎交互,同时通过冷启动监督和 RL 训练将深度研究能力内化到 MLLM 中,从而形成强大的端到端多模态深度研究 MLLM。它的性能远远优于现有的多模态深度研究 MLLM,以及基于强大的闭源基础模型(例如 GPT-5、Gemini-2.5-pro 和 Claude-4-Sonnet)构建的工作流程。代码将在 https://github.com/Osilly/Vision-DeepResearch 发布。
FASA:频率感知稀疏注意力
- 标题: FASA: Frequency-aware Sparse Attention
- 作者: Yifei Wang, Yueqi Wang, Zhenrui Yue, Huimin Zeng, Yong Wang, Ismini Lourentzou, Zhengzhong Tu, Xiangxiang Chu, Julian McAuley
- 日期: 2026-02-03
- ArXiv主页: https://arxiv.org/abs/2602.03152
- 论文链接: https://arxiv.org/pdf/2602.03152
英文摘要
The deployment of Large Language Models (LLMs) faces a critical bottleneck when handling lengthy inputs: the prohibitive memory footprint of the Key Value (KV) cache. To address this bottleneck, the token pruning paradigm leverages attention sparsity to selectively retain a small, critical subset of tokens. However, existing approaches fall short, with static methods risking irreversible information loss and dynamic strategies employing heuristics that insufficiently capture the query-dependent nature of token importance. We propose FASA, a novel framework that achieves query-aware token eviction by dynamically predicting token importance. FASA stems from a novel insight into RoPE: the discovery of functional sparsity at the frequency-chunk (FC) level. Our key finding is that a small, identifiable subset of “dominant” FCs consistently exhibits high contextual agreement with the full attention head. This provides a robust and computationally free proxy for identifying salient tokens. %making them a powerful and efficient proxy for token importance. Building on this insight, FASA first identifies a critical set of tokens using dominant FCs, and then performs focused attention computation solely on this pruned subset. % Since accessing only a small fraction of the KV cache, FASA drastically lowers memory bandwidth requirements and computational cost. Across a spectrum of long-context tasks, from sequence modeling to complex CoT reasoning, FASA consistently outperforms all token-eviction baselines and achieves near-oracle accuracy, demonstrating remarkable robustness even under constraint budgets. Notably, on LongBench-V1, FASA reaches nearly 100% of full-KV performance when only keeping 256 tokens, and achieves 2.56times speedup using just 18.9% of the cache on AIME24.
中文摘要
大型语言模型 (LLM) 的部署在处理冗长的输入时面临着一个关键瓶颈:键值 (KV) 缓存的内存占用过高。为了解决这个瓶颈,令牌修剪范例利用注意力稀疏性来选择性地保留一小部分关键的令牌子集。然而,现有方法存在不足,静态方法存在不可逆信息丢失的风险,而动态策略采用启发式方法,无法充分捕获令牌重要性的查询依赖性质。我们提出了 FASA,这是一种新颖的框架,通过动态预测令牌重要性来实现查询感知令牌驱逐。FASA 源于对 RoPE 的新颖见解:发现频率块 (FC) 级别的功能稀疏性。我们的主要发现是,一小部分可识别的“主导”FC 始终表现出与全注意力头的高度上下文一致性。这提供了一个强大且无需计算的代理来识别显着标记。%使它们成为代币重要性的强大且高效的代理。基于这一见解,FASA 首先使用主导 FC 识别一组关键令牌,然后仅对这个修剪后的子集执行集中注意力计算。由于仅访问 KV 缓存的一小部分,FASA 大大降低了内存带宽要求和计算成本。在一系列长上下文任务中,从序列建模到复杂的 CoT 推理,FASA 始终优于所有令牌驱逐基线,并实现接近预言机的准确性,即使在预算有限的情况下也表现出卓越的稳健性。值得注意的是,在 LongBench-V1 上,FASA 仅保留 256 个令牌即可达到近 100% 的全 KV 性能,并且在 AIME24 上仅使用 18.9% 的缓存即可实现 2.56 倍的加速。
Vision-DeepResearch 基准:重新思考多模态大语言模型的视觉和文本搜索
- 标题: Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
- 作者: Yu Zeng, Wenxuan Huang, Zhen Fang, Shuang Chen, Yufan Shen, Yishuo Cai, Xiaoman Wang, Zhenfei Yin, Lin Chen, Zehui Chen, Shiting Huang, Yiming Zhao, Yao Hu, Philip Torr, Wanli Ouyang, Shaosheng Cao
- 日期: 2026-02-02
- ArXiv主页: https://arxiv.org/abs/2602.02185
- 论文链接: https://arxiv.org/pdf/2602.02185
- 项目链接: https://osilly.github.io/Vision-DeepResearch/
英文摘要
Multimodal Large Language Models (MLLMs) have advanced VQA and now support Vision-DeepResearch systems that use search engines for complex visual-textual fact-finding. However, evaluating these visual and textual search abilities is still difficult, and existing benchmarks have two major limitations. First, existing benchmarks are not visual search-centric: answers that should require visual search are often leaked through cross-textual cues in the text questions or can be inferred from the prior world knowledge in current MLLMs. Second, overly idealized evaluation scenario: On the image-search side, the required information can often be obtained via near-exact matching against the full image, while the text-search side is overly direct and insufficiently challenging. To address these issues, we construct the Vision-DeepResearch benchmark (VDR-Bench) comprising 2,000 VQA instances. All questions are created via a careful, multi-stage curation pipeline and rigorous expert review, designed to assess the behavior of Vision-DeepResearch systems under realistic real-world conditions. Moreover, to address the insufficient visual retrieval capabilities of current MLLMs, we propose a simple multi-round cropped-search workflow. This strategy is shown to effectively improve model performance in realistic visual retrieval scenarios. Overall, our results provide practical guidance for the design of future multimodal deep-research systems. The code will be released in https://github.com/Osilly/Vision-DeepResearch.
中文摘要
多模态大型语言模型 (MLLM) 具有先进的 VQA,现在支持使用搜索引擎进行复杂的视觉文本事实调查的 Vision-DeepResearch 系统。然而,评估这些视觉和文本搜索能力仍然很困难,并且现有的基准有两个主要限制。首先,现有的基准不是以视觉搜索为中心的:需要视觉搜索的答案通常是通过文本问题中的跨文本线索泄露的,或者可以从当前 MLLM 中的先验知识中推断出来。其次,过于理想化的评估场景:在图像搜索方面,往往可以通过与全图近乎精确的匹配来获得所需的信息,而在文本搜索方面则过于直接,缺乏挑战性。为了解决这些问题,我们构建了由 2,000 个 VQA 实例组成的 Vision-DeepResearch 基准测试 (VDR-Bench)。所有问题都是通过仔细、多阶段的管理流程和严格的专家审查创建的,旨在评估 Vision-DeepResearch 系统在现实条件下的行为。此外,为了解决当前 MLLM 视觉检索能力不足的问题,我们提出了一种简单的多轮裁剪搜索工作流程。该策略被证明可以有效提高现实视觉检索场景中的模型性能。总的来说,我们的结果为未来多模态深度研究系统的设计提供了实用指导。代码将在 https://github.com/Osilly/Vision-DeepResearch 发布。
Golden Goose:从不可验证的互联网文本合成无限 RLVR 任务的简单技巧
- 标题: Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
- 作者: Ximing Lu, David Acuna, Jaehun Jung, Jian Hu, Di Zhang, Shizhe Diao, Yunheng Zou, Shaokun Zhang, Brandon Cui, Mingjie Liu, Hyunwoo Kim, Prithviraj Ammanabrolu, Jan Kautz, Yi Dong, Yejin Choi
- 日期: 2026-01-30
- ArXiv主页: https://arxiv.org/abs/2601.22975
- 论文链接: https://arxiv.org/pdf/2601.22975
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) has become a cornerstone for unlocking complex reasoning in Large Language Models (LLMs). Yet, scaling up RL is bottlenecked by limited existing verifiable data, where improvements increasingly saturate over prolonged training. To overcome this, we propose Golden Goose, a simple trick to synthesize unlimited RLVR tasks from unverifiable internet text by constructing a multiple-choice question-answering version of the fill-in-the-middle task. Given a source text, we prompt an LLM to identify and mask key reasoning steps, then generate a set of diverse, plausible distractors. This enables us to leverage reasoning-rich unverifiable corpora typically excluded from prior RLVR data construction (e.g., science textbooks) to synthesize GooseReason-0.7M, a large-scale RLVR dataset with over 0.7 million tasks spanning mathematics, programming, and general scientific domains. Empirically, GooseReason effectively revives models saturated on existing RLVR data, yielding robust, sustained gains under continuous RL and achieving new state-of-the-art results for 1.5B and 4B-Instruct models across 15 diverse benchmarks. Finally, we deploy Golden Goose in a real-world setting, synthesizing RLVR tasks from raw FineWeb scrapes for the cybersecurity domain, where no prior RLVR data exists. Training Qwen3-4B-Instruct on the resulting data GooseReason-Cyber sets a new state-of-the-art in cybersecurity, surpassing a 7B domain-specialized model with extensive domain-specific pre-training and post-training. This highlights the potential of automatically scaling up RLVR data by exploiting abundant, reasoning-rich, unverifiable internet text.
中文摘要
具有可验证奖励的强化学习 (RLVR) 已成为解锁大型语言模型 (LLM) 中复杂推理的基石。然而,扩大强化学习受到现有可验证数据有限的瓶颈,随着长时间的训练,改进会逐渐饱和。为了克服这个问题,我们提出了 Golden Goose,这是一种简单的技巧,通过构建填空任务的多项选择问答版本,从无法验证的互联网文本中合成无限的 RLVR 任务。给定源文本,我们提示法学硕士识别并掩盖关键推理步骤,然后生成一组多样化的、合理的干扰因素。这使我们能够利用通常被排除在先前 RLVR 数据构建(例如科学教科书)之外的推理丰富且无法验证的语料库来合成 GooseReason-0.7M,这是一个大规模 RLVR 数据集,包含超过 70 万个任务,涵盖数学、编程和一般科学领域。根据经验,GooseReason 有效地恢复了现有 RLVR 数据饱和的模型,在连续 RL 下产生强劲、持续的收益,并在 15 个不同的基准测试中为 1.5B 和 4B-Instruct 模型取得了新的最先进的结果。最后,我们在现实环境中部署 Golden Goose,从网络安全领域的原始 FineWeb 抓取中合成 RLVR 任务,其中不存在先前的 RLVR 数据。使用生成的数据训练 Qwen3-4B-Instruct GooseReason-Cyber 在网络安全领域树立了新的最先进水平,超越了 7B 领域专用模型,具有广泛的特定领域预训练和后训练。这凸显了通过利用丰富、推理丰富、无法验证的互联网文本来自动扩展 RLVR 数据的潜力。
WideSeek-R1:通过多智能体强化学习探索宽度缩放以实现广泛的信息搜索
- 标题: WideSeek-R1: Exploring Width Scaling for Broad Information Seeking via Multi-Agent Reinforcement Learning
- 作者: Zelai Xu, Zhexuan Xu, Ruize Zhang, Chunyang Zhu, Shi Yu, Weilin Liu, Quanlu Zhang, Wenbo Ding, Chao Yu, Yu Wang
- 日期: 2026-02-04
- ArXiv主页: https://arxiv.org/abs/2602.04634
- 论文链接: https://arxiv.org/pdf/2602.04634
- 项目链接: https://wideseek-r1.github.io/
- gitHub仓库: https://github.com/RLinf/RLinf
英文摘要
Recent advancements in Large Language Models (LLMs) have largely focused on depth scaling, where a single agent solves long-horizon problems with multi-turn reasoning and tool use. However, as tasks grow broader, the key bottleneck shifts from individual competence to organizational capability. In this work, we explore a complementary dimension of width scaling with multi-agent systems to address broad information seeking. Existing multi-agent systems often rely on hand-crafted workflows and turn-taking interactions that fail to parallelize work effectively. To bridge this gap, we propose WideSeek-R1, a lead-agent-subagent framework trained via multi-agent reinforcement learning (MARL) to synergize scalable orchestration and parallel execution. By utilizing a shared LLM with isolated contexts and specialized tools, WideSeek-R1 jointly optimizes the lead agent and parallel subagents on a curated dataset of 20k broad information-seeking tasks. Extensive experiments show that WideSeek-R1-4B achieves an item F1 score of 40.0% on the WideSearch benchmark, which is comparable to the performance of single-agent DeepSeek-R1-671B. Furthermore, WideSeek-R1-4B exhibits consistent performance gains as the number of parallel subagents increases, highlighting the effectiveness of width scaling.
中文摘要
大型语言模型 (LLM) 的最新进展主要集中在深度扩展上,其中单个代理通过多轮推理和工具使用来解决长期问题。然而,随着任务变得越来越广泛,关键瓶颈从个人能力转向组织能力。在这项工作中,我们探索了多智能体系统宽度缩放的补充维度,以解决广泛的信息搜索问题。现有的多代理系统通常依赖于手工设计的工作流程和轮流交互,无法有效地并行工作。为了弥补这一差距,我们提出了 WideSeek-R1,这是一种通过多代理强化学习 (MARL) 进行训练的主导-代理-子代理框架,以协同可扩展的编排和并行执行。通过利用具有隔离上下文和专用工具的共享 LLM,WideSeek-R1 在 20k 广泛信息搜索任务的精选数据集上联合优化主导代理和并行子代理。大量实验表明,WideSeek-R1-4B 在 WideSearch 基准上实现了 40.0% 的项目 F1 分数,与单代理 DeepSeek-R1-671B 的性能相当。此外,随着并行子代理数量的增加,WideSeek-R1-4B 表现出一致的性能增益,突出了宽度缩放的有效性。
CodeOCR:论视觉语言模型在代码理解中的有效性
-
标题: CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding
-
作者: Yuling Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen, Chenxu Zhang, Longfei Yun, Chengcheng Wan, Hongyu Zhang, David Lo, Xiaodong Gu
-
日期: 2026-02-02
-
ArXiv主页: https://arxiv.org/abs/2602.01785
-
gitHub仓库: https://github.com/YerbaPage/CodeOCR
英文摘要
Large Language Models (LLMs) have achieved remarkable success in source code understanding, yet as software systems grow in scale, computational efficiency has become a critical bottleneck. Currently, these models rely on a text-based paradigm that treats source code as a linear sequence of tokens, which leads to a linear increase in context length and associated computational costs. The rapid advancement of Multimodal LLMs (MLLMs) introduces an opportunity to optimize efficiency by representing source code as rendered images. Unlike text, which is difficult to compress without losing semantic meaning, the image modality is inherently suitable for compression. By adjusting resolution, images can be scaled to a fraction of their original token cost while remaining recognizable to vision-capable models. To explore the feasibility of this approach, we conduct the first systematic study on the effectiveness of MLLMs for code understanding. Our experiments reveal that: (1) MLLMs can effectively understand code with substantial token reduction, achieving up to 8x compression; (2) MLLMs can effectively leverage visual cues such as syntax highlighting, improving code completion performance under 4x compression; and (3) Code-understanding tasks like clone detection exhibit exceptional resilience to visual compression, with some compression ratios even slightly outperforming raw text inputs. Our findings highlight both the potential and current limitations of MLLMs in code understanding, which points out a shift toward image-modality code representation as a pathway to more efficient inference.
中文摘要
大型语言模型(LLM)在源代码理解方面取得了显着的成功,但随着软件系统规模的增长,计算效率已成为关键瓶颈。目前,这些模型依赖于基于文本的范例,将源代码视为令牌的线性序列,这导致上下文长度和相关计算成本线性增加。多模态 LLM (MLLM) 的快速发展带来了通过将源代码表示为渲染图像来优化效率的机会。与难以在不丢失语义的情况下压缩的文本不同,图像模态本质上适合压缩。通过调整分辨率,图像可以缩放到原始代币成本的一小部分,同时保持视觉模型的可识别性。为了探索这种方法的可行性,我们对 MLLM 在代码理解方面的有效性进行了首次系统研究。我们的实验表明:(1)MLLM 可以有效地理解代码,并大幅减少标记,实现高达 8 倍的压缩;(2) MLLM 可以有效利用语法突出显示等视觉提示,提高 4 倍压缩下的代码完成性能;(3) 克隆检测等代码理解任务对视觉压缩表现出卓越的弹性,某些压缩率甚至略优于原始文本输入。我们的研究结果强调了 MLLM 在代码理解方面的潜在和当前局限性,这指出了向图像模态代码表示的转变作为更有效推理的途径。
AOrchestra:自动创建子代理以进行代理编排
-
标题: AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
-
作者: Jianhao Ruan, Zhihao Xu, Yiran Peng, Fashen Ren, Zhaoyang Yu, Xinbing Liang, Jinyu Xiang, Bang Liu, Chenglin Wu, Yuyu Luo, Jiayi Zhang
-
日期: 2026-02-03
-
ArXiv主页: https://arxiv.org/abs/2602.03786
英文摘要
Language agents have shown strong promise for task automation. Realizing this promise for increasingly complex, long-horizon tasks has driven the rise of a sub-agent-as-tools paradigm for multi-turn task solving. However, existing designs still lack a dynamic abstraction view of sub-agents, thereby hurting adaptability. We address this challenge with a unified, framework-agnostic agent abstraction that models any agent as a tuple Instruction, Context, Tools, Model. This tuple acts as a compositional recipe for capabilities, enabling the system to spawn specialized executors for each task on demand. Building on this abstraction, we introduce an agentic system AOrchestra, where the central orchestrator concretizes the tuple at each step: it curates task-relevant context, selects tools and models, and delegates execution via on-the-fly automatic agent creation. Such designs enable reducing human engineering efforts, and remain framework-agnostic with plug-and-play support for diverse agents as task executors. It also enables a controllable performance-cost trade-off, allowing the system to approach Pareto-efficient. Across three challenging benchmarks (GAIA, SWE-Bench, Terminal-Bench), AOrchestra achieves 16.28% relative improvement against the strongest baseline when paired with Gemini-3-Flash. The code is available at: https://github.com/FoundationAgents/AOrchestra
中文摘要
语言代理在任务自动化方面表现出了强大的前景。实现日益复杂、长期任务的这一承诺推动了用于多轮任务解决的子代理即工具范式的兴起。然而,现有的设计仍然缺乏子代理的动态抽象视图,从而损害了适应性。我们通过统一的、与框架无关的代理抽象来应对这一挑战,该抽象将任何代理建模为元组指令、上下文、工具、模型。该元组充当功能的组合配方,使系统能够根据需要为每个任务生成专门的执行器。在此抽象的基础上,我们引入了一个代理系统 AOrchestra,其中中央编排器在每个步骤中具体化元组:它管理与任务相关的上下文,选择工具和模型,并通过动态自动代理创建来委托执行。此类设计可以减少人类工程工作,并保持与框架无关的状态,为作为任务执行者的不同代理提供即插即用支持。它还实现了可控的性能成本权衡,使系统接近帕累托效率。在三个具有挑战性的基准测试(GAIA、SWE-Bench、Terminal-Bench)中,与 Gemini-3-Flash 配合使用时,AOrchestra 相对于最强基准实现了 16.28% 的相对改进。代码位于:https://github.com/FoundationAgents/AOrchestra
CAR-bench:评估现实世界不确定性下 LLM 代理的一致性和限制意识
- 标题: CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
- 作者: Johannes Kirmayr, Lukas Stappen, Elisabeth André
- 日期: 2026-01-29
- ArXiv主页: https://arxiv.org/abs/2601.22027
- 论文链接: https://arxiv.org/pdf/2601.22027
- 项目链接: https://car-bench.github.io/car-bench/
- gitHub仓库: https://github.com/CAR-bench/car-bench
英文摘要
Existing benchmarks for Large Language Model (LLM) agents focus on task completion under idealistic settings but overlook reliability in real-world, user-facing applications. In domains, such as in-car voice assistants, users often issue incomplete or ambiguous requests, creating intrinsic uncertainty that agents must manage through dialogue, tool use, and policy adherence. We introduce CAR-bench, a benchmark for evaluating consistency, uncertainty handling, and capability awareness in multi-turn, tool-using LLM agents in an in-car assistant domain. The environment features an LLM-simulated user, domain policies, and 58 interconnected tools spanning navigation, productivity, charging, and vehicle control. Beyond standard task completion, CAR-bench introduces Hallucination tasks that test agents’ limit-awareness under missing tools or information, and Disambiguation tasks that require resolving uncertainty through clarification or internal information gathering. Baseline results reveal large gaps between occasional and consistent success on all task types. Even frontier reasoning LLMs achieve less than 50% consistent pass rate on Disambiguation tasks due to premature actions, and frequently violate policies or fabricate information to satisfy user requests in Hallucination tasks, underscoring the need for more reliable and self-aware LLM agents in real-world settings.
中文摘要
大型语言模型 (LLM) 代理的现有基准侧重于理想设置下的任务完成情况,但忽视了现实世界中面向用户的应用程序的可靠性。在车载语音助手等领域,用户经常发出不完整或模糊的请求,从而产生了代理必须通过对话、工具使用和政策遵守来管理的内在不确定性。我们推出了 CAR-bench,这是一个用于评估车载助理领域中多轮、使用工具的 LLM 代理的一致性、不确定性处理和能力意识的基准。该环境具有法学硕士模拟用户、域策略和 58 个涵盖导航、生产力、充电和车辆控制的互连工具。除了标准任务完成之外,CAR-bench 还引入了幻觉任务,用于测试代理在缺少工具或信息的情况下的限制意识,以及消歧任务,需要通过澄清或内部信息收集来解决不确定性。基线结果显示,所有任务类型的偶尔成功和持续成功之间存在巨大差距。即使是前沿推理法学硕士,由于不成熟的行动,在消歧任务上的一致通过率也低于 50%,并且在幻觉任务中经常违反政策或捏造信息以满足用户请求,这凸显了在现实环境中需要更可靠和具有自我意识的法学硕士代理。
闭环:使用 RPG 编码器表示通用存储库
- 标题: Closing the Loop: Universal Repository Representation with RPG-Encoder
- 作者: Jane Luo, Chengyu Yin, Xin Zhang, Qingtao Li, Steven Liu, Yiming Huang, Jie Wu, Hao Liu, Yangyu Huang, Yu Kang, Fangkai Yang, Ying Xin, Scarlett Li
- 日期: 2026-02-02
- ArXiv主页: https://arxiv.org/abs/2602.02084
- 论文链接: https://arxiv.org/pdf/2602.02084
- 项目链接: https://ayanami2003.github.io/RPG-Encoder/
- gitHub仓库: https://github.com/microsoft/RPG-ZeroRepo
英文摘要
Current repository agents encounter a reasoning disconnect due to fragmented representations, as existing methods rely on isolated API documentation or dependency graphs that lack semantic depth. We consider repository comprehension and generation to be inverse processes within a unified cycle: generation expands intent into implementation, while comprehension compresses implementation back into intent. To address this, we propose RPG-Encoder, a framework that generalizes the Repository Planning Graph (RPG) from a static generative blueprint into a unified, high-fidelity representation. RPG-Encoder closes the reasoning loop through three mechanisms: (1) Encoding raw code into the RPG that combines lifted semantic features with code dependencies; (2) Evolving the topology incrementally to decouple maintenance costs from repository scale, reducing overhead by 95.7%; and (3) Operating as a unified interface for structure-aware navigation. In evaluations, RPG-Encoder establishes state-of-the-art repository understanding on SWE-bench Verified with 93.7% Acc@5 and exceeds the best baseline by over 10% on SWE-bench Live Lite. These results highlight our superior fine-grained localization accuracy in complex codebases. Furthermore, it achieves 98.5% reconstruction coverage on RepoCraft, confirming RPG’s high-fidelity capacity to mirror the original codebase and closing the loop between intent and implementation.
中文摘要
当前的存储库代理由于碎片化的表示而遇到推理脱节,因为现有方法依赖于孤立的 API 文档或缺乏语义深度的依赖关系图。我们认为存储库理解和生成是统一循环内的逆过程:生成将意图扩展为实现,而理解将实现压缩回意图。为了解决这个问题,我们提出了 RPG-Encoder,这是一个将存储库规划图(RPG)从静态生成蓝图概括为统一的高保真表示的框架。RPG-Encoder 通过三种机制关闭推理循环:(1)将原始代码编码到 RPG 中,将提升的语义特征与代码依赖关系相结合;(2) 逐步改进拓扑,将维护成本与存储库规模解耦,开销降低 95.7%;(3) 作为结构感知导航的统一界面进行操作。在评估中,RPG-Encoder 在 SWE-bench Verified 上以 93.7% Acc@5 建立了最先进的存储库理解,并在 SWE-bench Live Lite 上超出了最佳基线 10% 以上。这些结果凸显了我们在复杂代码库中卓越的细粒度定位准确性。此外,它在 RepoCraft 上实现了 98.5% 的重建覆盖率,证实了 RPG 镜像原始代码库的高保真能力,并关闭了意图和实现之间的循环。
DFlash:用于闪存推测解码的块扩散
- 标题: DFlash: Block Diffusion for Flash Speculative Decoding
- 作者: Jian Chen, Yesheng Liang, Zhijian Liu
- 日期: 2026-02-05
- ArXiv主页: https://arxiv.org/abs/2602.06036
- 论文链接: https://arxiv.org/pdf/2602.06036
- 项目链接: https://z-lab.ai/projects/dflash/
- gitHub仓库: https://github.com/z-lab/dflash
英文摘要
Autoregressive large language models (LLMs) deliver strong performance but require inherently sequential decoding, leading to high inference latency and poor GPU utilization. Speculative decoding mitigates this bottleneck by using a fast draft model whose outputs are verified in parallel by the target LLM; however, existing methods still rely on autoregressive drafting, which remains sequential and limits practical speedups. Diffusion LLMs offer a promising alternative by enabling parallel generation, but current diffusion models typically underperform compared with autoregressive models. In this paper, we introduce DFlash, a speculative decoding framework that employs a lightweight block diffusion model for parallel drafting. By generating draft tokens in a single forward pass and conditioning the draft model on context features extracted from the target model, DFlash enables efficient drafting with high-quality outputs and higher acceptance rates. Experiments show that DFlash achieves over 6x lossless acceleration across a range of models and tasks, delivering up to 2.5x higher speedup than the state-of-the-art speculative decoding method EAGLE-3.
中文摘要
自回归大型语言模型 (LLM) 可提供强大的性能,但需要固有的顺序解码,从而导致推理延迟高和 GPU 利用率低。推测性解码通过使用快速草稿模型来缓解这一瓶颈,该模型的输出由目标 LLM 并行验证;然而,现有的方法仍然依赖于自回归绘图,这仍然是连续的并限制了实际的加速。扩散法学硕士通过实现并行生成提供了一种有前途的替代方案,但当前的扩散模型通常比自回归模型表现不佳。在本文中,我们介绍了 DFlash,这是一种推测性解码框架,它采用轻量级块扩散模型进行并行绘图。通过在单次前向传递中生成草稿令牌,并根据从目标模型中提取的上下文特征来调节草稿模型,DFlash 能够实现高效的草稿,并提供高质量的输出和更高的接受率。实验表明,DFlash 在一系列模型和任务中实现了超过 6 倍的无损加速,比最先进的推测解码方法 EAGLE-3 提供高达 2.5 倍的加速。
多模式过程奖励模型中的训练数据效率
-
标题: Training Data Efficiency in Multimodal Process Reward Models
-
作者: Jinyuan Li, Chengsong Huang, Langlin Huang, Shaoyang Xu, Haolin Liu, Wenxuan Zhang, Jiaxin Huang
-
日期: 2026-02-04
-
ArXiv主页: https://arxiv.org/abs/2602.04145
-
gitHub仓库: https://github.com/JinYuanLi0012/Balanced-Info-MPRM
英文摘要
Multimodal Process Reward Models (MPRMs) are central to step-level supervision for visual reasoning in MLLMs. Training MPRMs typically requires large-scale Monte Carlo (MC)-annotated corpora, incurring substantial training cost. This paper studies the data efficiency for MPRM training.Our preliminary experiments reveal that MPRM training quickly saturates under random subsampling of the training data, indicating substantial redundancy within existing MC-annotated corpora.To explain this, we formalize a theoretical framework and reveal that informative gradient updates depend on two factors: label mixtures of positive/negative steps and label reliability (average MC scores of positive steps). Guided by these insights, we propose the Balanced-Information Score (BIS), which prioritizes both mixture and reliability based on existing MC signals at the rollout level, without incurring any additional cost. Across two backbones (InternVL2.5-8B and Qwen2.5-VL-7B) on VisualProcessBench, BIS-selected subsets consistently match and even surpass the full-data performance at small fractions. Notably, the BIS subset reaches full-data performance using only 10% of the training data, improving over random subsampling by a relative 4.1%.
中文摘要
多模态过程奖励模型 (MPRM) 是 MLLM 中视觉推理的步骤级监督的核心。训练 MPRM 通常需要大规模的蒙特卡罗 (MC) 注释语料库,从而产生大量的训练成本。本文研究了 MPRM 训练的数据效率。我们的初步实验表明,MPRM 训练在训练数据的随机子采样下很快饱和,这表明现有 MC 注释语料库中存在大量冗余。为了解释这一点,我们形式化了一个理论框架,并揭示了信息梯度更新取决于两个因素:正/负步骤的标签混合和标签可靠性(正步骤的平均 MC 分数)。在这些见解的指导下,我们提出了平衡信息评分 (BIS),它根据部署级别的现有 MC 信号优先考虑混合和可靠性,而不会产生任何额外成本。在 VisualProcessBench 上的两个主干网(InternVL2.5-8B 和 Qwen2.5-VL-7B)中,BIS 选择的子集始终匹配甚至超越了小部分的全数据性能。值得注意的是,BIS 子集仅使用 10% 的训练数据即可达到全数据性能,相对随机二次采样提高了 4.1%。
UniReason 1.0:世界知识对齐图像生成和编辑的统一推理框架
-
标题: UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
-
作者: Dianyi Wang, Chaofan Ma, Feng Han, Size Wu, Wei Song, Yibin Wang, Zhixiong Zhang, Tianhang Wang, Siyuan Wang, Zhongyu Wei, Jiaqi Wang
-
日期: 2026-02-02
-
ArXiv主页: https://arxiv.org/abs/2602.02437
英文摘要
Unified multimodal models often struggle with complex synthesis tasks that demand deep reasoning, and typically treat text-to-image generation and image editing as isolated capabilities rather than interconnected reasoning steps. To address this, we propose UniReason, a unified framework that harmonizes these two tasks through a dual reasoning paradigm. We formulate generation as world knowledge-enhanced planning to inject implicit constraints, and leverage editing capabilities for fine-grained visual refinement to further correct visual errors via self-reflection. This approach unifies generation and editing within a shared representation, mirroring the human cognitive process of planning followed by refinement. We support this framework by systematically constructing a large-scale reasoning-centric dataset (~300k samples) covering five major knowledge domains (e.g., cultural commonsense, physics, etc.) for planning, alongside an agent-generated corpus for visual self-correction. Extensive experiments demonstrate that UniReason achieves advanced performance on reasoning-intensive benchmarks such as WISE, KrisBench and UniREditBench, while maintaining superior general synthesis capabilities.
中文摘要
统一的多模态模型通常难以应对需要深度推理的复杂合成任务,并且通常将文本到图像的生成和图像编辑视为独立的功能,而不是互连的推理步骤。为了解决这个问题,我们提出了 UniReason,一个统一的框架,通过双重推理范式协调这两项任务。我们将生成制定为世界知识增强规划,以注入隐式约束,并利用编辑功能进行细粒度视觉细化,以通过自我反思进一步纠正视觉错误。这种方法将生成和编辑统一在一个共享的表示中,反映了人类规划和细化的认知过程。我们通过系统地构建一个涵盖五个主要知识领域(例如文化常识、物理等)的大规模以推理为中心的数据集(约 30 万个样本)来支持该框架,并构建一个用于视觉自我校正的代理生成的语料库。大量实验表明,UniReason 在 WISE、KrisBench 和 UniREditBench 等推理密集型基准测试中实现了先进的性能,同时保持了卓越的通用综合能力。
思想链中没有全局计划:揭示法学硕士的潜在规划视野
-
标题: No Global Plan in Chain-of-Thought: Uncover the Latent Planning Horizon of LLMs
-
作者: Liyan Xu, Mo Yu, Fandong Meng, Jie Zhou
-
日期: 2026-02-02
-
ArXiv主页: https://arxiv.org/abs/2602.02103
-
gitHub仓库: https://github.com/lxucs/tele-lens
英文摘要
This work stems from prior complementary observations on the dynamics of Chain-of-Thought (CoT): Large Language Models (LLMs) is shown latent planning of subsequent reasoning prior to CoT emergence, thereby diminishing the significance of explicit CoT; whereas CoT remains critical for tasks requiring multi-step reasoning. To deepen the understanding between LLM’s internal states and its verbalized reasoning trajectories, we investigate the latent planning strength of LLMs, through our probing method, Tele-Lens, applying to hidden states across diverse task domains. Our empirical results indicate that LLMs exhibit a myopic horizon, primarily conducting incremental transitions without precise global planning. Leveraging this characteristic, we propose a hypothesis on enhancing uncertainty estimation of CoT, which we validate that a small subset of CoT positions can effectively represent the uncertainty of the entire path. We further underscore the significance of exploiting CoT dynamics, and demonstrate that automatic recognition of CoT bypass can be achieved without performance degradation. Our code, data and models are released at https://github.com/lxucs/tele-lens.
中文摘要
这项工作源于先前对思想链 (CoT) 动态的补充观察:大型语言模型 (LLM) 在 CoT 出现之前就显示出后续推理的潜在规划,从而削弱了显式 CoT 的重要性;而 CoT 对于需要多步骤推理的任务仍然至关重要。为了加深对 LLM 内部状态及其语言推理轨迹之间的理解,我们通过我们的探测方法 Tele-Lens 来研究 LLM 的潜在规划强度,并将其应用于不同任务领域的隐藏状态。我们的实证结果表明,法学硕士表现出短视的视野,主要是在没有精确的全球规划的情况下进行增量过渡。利用这一特性,我们提出了增强 CoT 不确定性估计的假设,并验证了 CoT 位置的一个小子集可以有效地表示整个路径的不确定性。我们进一步强调了利用 CoT 动态的重要性,并证明可以在不降低性能的情况下实现 CoT 旁路的自动识别。我们的代码、数据和模型发布在 https://github.com/lxucs/tele-lens。
Spider-Sense:通过分层自适应筛选进行有效代理防御的内在风险感知
-
标题: Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening
-
作者: Zhenxiong Yu, Zhi Yang, Zhiheng Jin, Shuhe Wang, Heng Zhang, Yanlin Fei, Lingfeng Zeng, Fangqi Lou, Shuo Zhang, Tu Hu, Jingping Liu, Rongze Chen, Xingyu Zhu, Kunyi Wang, Chaofa Yuan, Xin Guo, Zhaowei Liu, Feipeng Zhang, Jie Huang, Huacan Wang, Ronghao Chen, Liwen Zhang
-
日期: 2026-02-05
-
ArXiv主页: https://arxiv.org/abs/2602.05386
-
gitHub仓库: https://github.com/aifinlab/Spider-Sense
英文摘要
As large language models (LLMs) evolve into autonomous agents, their real-world applicability has expanded significantly, accompanied by new security challenges. Most existing agent defense mechanisms adopt a mandatory checking paradigm, in which security validation is forcibly triggered at predefined stages of the agent lifecycle. In this work, we argue that effective agent security should be intrinsic and selective rather than architecturally decoupled and mandatory. We propose Spider-Sense framework, an event-driven defense framework based on Intrinsic Risk Sensing (IRS), which allows agents to maintain latent vigilance and trigger defenses only upon risk perception. Once triggered, the Spider-Sense invokes a hierarchical defence mechanism that trades off efficiency and precision: it resolves known patterns via lightweight similarity matching while escalating ambiguous cases to deep internal reasoning, thereby eliminating reliance on external models. To facilitate rigorous evaluation, we introduce S^2Bench, a lifecycle-aware benchmark featuring realistic tool execution and multi-stage attacks. Extensive experiments demonstrate that Spider-Sense achieves competitive or superior defense performance, attaining the lowest Attack Success Rate (ASR) and False Positive Rate (FPR), with only a marginal latency overhead of 8.3%.
中文摘要
随着大型语言模型(LLM)演变成自主代理,它们的现实世界适用性显着扩展,并伴随着新的安全挑战。现有的代理防御机制大多采用强制检查范式,在代理生命周期的预定义阶段强制触发安全验证。在这项工作中,我们认为有效的代理安全应该是内在的和选择性的,而不是在架构上解耦和强制性的。我们提出了Spider-Sense框架,这是一种基于内在风险感知(IRS)的事件驱动的防御框架,它允许代理保持潜在的警惕性并仅在风险感知时触发防御。一旦触发,Spider-Sense 就会调用一种权衡效率和精度的分层防御机制:它通过轻量级相似性匹配来解决已知模式,同时将模棱两可的案例升级为深层内部推理,从而消除对外部模型的依赖。为了促进严格的评估,我们引入了 S^2Bench,这是一种生命周期感知的基准测试,具有真实的工具执行和多阶段攻击的特点。大量实验表明,Spider-Sense 实现了具有竞争力或卓越的防御性能,实现了最低的攻击成功率 (ASR) 和误报率 (FPR),且边际延迟开销仅为 8.3%。
MARS:用于自动化人工智能研究的具有反射搜索的模块化代理
-
标题: MARS: Modular Agent with Reflective Search for Automated AI Research
-
作者: Jiefeng Chen, Bhavana Dalvi Mishra, Jaehyun Nam, Rui Meng, Tomas Pfister, Jinsung Yoon
-
日期: 2026-02-02
-
ArXiv主页: https://arxiv.org/abs/2602.02660
-
gitHub仓库: https://github.com/jfc43/MARS
英文摘要
Automating AI research differs from general software engineering due to computationally expensive evaluation (e.g., model training) and opaque performance attribution. Current LLM-based agents struggle here, often generating monolithic scripts that ignore execution costs and causal factors. We introduce MARS (Modular Agent with Reflective Search), a framework optimized for autonomous AI research. MARS relies on three pillars: (1) Budget-Aware Planning via cost-constrained Monte Carlo Tree Search (MCTS) to explicitly balance performance with execution expense; (2) Modular Construction, employing a “Design-Decompose-Implement” pipeline to manage complex research repositories; and (3) Comparative Reflective Memory, which addresses credit assignment by analyzing solution differences to distill high-signal insights. MARS achieves state-of-the-art performance among open-source frameworks on MLE-Bench under comparable settings, maintaining competitiveness with the global leaderboard’s top methods. Furthermore, the system exhibits qualitative “Aha!” moments, where 63% of all utilized lessons originate from cross-branch transfer, demonstrating that the agent effectively generalizes insights across search paths.
中文摘要
由于计算成本高昂的评估(例如模型训练)和不透明的性能归因,自动化人工智能研究与一般软件工程不同。当前基于 LLM 的代理在这方面举步维艰,通常会生成忽略执行成本和因果因素的整体脚本。我们推出 MARS(带有反射搜索的模块化代理),这是一个针对自主人工智能研究优化的框架。MARS 依赖于三个支柱:(1)通过成本受限的蒙特卡罗树搜索(MCTS)进行预算感知规划,以明确平衡性能与执行费用;(2)模块化构建,采用“设计-分解-实施”流程来管理复杂的研究存储库;(3) 比较反思记忆,通过分析解决方案差异来提取高信号洞察来解决学分分配问题。MARS在MLE-Bench上同类设置下的开源框架中实现了最先进的性能,保持了与全球排行榜顶级方法的竞争力。此外,该系统表现出定性的“啊哈!”时刻,所有使用的课程中有 63% 来自跨部门转移,这表明代理有效地概括了跨搜索路径的见解。
用于生成视图自适应人类视频的 3D 感知隐式运动控制
- 标题: 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
- 作者: Zhixue Fang, Xu He, Songlin Tang, Haoxian Zhang, Qingfeng Li, Xiaoqiang Liu, Pengfei Wan, Kun Gai
- 日期: 2026-02-03
- ArXiv主页: https://arxiv.org/abs/2602.03796
- 论文链接: https://arxiv.org/pdf/2602.03796
- 项目链接: https://hjrphoebus.github.io/3DiMo/
英文摘要
Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding novel-view synthesis. Explicit 3D models, though structurally informative, suffer from inherent inaccuracies (e.g., depth ambiguity and inaccurate dynamics) which, when used as a strong constraint, override the powerful intrinsic 3D awareness of large-scale video generators. In this work, we revisit motion control from a 3D-aware perspective, advocating for an implicit, view-agnostic motion representation that naturally aligns with the generator’s spatial priors rather than depending on externally reconstructed constraints. We introduce 3DiMo, which jointly trains a motion encoder with a pretrained video generator to distill driving frames into compact, view-agnostic motion tokens, injected semantically via cross-attention. To foster 3D awareness, we train with view-rich supervision (i.e., single-view, multi-view, and moving-camera videos), forcing motion consistency across diverse viewpoints. Additionally, we use auxiliary geometric supervision that leverages SMPL only for early initialization and is annealed to zero, enabling the model to transition from external 3D guidance to learning genuine 3D spatial motion understanding from the data and the generator’s priors. Experiments confirm that 3DiMo faithfully reproduces driving motions with flexible, text-driven camera control, significantly surpassing existing methods in both motion fidelity and visual quality.
中文摘要
视频生成中现有的人体运动控制方法通常依赖 2D 姿势或显式 3D 参数模型(例如 SMPL)作为控制信号。然而,2D 将运动严格绑定到驾驶视点,从而妨碍了新颖的视图合成。显式 3D 模型虽然在结构上信息丰富,但存在固有的不准确性(例如,深度模糊性和不准确的动态),当用作强约束时,会覆盖大型视频生成器强大的内在 3D 感知能力。在这项工作中,我们从 3D 感知的角度重新审视运动控制,提倡一种隐式的、与视图无关的运动表示,它自然地与生成器的空间先验对齐,而不是依赖于外部重建的约束。我们引入了 3DiMo,它联合训练运动编码器和预先训练的视频生成器,将驱动帧提取为紧凑的、与视图无关的运动标记,并通过交叉注意力进行语义注入。为了培养 3D 意识,我们通过丰富视图的监督(即单视图、多视图和移动摄像机视频)进行训练,强制不同视点之间的运动一致性。此外,我们使用辅助几何监督,仅利用 SMPL 进行早期初始化,并退火至零,使模型能够从外部 3D 指导过渡到从数据和生成器先验中学习真正的 3D 空间运动理解。实验证实,3DiMo 通过灵活的文本驱动摄像头控制忠实地再现了驾驶动作,在运动保真度和视觉质量方面都显着超越了现有方法。
MemSkill:自我进化智能体的学习和进化记忆技能
- 标题: MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
- 作者: Haozhen Zhang, Quanyu Long, Jianzhu Bao, Tao Feng, Weizhi Zhang, Haodong Yue, Wenya Wang
- 日期: 2026-02-02
- ArXiv主页: https://arxiv.org/abs/2602.02474
- 论文链接: https://arxiv.org/pdf/2602.02474
- 项目链接: https://viktoraxelsen.github.io/MemSkill/
- gitHub仓库: https://github.com/ViktorAxelsen/MemSkill
英文摘要
Most Large Language Model (LLM) agent memory systems rely on a small set of static, hand-designed operations for extracting memory. These fixed procedures hard-code human priors about what to store and how to revise memory, making them rigid under diverse interaction patterns and inefficient on long histories. To this end, we present MemSkill, which reframes these operations as learnable and evolvable memory skills, structured and reusable routines for extracting, consolidating, and pruning information from interaction traces. Inspired by the design philosophy of agent skills, MemSkill employs a controller that learns to select a small set of relevant skills, paired with an LLM-based executor that produces skill-guided memories. Beyond learning skill selection, MemSkill introduces a designer that periodically reviews hard cases where selected skills yield incorrect or incomplete memories, and evolves the skill set by proposing refinements and new skills. Together, MemSkill forms a closed-loop procedure that improves both the skill-selection policy and the skill set itself. Experiments on LoCoMo, LongMemEval, HotpotQA, and ALFWorld demonstrate that MemSkill improves task performance over strong baselines and generalizes well across settings. Further analyses shed light on how skills evolve, offering insights toward more adaptive, self-evolving memory management for LLM agents.
中文摘要
大多数大型语言模型 (LLM) 代理内存系统依赖于一小组静态的、手工设计的操作来提取内存。这些固定的程序硬编码了人类关于存储什么以及如何修改记忆的先验知识,使得它们在不同的交互模式下变得僵化,并且在长期历史中效率低下。为此,我们提出了 MemSkill,它将这些操作重新定义为可学习和可进化的记忆技能、结构化和可重用的例程,用于从交互痕迹中提取、合并和修剪信息。受代理技能设计理念的启发,MemSkill 采用了一个控制器,可以学习选择一小组相关技能,并与一个基于 LLM 的执行器配合使用,可以生成技能引导的记忆。除了学习技能选择之外,MemSkill 还引入了一位设计师,该设计师会定期审查所选技能产生不正确或不完整记忆的困难案例,并通过提出改进和新技能来发展技能集。MemSkill 共同形成了一个闭环程序,可以改进技能选择策略和技能集本身。LoCoMo、LongMemEval、HotpotQA 和 ALFWorld 上的实验表明,MemSkill 在强基线上提高了任务性能,并且在不同设置中具有良好的泛化能力。进一步的分析揭示了技能如何演变,为法学硕士代理人的更具适应性、自我进化的记忆管理提供了见解。
SWE-Universe:将真实世界的可验证环境扩展到数百万个
- 标题: SWE-Universe: Scale Real-World Verifiable Environments to Millions
- 作者: Mouxiang Chen, Lei Zhang, Yunlong Feng, Xuwu Wang, Wenting Zhao, Ruisheng Cao, Jiaxi Yang, Jiawei Chen, Mingze Li, Zeyao Ma, Hao Ge, Zongmeng Zhang, Zeyu Cui, Dayiheng Liu, Jingren Zhou, Jianling Sun, Junyang Lin, Binyuan Hui
- 日期: 2026-02-02
- ArXiv主页: https://arxiv.org/abs/2602.02361
- 论文链接: https://arxiv.org/pdf/2602.02361
英文摘要
We propose SWE-Universe, a scalable and efficient framework for automatically constructing real-world software engineering (SWE) verifiable environments from GitHub pull requests (PRs). To overcome the prevalent challenges of automatic building, such as low production yield, weak verifiers, and prohibitive cost, our framework utilizes a building agent powered by an efficient custom-trained model. This agent employs iterative self-verification and in-loop hacking detection to ensure the reliable generation of high-fidelity, verifiable tasks. Using this method, we scale the number of real-world multilingual SWE environments to a million scale (807,693). We demonstrate the profound value of our environments through large-scale agentic mid-training and reinforcement learning. Finally, we applied this technique to Qwen3-Max-Thinking and achieved a score of 75.3% on SWE-Bench Verified. Our work provides both a critical resource and a robust methodology to advance the next generation of coding agents.
中文摘要
我们提出了 SWE-Universe,这是一个可扩展且高效的框架,用于从 GitHub 拉取请求 (PR) 自动构建真实世界的软件工程 (SWE) 可验证环境。为了克服自动构建的普遍挑战,例如低产量、弱验证器和高昂的成本,我们的框架采用了由高效的定制训练模型支持的构建代理。该代理采用迭代自我验证和内循环黑客检测来确保可靠生成高保真、可验证的任务。使用这种方法,我们将现实世界的多语言 SWE 环境的数量扩展到百万级 (807,693)。我们通过大规模代理中期训练和强化学习展示了我们环境的深刻价值。最后,我们将该技术应用于 Qwen3-Max-Thinking,并在 SWE-Bench Verified 上取得了 75.3% 的分数。我们的工作提供了关键资源和强大的方法来推进下一代编码代理。
四方二:通过改进的无偏梯度估计在 NVFP4 中进行准确的 LLM 预训练
-
标题: Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
-
作者: Andrei Panferov, Erik Schultheis, Soroush Tabesh, Dan Alistarh
-
日期: 2026-01-30
-
ArXiv主页: https://arxiv.org/abs/2601.22813
-
gitHub仓库: https://github.com/IST-DASLab/Quartet-II
英文摘要
The NVFP4 lower-precision format, supported in hardware by NVIDIA Blackwell GPUs, promises to allow, for the first time, end-to-end fully-quantized pre-training of massive models such as LLMs. Yet, existing quantized training methods still sacrifice some of the representation capacity of this format in favor of more accurate unbiased quantized gradient estimation by stochastic rounding (SR), losing noticeable accuracy relative to standard FP16 and FP8 training. In this paper, improve the state of the art for quantized training in NVFP4 via a novel unbiased quantization routine for micro-scaled formats, called MS-EDEN, that has more than 2x lower quantization error than SR. We integrate it into a novel fully-NVFP4 quantization scheme for linear layers, called Quartet II. We show analytically that Quartet II achieves consistently better gradient estimation across all major matrix multiplications, both on the forward and on the backward passes. In addition, our proposal synergizes well with recent training improvements aimed specifically at NVFP4. We further validate Quartet II on end-to-end LLM training with up to 1.9B parameters on 38B tokens. We provide kernels for execution on NVIDIA Blackwell GPUs with up to 4.2x speedup over BF16. Our code is available at https://github.com/IST-DASLab/Quartet-II .
中文摘要
NVFP4 低精度格式由 NVIDIA Blackwell GPU 硬件支持,有望首次实现 LLM 等大规模模型的端到端完全量化预训练。然而,现有的量化训练方法仍然牺牲了这种格式的一些表示能力,以支持通过随机舍入 (SR) 进行更准确的无偏量化梯度估计,相对于标准 FP16 和 FP8 训练,其准确性明显下降。在本文中,通过一种用于微尺度格式的新型无偏量化例程(称为 MS-EDEN)提高了 NVFP4 量化训练的技术水平,该例程的量化误差比 SR 低 2 倍以上。我们将其集成到一种新颖的线性层全 NVFP4 量化方案中,称为 Quartet II。我们通过分析表明,Quartet II 在所有主要矩阵乘法中(无论是前向传播还是后向传播)都始终实现了更好的梯度估计。此外,我们的建议与最近专门针对 NVFP4 的培训改进具有良好的协同作用。我们进一步在 38B 代币上使用高达 1.9B 参数的端到端 LLM 训练验证 Quartet II。我们提供在 NVIDIA Blackwell GPU 上执行的内核,其速度比 BF16 提升高达 4.2 倍。我们的代码可在 https://github.com/IST-DASLab/Quartet-II 获取。
ASTRA:代理轨迹和强化竞技场的自动合成
- 标题: ASTRA: Automated Synthesis of agentic Trajectories and Reinforcement Arenas
- 作者: Xiaoyu Tian, Haotian Wang, Shuaiting Chen, Hao Zhou, Kaichi Yu, Yudian Zhang, Jade Ouyang, Junxi Yin, Jiong Chen, Baoyan Guo, Lei Zhang, Junjie Tao, Yuansheng Song, Ming Cui, Chengwei Liu
- 日期: 2026-01-29
- ArXiv主页: https://arxiv.org/abs/2601.21558
- 论文链接: https://arxiv.org/pdf/2601.21558
- 项目链接: https://lianjiatech.github.io/astra.blog/
- gitHub仓库: https://github.com/LianjiaTech/astra
英文摘要
Large language models (LLMs) are increasingly used as tool-augmented agents for multi-step decision making, yet training robust tool-using agents remains challenging. Existing methods still require manual intervention, depend on non-verifiable simulated environments, rely exclusively on either supervised fine-tuning (SFT) or reinforcement learning (RL), and struggle with stable long-horizon, multi-turn learning. To address these challenges, we introduce ASTRA, a fully automated end-to-end framework for training tool-augmented language model agents via scalable data synthesis and verifiable reinforcement learning. ASTRA integrates two complementary components. First, a pipeline that leverages the static topology of tool-call graphs synthesizes diverse, structurally grounded trajectories, instilling broad and transferable tool-use competence. Second, an environment synthesis framework that captures the rich, compositional topology of human semantic reasoning converts decomposed question-answer traces into independent, code-executable, and rule-verifiable environments, enabling deterministic multi-turn RL. Based on this method, we develop a unified training methodology that integrates SFT with online RL using trajectory-level rewards to balance task completion and interaction efficiency. Experiments on multiple agentic tool-use benchmarks demonstrate that ASTRA-trained models achieve state-of-the-art performance at comparable scales, approaching closed-source systems while preserving core reasoning ability. We release the full pipelines, environments, and trained models at https://github.com/LianjiaTech/astra.
中文摘要
大型语言模型(LLM)越来越多地用作多步骤决策的工具增强代理,但训练强大的工具使用代理仍然具有挑战性。现有方法仍然需要人工干预,依赖于不可验证的模拟环境,完全依赖监督微调(SFT)或强化学习(RL),并且难以实现稳定的长视野、多轮学习。为了应对这些挑战,我们引入了 ASTRA,这是一个完全自动化的端到端框架,用于通过可扩展的数据合成和可验证的强化学习来训练工具增强的语言模型代理。ASTRA 集成了两个互补的组件。首先,利用工具调用图的静态拓扑的管道综合了多样化的、有结构基础的轨迹,灌输了广泛且可转移的工具使用能力。其次,环境综合框架捕获了人类语义推理的丰富的组合拓扑,将分解的问答轨迹转换为独立的、代码可执行的和规则可验证的环境,从而实现确定性多轮强化学习。基于这种方法,我们开发了一种统一的训练方法,将 SFT 与在线 RL 相结合,使用轨迹级奖励来平衡任务完成和交互效率。对多个代理工具使用基准的实验表明,ASTRA 训练的模型在可比较的规模上实现了最先进的性能,接近闭源系统,同时保留了核心推理能力。我们在 https://github.com/LianjiaTech/astra 发布了完整的管道、环境和训练模型。
长度无偏序列策略优化:揭示和控制 RLVR 中的响应长度变化
-
标题: Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
-
作者: Fanfan Liu, Youyang Yin, Peng Shi, Siqi Yang, Zhixiong Zeng, Haibo Qiu
-
日期: 2026-02-05
-
ArXiv主页: https://arxiv.org/abs/2602.05261
-
gitHub仓库: https://github.com/murphy4122/LUSPO
英文摘要
Recent applications of Reinforcement Learning with Verifiable Rewards (RLVR) to Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated significant success in enhancing reasoning capabilities for complex tasks. During RLVR training, an increase in response length is often regarded as a key factor contributing to the growth of reasoning ability. However, the patterns of change in response length vary significantly across different RLVR algorithms during the training process. To provide a fundamental explanation for these variations, this paper conducts an in-depth analysis of the components of mainstream RLVR algorithms. We present a theoretical analysis of the factors influencing response length and validate our theory through extensive experimentation. Building upon these theoretical findings, we propose the Length-Unbiased Sequence Policy Optimization (LUSPO) algorithm. Specifically, we rectify the length bias inherent in Group Sequence Policy Optimization (GSPO), rendering its loss function unbiased with respect to response length and thereby resolving the issue of response length collapse. We conduct extensive experiments across mathematical reasoning benchmarks and multimodal reasoning scenarios, where LUSPO consistently achieves superior performance. Empirical results demonstrate that LUSPO represents a novel, state-of-the-art optimization strategy compared to existing methods such as GRPO and GSPO.
中文摘要
最近,具有可验证奖励的强化学习 (RLVR) 在大型语言模型 (LLM) 和视觉语言模型 (VLM) 中的应用在增强复杂任务的推理能力方面取得了巨大成功。在RLVR训练过程中,反应长度的增加通常被认为是促进推理能力增长的关键因素。然而,在训练过程中,不同 RLVR 算法的响应长度变化模式存在显着差异。为了从根本上解释这些变化,本文对主流 RLVR 算法的组成部分进行了深入分析。我们对影响响应长度的因素进行了理论分析,并通过广泛的实验验证了我们的理论。基于这些理论发现,我们提出了长度无偏序列策略优化(LUSPO)算法。具体来说,我们纠正了组序列策略优化(GSPO)中固有的长度偏差,使其损失函数相对于响应长度无偏差,从而解决了响应长度崩溃的问题。我们在数学推理基准和多模态推理场景中进行了广泛的实验,其中 LUSPO 始终实现卓越的性能。实证结果表明,与 GRPO 和 GSPO 等现有方法相比,LUSPO 代表了一种新颖、最先进的优化策略。
daVinci-Agency:有效解锁长期机构数据
-
标题: daVinci-Agency: Unlocking Long-Horizon Agency Data-Efficiently
-
作者: Mohan Jiang, Dayuan Fu, Junhao Shi, Ji Zeng, Weiye Si, Keyu Li, Xuefeng Li, Yang Xiao, Wenjie Li, Dequan Wang, Pengfei Liu
-
日期: 2026-02-02
-
ArXiv主页: https://arxiv.org/abs/2602.02619
英文摘要
While Large Language Models (LLMs) excel at short-term tasks, scaling them to long-horizon agentic workflows remains challenging. The core bottleneck lies in the scarcity of training data that captures authentic long-dependency structures and cross-stage evolutionary dynamics–existing synthesis methods either confine to single-feature scenarios constrained by model distribution, or incur prohibitive human annotation costs, failing to provide scalable, high-quality supervision. We address this by reconceptualizing data synthesis through the lens of real-world software evolution. Our key insight: Pull Request (PR) sequences naturally embody the supervision signals for long-horizon learning. They decompose complex objectives into verifiable submission units, maintain functional coherence across iterations, and encode authentic refinement patterns through bug-fix histories. Building on this, we propose daVinci-Agency, which systematically mines structured supervision from chain-of-PRs through three interlocking mechanisms: (1) progressive task decomposition via continuous commits, (2) long-term consistency enforcement through unified functional objectives, and (3) verifiable refinement from authentic bug-fix trajectories. Unlike synthetic trajectories that treat each step independently, daVinci-Agency’s PR-grounded structure inherently preserves the causal dependencies and iterative refinements essential for teaching persistent goal-directed behavior and enables natural alignment with project-level, full-cycle task modeling. The resulting trajectories are substantial–averaging 85k tokens and 116 tool calls–yet remarkably data-efficient: fine-tuning GLM-4.6 on 239 daVinci-Agency samples yields broad improvements across benchmarks, notably achieving a 47% relative gain on Toolathlon. Beyond benchmark performance, our analysis confirms…
中文摘要
虽然大型语言模型 (LLM) 擅长短期任务,但将其扩展到长期代理工作流程仍然具有挑战性。核心瓶颈在于缺乏真实的长依赖结构和跨阶段进化动态的训练数据——现有的合成方法要么局限于受模型分布约束的单一特征场景,要么招致高昂的人工注释成本,无法提供可扩展的高质量监督。我们通过现实世界软件演化的视角重新概念化数据合成来解决这个问题。我们的主要见解:拉取请求(PR)序列自然地体现了长期学习的监督信号。他们将复杂的目标分解为可验证的提交单元,在迭代中保持功能的一致性,并通过错误修复历史记录真实的细化模式。在此基础上,我们提出了daVinci-Agency,它通过三种连锁机制从PR链中系统地挖掘结构化监督:(1)通过连续提交进行渐进式任务分解,(2)通过统一的功能目标执行长期一致性,以及(3)从真实的错误修复轨迹中进行可验证的细化。与独立处理每个步骤的合成轨迹不同,daVinci-Agency 以 PR 为基础的结构本质上保留了因果依赖性和迭代细化,这对于教授持久的目标导向行为至关重要,并能够与项目级、全周期任务建模自然对齐。由此产生的轨迹是巨大的——平均 85k 个代币和 116 个工具调用——但数据效率非常高:在 239 个 daVinci-Agency 样本上微调 GLM-4.6 在基准测试中产生了广泛的改进,特别是在 Toolathlon 上实现了 47% 的相对增益。除了基准性能之外,我们的分析还证实…
DAMO开发者矩阵,由阿里巴巴达摩院和中国互联网协会联合发起,致力于探讨最前沿的技术趋势与应用成果,搭建高质量的交流与分享平台,推动技术创新与产业应用链接,围绕“人工智能与新型计算”构建开放共享的开发者生态。
更多推荐

所有评论(0)