CVPR 2026

CVPR 2026 于 2026 年 6 月 3 日至 7 日在美国丹佛 Colorado Convention Center 举行,其中主会时间为 6 月 5 日至 7 日。

据公开统计,CVPR 2026 共有 16092 篇有效投稿进入评审流程,约 4090 篇论文被接收,录用率约 25.42%。本文件基于 CVPR 2026 Open Access 主会论文列表,汇总标题或内容与**目标检测(Object Detection)**相关的论文。

目标检测是计算机视觉的基础任务之一,旨在定位图像/视频中的目标并识别其类别。随着 Transformer、多模态大模型、三维感知与开放词汇学习的发展,目标检测已从封闭类别的边界框回归,扩展到开放世界、跨域部署、三维场景理解与行业专用场景。

说明:

  1. 主分类优先依据论文标题中的目标检测相关表述(如 object detection / detector / YOLO / DETR / open-vocabulary detection 等)。
  2. 标题完全不含目标检测相关表述、或属于异常检测/伪造检测/OOD/变化检测等相邻任务的论文,归入“其他”。
  3. Paper 链接优先 arXiv;若无 arXiv 则回退到 Open Access PDF。Code/Blog 以公开可检索信息为准,未能确认则留空。
  4. Team 在缺少机构主页时,以作者列表缩写作为团队线索。

现将目标检测方向上接收的论文汇总如下(主分类 84 篇;其他相关检测 167 篇)。

通用目标检测框架与方法

  1. AntiStyler: Defending Object Detection Models Against Adversarial Patch Attacks Using Style Removal
  1. Explaining Object Detectors via Collective Contribution of Pixels
  • Paper: https://arxiv.org/abs/2412.00666
  • Code: https://github.com/tttt-0814/VX-CODE
  • Keywords: 通用框架与方法
  • Features: Visual explanations for object detectors are crucial for enhancing their reliability. Object detectors identify and localize instances by assessing multiple visual features collectively.
  • Blog:
  • Team: Toshinori Yamauchi,Hiroshi Kera,Kazuhiko Kawamoto
  1. Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection
  • Paper: https://arxiv.org/abs/2603.24166
  • Code:
  • Keywords: 通用框架与方法
  • Features: Most referring object detection (ROD) models, especially the modern grounding detectors, are designed for data-rich conditions, yet many practical deployments, such as robotics, augmented reality, and other specialized d…
  • Blog:
  • Team: Xu Zhang,Zhe Chen,Jing Zhang,Dacheng Tao
  1. InsCal: Calibrated Multi-Source Fully Test-Time Prompt Tuning for Object Detection
  1. Mind the Gap: Transferring Labels to Align Object Detection Datasets
  1. PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection
  • Paper: https://arxiv.org/abs/2603.06917
  • Code:
  • Keywords: DETR
  • Features: 提出 PaQ-DETR;Detection Transformer (DETR) has redefined object detection by casting it as a set prediction task within an end-to-end framework. Despite its elegance, DETR and its variants still rely on fixed learnable queries and suf…
  • Blog:
  • Team: Zhengjian Kang,Jun Zhuang,Kangtong Mo 等
  1. RAVEN: Radar Adaptive Vision Encoders for Efficient Chirp-wise Object Detection and Segmentation
  • Paper: https://arxiv.org/abs/2604.04490
  • Code:
  • Keywords: Radar
  • Features: 提出 RAVEN;This paper presents RAVEN, a computationally efficient deep learning architecture for FMCW radar perception. The method processes raw ADC data in a chirp-wise streaming manner, preserves MIMO structure through independen…
  • Blog:
  • Team: Anuvab Sen,Mir Sayeed Mohammad,Saibal Mukhopadhyay
  1. Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding

实时/YOLO/高效检测

  1. AKCMamba-YOLO: Selective State Space Models For Real-Time Object Detection
  1. Does YOLO Really Need to See Every Training Image in Every Epoch?
  • Paper: https://arxiv.org/abs/2603.17684
  • Code:
  • Keywords: YOLO
  • Features: YOLO detectors are known for their fast inference speed, yet training them remains unexpectedly time-consuming due to their exhaustive pipeline that processes every training image in every epoch, even when many images ha…
  • Blog:
  • Team: Xingxing Xie,Jiahua Dong,Junwei Han,Gong Cheng
  1. YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection
  • Paper: https://arxiv.org/abs/2512.23273
  • Code:
  • Keywords: YOLO, Transformer, Real-Time
  • Features: 提出 YOLO-Master;Existing Real-Time Object Detection (RTOD) methods commonly adopt YOLO-like architectures for their favorable trade-off between accuracy and speed. However, these models rely on static dense computation that applies unif…
  • Blog:
  • Team: Xu Lin,Jinlong Peng,Zhenye Gan 等
  1. YOLO-ULM: Ultra-Lightweight Models for Real-Time Object Detection

开放词汇/开放世界检测

  1. Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency Learning
  • Paper: https://arxiv.org/abs/2603.26179
  • Code: https://github.com/bozhao-li/CCL
  • Keywords: Open-Vocabulary, Robustness
  • Features: Recent advances in open-vocabulary object detection focus primarily on two aspects: scaling up datasets and leveraging contrastive learning to align language and vision modalities. However, these approaches often neglect…
  • Blog:
  • Team: Bozhao Li,Shaocong Wu,Tong Shao 等
  1. Detecting Unknown Objects via Energy-based Separation for Open World Object Detection
  • Paper: https://arxiv.org/abs/2603.29954
  • Code:
  • Keywords: Open-World
  • Features: In this work, we tackle the problem of Open World Object Detection (OWOD). This challenging scenario requires the detector to incrementally learn to classify known objects without forgetting while identifying unknown obj…
  • Blog:
  • Team: Jun-Woo Heo,Keonhee Park,Gyeong-Moon Park
  1. NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection
  • Paper: https://arxiv.org/abs/2603.21069
  • Code:
  • Keywords: Open-Vocabulary
  • Features: 提出 NoOVD;Despite the remarkable progress in open-vocabulary object detection (OVD), a significant gap remains between the training and testing phases. During training, the RPN and RoI heads often misclassify unlabeled novel-categ…
  • Blog:
  • Team: Yupeng Zhang,Ruize Han,Zhiwei Chen 等
  1. Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
  • Paper: https://arxiv.org/abs/2604.04444
  • Code:
  • Keywords: Open-Vocabulary
  • Features: Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-training, existing OVOD methods achieve strong detection performance on general…
  • Blog:
  • Team: Weihao Cao,Runqi Wang,Xiaoyue Duan 等
  1. Prompt-Free Unknown Label Generation for Open World Detection in Remote Sensing
  1. SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection
  • Paper: https://arxiv.org/abs/2603.26109
  • Code:
  • Keywords: Open-Vocabulary, Camouflaged Object
  • Features: 提出 SDDF;Open-vocabulary object detection (OVOD) aims to detect known and unknown objects in the open world by leveraging text prompts. Benefiting from the emergence of large-scale vision–language pre-trained models, OVOD has de…
  • Blog:
  • Team: Jiaming Liang,Yifeng Zhan,Chunlin Liu 等
  1. SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Names
  1. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
  • Paper: https://arxiv.org/abs/2605.10130
  • Code:
  • Keywords: Open-Vocabulary, Multimodal, Vision-Language, Thermal
  • Features: 提出 Thermal-Det;Existing open-vocabulary detectors focus on RGB images and fail to generalize to thermal imagery, where low texture and emissivity variations challenge RGB-based semantics. We present Thermal-Det, the first large languag…
  • Blog:
  • Team: Yasiru Ranasinghe,Elim Schenck,Florence Yellin 等
  1. ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detection
  1. WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
  • Paper: https://arxiv.org/abs/2512.12309
  • Code:
  • Keywords: Open-Vocabulary
  • Features: 提出 WeDetect;Open-vocabulary object detection aims to detect arbitrary classes via text prompts. Methods without cross-modal fusion layers (non-fusion) offer faster inference by treating recognition as a retrieval problem, \ie, match…
  • Blog:
  • Team: Shenghao Fu,Yukun Su,Fengyun Rao 等

少样本/增量/零样本检测

  1. A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps
  • Paper: https://arxiv.org/abs/2603.28182
  • Code: https://github.com/Intellindust-AI-Lab/FT-FSOD
  • Keywords: Few-Shot
  • Features: Few-shot object detection (FSOD) is challenging due to unstable optimization and limited generalization arising from the scarcity of training samples. To address these issues, we propose a hybrid ensemble decoder that en…
  • Blog:
  • Team: Xuanlong Yu,Youyang Sha,Longfei Liu 等
  1. AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detection
  1. Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detection
  • Paper: https://arxiv.org/abs/2603.02286
  • Code: https://github.com/zyt95579/PDP
  • Keywords: Incremental Learning
  • Features: Incremental Object Detection (IOD) aims to continuously learn new object categories without forgetting previously learned ones. Recently, prompt-based methods have gained popularity for their replay-free design and param…
  • Blog:
  • Team: Yaoteng Zhang,Qing Zhou,Junyu Gao,Qi Wang
  1. Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detection
  1. EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
  • Paper: https://arxiv.org/abs/2602.20985
  • Code:
  • Keywords: Incremental Learning, DETR, Transformer
  • Features: 提出 EW-DETR;Real-world object detection must operate in evolving environments where new classes emerge, domains shift, and unseen objects must be identified as "unknown": all without accessing prior data. We introduce Evolvi…
  • Blog:
  • Team: Munish Monga,Vishal Chudasama,Pankaj Wasnik,C.V. Jawahar
  1. Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation
  1. Parameterized Prompt for Incremental Object Detection
  • Paper: https://arxiv.org/abs/2510.27316
  • Code: https://github.com/EMLS-ICTCAS/P2IOD
  • Keywords: Incremental Learning
  • Features: Recent studies have demonstrated that incorporating trainable prompts into pretrained models enables effective incremental learning. However, the application of prompts in incremental object detection (IOD) remains under…
  • Blog:
  • Team: Zijia An,Boyu Diao,Ruiqi Liu 等
  1. Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection
  • Paper: https://arxiv.org/abs/2603.18541
  • Code:
  • Keywords: Few-Shot
  • Features: Cross-domain few-shot object detection (CD-FSOD) aims to adapt pretrained detectors from a source domain to target domains with limited annotations, suffering from severe domain shifts and data scarcity problems. In this…
  • Blog:
  • Team: Yongwei Jiang,Yixiong Zou,Yuhua Li,Ruixuan Li

弱监督/半监督/主动学习

  1. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection
  1. Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
  • Paper: https://arxiv.org/abs/2511.14197
  • Code:
  • Keywords: 弱监督/半监督/主动学习
  • Features: High-quality data has become a primary driver of progress under scale laws, with curated datasets often outperforming much larger unfiltered ones at lower cost. Online data curation extends this idea by dynamically selec…
  • Blog:
  • Team: Zitang Sun,Masakazu Yoshimura,Junji Otsuka 等
  1. Partial Weakly-Supervised Oriented Object Detection
  1. Portable Active Learning for Object Detection
  • Paper: https://arxiv.org/abs/2605.10349
  • Code:
  • Keywords: Active Learning
  • Features: Annotating bounding boxes is costly and limits the scalability of object detection. This challenge is compounded by the need to preserve high accuracy while minimizing manual effort in real-world applications.
  • Blog:
  • Team: Rashi Sharma,Justin Timothy C. Bersamin,Karthikk Subramanian
  1. TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement
  • Paper: https://arxiv.org/abs/2602.23120
  • Code:
  • Keywords: Weakly-Supervised
  • Features: 提出 TriLite;Weakly supervised object localization (WSOL) aims to localize target objects in images using only image-level labels. Despite recent progress, many approaches still rely on multi-stage pipelines or full fine-tuning of la…
  • Blog:
  • Team: Arian Sabaghi,Jose Oramas

域适应/域泛化

  1. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compression
  1. CD-Buffer: Complementary Dual-Buffer Framework for Test-Time Adaptation in Adverse Weather Object Detection
  • Paper: https://arxiv.org/abs/2603.26092
  • Code:
  • Keywords: Test-Time Adaptation
  • Features: 提出 CD-Buffer;Test-Time Adaptation (TTA) enables real-time adaptation to domain shifts without off-line retraining. Recent TTA methods have predominantly explored additive approaches that introduce lightweight modules for feature refi…
  • Blog:
  • Team: Youngjun Song,Hyeongyu Kim,Dosik Hwang
  1. DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection
  • Paper: https://arxiv.org/abs/2603.18757
  • Code: https://github.com/enesdoruk/DA-Mamba
  • Keywords: Domain Adaptation, Mamba
  • Features: 提出 DA-Mamba;Domain Adaptive Object Detection (DAOD) aims to transfer detectors from a labeled source domain to an unlabeled target domain. Existing DAOD methods employ multi-granularity feature alignment to learn domain-invariant re…
  • Blog:
  • Team: Haochen Li,Rui Zhang,Hantao Yao 等
  1. Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object Detection
  1. Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
  • Paper: https://arxiv.org/abs/2512.17514
  • Code:
  • Keywords: 域适应/域泛化
  • Features: Current state-of-the-art approaches in Source-Free Object Detection (SFOD) typically rely on Mean-Teacher self-labeling. However, domain shift often reduces the detector's ability to maintain strong object-focused re…
  • Blog:
  • Team: Sairam VCR,Rishabh Lalla,Aveen Dayal 等

3D目标检测

  1. CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
  • Paper: https://arxiv.org/abs/2603.23276
  • Code: https://github.com/IMPL-Lab/CCF
  • Keywords: Domain Generalization, 3D Object Detection, Multimodal
  • Features: 提出 CCF;Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training.
  • Blog:
  • Team: Yuchen Wu,Kun Wang,Yining Pan,Na Zhao
  1. CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
  • Paper: https://arxiv.org/abs/2603.05042
  • Code: https://github.com/kwong292521/CoIn3D
  • Keywords: 3D Object Detection
  • Features: 提出 CoIn3D;Multi-camera 3D object detection (MC3D) has attracted increasing attention with the growing deployment of multi-sensor physical agents, such as robots and autonomous vehicles. However, MC3D models still struggle to gener…
  • Blog:
  • Team: Zhaonian Kuang,Rui Ding,Haotian Wang 等
  1. Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments
  • Paper: https://arxiv.org/abs/2604.07997
  • Code: https://github.com/zyrant/FI3Det
  • Keywords: Few-Shot, Incremental Learning, 3D Object Detection
  • Features: Incremental 3D object perception is a critical step toward embodied intelligence in dynamic indoor environments. However, existing incremental 3D detection methods rely on extensive annotations of novel classes for satis…
  • Blog:
  • Team: Yun Zhu,Jianjun Qian,Jian Yang 等
  1. H2A2: Homogeneity-Aware and Heterogeneity-Aware Feature Perception for Unified Indoor 3D Object Detection
  1. MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
  • Paper: https://arxiv.org/abs/2604.01646
  • Code: https://github.com/VisualAIKHU/MonoSAOD
  • Keywords: 3D Object Detection, Monocular
  • Features: 提出 MonoSAOD;Monocular 3D object detection has achieved impressive performance on densely annotated datasets. However, it struggles when only a fraction of objects are labeled due to the high cost of 3D annotation.
  • Blog:
  • Team: Junyoung Jung,Seokwon Kim,Jung Uk Kim
  1. Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift
  • Paper: https://arxiv.org/abs/2605.05328
  • Code:
  • Keywords: 3D Object Detection, Robustness
  • Features: 提出 Query2Uncertainty;Reliable uncertainty estimation for 3D object detection is critical for deploying safe autonomous systems, yet modern detectors remain poorly calibrated, especially under distribution shifts. Although post-hoc calibratio…
  • Blog:
  • Team: Till Beemelmanns,Alexey Nekrasov,Stefan Vilceanu 等
  1. R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection
  • Paper: https://arxiv.org/abs/2603.11566
  • Code: https://github.com/VDIGPKU/R4Det
  • Keywords: 3D Object Detection, Radar
  • Features: 提出 R4Det;4D radar-camera sensing configuration has gained increasing importance in autonomous driving. However, existing 3D object detection methods that fuse 4D Radar and camera data confront several challenges.
  • Blog:
  • Team: Zhongyu Xia,Yousen Tang,Yongtao Wang 等
  1. RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cue for 3D Object Detection
  • Paper: https://arxiv.org/abs/2507.19856
  • Code:
  • Keywords: 3D Object Detection, Monocular, Radar
  • Features: 提出 RaGS;4D millimeter-wave radar is a promising sensing modality for autonomous driving, yet effective 3D object detection from 4D radar and monocular images remains challenging. Existing fusion approaches either rely on instanc…
  • Blog:
  • Team: Xiaokai Bai,Chenxu Zhou,Lianqing Zheng 等
  1. RARE: Learn to RAnk and REtrieve for Monocular 3D Object Detection
  1. Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
  • Paper: https://arxiv.org/abs/2604.14563
  • Code: https://github.com/Mingqj/SEPatch3D
  • Keywords: 3D Object Detection, Multi-View, Transformer
  • Features: Vision Transformer (ViT)-based sparse multi-view 3D object detectors have achieved remarkable accuracy but still suffer from high inference latency due to heavy token processing. To accelerate these models, token compres…
  • Blog:
  • Team: Mingqian Ji,Shanshan Zhang,Jian Yang
  1. SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors
  • Paper: https://arxiv.org/abs/2505.22499
  • Code:
  • Keywords: BEV, Robustness
  • Features: 提出 SABER;Adversarial robustness of BEV 3D object detectors is critical for autonomous driving (AD). Existing invasive attacks require altering the target vehicle itself (e.g.
  • Blog:
  • Team: Aixuan Li,Mochu Xiang,Bosen Hou 等
  1. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
  • Paper: https://arxiv.org/abs/2604.18476
  • Code:
  • Keywords: 3D Object Detection
  • Features: 提出 SemLT3D;Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily prioritize overall performance while overlooking the severe long-ta…
  • Blog:
  • Team: Hao Vo,Khoa Vo,Thinh Phan 等
  1. SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection
  • Paper: https://arxiv.org/abs/2511.06702
  • Code:
  • Keywords: 3D Object Detection, Monocular
  • Features: 提出 SPAN;Existing monocular 3D detectors typically tame the pronounced nonlinear regression of 3D bounding box through decoupled prediction paradigm, which employs multiple branches to estimate geometric center, depth, dimensions…
  • Blog:
  • Team: Yifan Wang,Yian Zhao,Fanqi Pu 等
  1. Spe-BEVHead: Rethinking the Detection Head Design for Bird’s-Eye-View Object Detection
  1. SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
  • Paper: https://arxiv.org/abs/2605.14110
  • Code:
  • Keywords: 3D Object Detection, Multi-View, Transformer
  • Features: 提出 SToRe3D;Vision Transformers (ViTs) enable strong multi-view 3D detection but are limited by high inference latency from dense token and query processing across multiple views and large 3D regions. Existing sparsity methods, desi…
  • Blog:
  • Team: Sandro Papais,Lezhou Feng,Charles Cossette,Lingting Ge
  1. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection
  1. TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection
  1. Towards Intrinsic-Aware Monocular 3D Object Detection
  • Paper: https://arxiv.org/abs/2603.27059
  • Code:
  • Keywords: 3D Object Detection, Monocular
  • Features: Monocular 3D object detection (Mono3D) aims to infer object locations and dimensions in 3D space from a single RGB image. Despite recent progress, existing methods remain highly sensitive to camera intrinsics and struggl…
  • Blog:
  • Team: Zhihao Zhang,Abhinav Kumar,Xiaoming Liu
  1. Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
  • Paper: https://arxiv.org/abs/2505.04594
  • Code:
  • Keywords: 3D Object Detection, Monocular
  • Features: Monocular 3D detection (Mono3D) aims to infer 3D bounding boxes from a single RGB image. Without auxiliary sensors such as LiDAR, this task is inherently ill-posed since the 3D-to-2D projection introduces depth ambiguity…
  • Blog:
  • Team: Zhihao Zhang,Abhinav Kumar,Girish Chandar Ganesan,Xiaoming Liu
  1. VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
  • Paper: https://arxiv.org/abs/2603.00912
  • Code: https://github.com/yangcaoai/VGGT-Det-CVPR2026
  • Keywords: 3D Object Detection, Multi-View
  • Features: 提出 VGGT-Det;Current multi-view indoor 3D object detectors rely on sensor geometry that is costly to obtain (i.e., precisely calibrated multi-view camera poses) to fuse multi-view information into a global scene representation, limit…
  • Blog:
  • Team: Yang Cao,Feize Wu,Dave Zhenyu Chen 等
  1. Zoo3D: Zero-Shot 3D Object Detection at Scene Level
  • Paper: https://arxiv.org/abs/2511.20253
  • Code: https://github.com/col14m/zoo3d
  • Keywords: Zero-Shot, 3D Object Detection
  • Features: 提出 Zoo3D;3D object detection is fundamental for spatial understanding. Real-world environments demand models capable of recognizing diverse, previously unseen objects, which remains a major limitation of closed-set methods.
  • Blog:
  • Team: Andrey Lemeshko,Bulat Gabdullin,Nikita Drozdov 等

视频目标检测

  1. D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Network
  1. When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detection

小目标检测

  1. BDNet:Bio-Inspired Dual-Backbone Small Object Detection Network
  1. DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection
  1. ElasticFormer: Detecting Objects in HRW Shots via Elastic Computing Vision Transformer
  1. Towards Persistence: Learning Topological Constraints for Event-based Small Object Detection

显著性/伪装目标检测

  1. Beyond Appearance: Camouflaged Object Detection via Geometric Structure
  1. Generalizable Co-Salient Object Detection via Mixed Content-Style Modulation
  1. M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection
  1. TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection
  • Paper: https://arxiv.org/abs/2604.00549
  • Code: https://github.com/hzz-yy/TF-SSD
  • Keywords: Salient Object
  • Features: 提出 TF-SSD;Co-salient Object Detection (CoSOD) aims to segment salient objects that consistently appear across a group of related images. Despite the notable progress achieved by recent training-based approaches, they still remain…
  • Blog:
  • Team: Zhijin He,Shuo Jin,Siyue Yu 等
  1. Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detection

遥感/旋转/无人机目标检测

  1. Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing Images
  • Paper: https://arxiv.org/abs/2512.24074
  • Code:
  • Keywords: Remote Sensing
  • Features: Fine-grained remote sensing datasets often use hierarchical label structures to differentiate objects in a coarse-to-fine manner, with each object annotated across multiple levels. However, embedding this semantic hierar…
  • Blog:
  • Team: Jingzhou Chen,Dexin Chen,Fengchao Xiong 等
  1. Fourier Angle Alignment for Oriented Object Detection in Remote Sensing
  • Paper: https://arxiv.org/abs/2602.23790
  • Code: https://github.com/gcy0423/Fourier-Angle-Alignment
  • Keywords: Remote Sensing, Oriented Detection
  • Features: In remote sensing rotated object detection, mainstream methods suffer from two bottlenecks, directional incoherence at detector neck and task conflict at detecting head. Ulitising fourier rotation equivariance, we introd…
  • Blog:
  • Team: Changyu Gu,Linwei Chen,Lin Gu,Ying Fu
  1. Rotation Invariant and Symmetry Aware Pixel Difference Network for Remote Sensing Object Detection
  1. Tri-Modal Fusion Transformers for UAV-based Object Detection
  • Paper: https://arxiv.org/abs/2604.16630
  • Code:
  • Keywords: Transformer, UAV
  • Features: Reliable UAV object detection requires robustness to illumination changes, motion blur, and scene dynamics that suppress RGB cues. Thermal long-wave infrared (LWIR) sensing preserves contrast in low light, and event came…
  • Blog:
  • Team: Craig Iaboni,Pramod Abichandani
  1. Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection
  • Paper: https://arxiv.org/abs/2604.02966
  • Code: https://github.com/Sirius-Li/UAVGen
  • Keywords: UAV
  • Features: Unmanned aerial vehicle (UAV) based object detection is a critical but challenging task, when applied in dynamically changing scenarios with limited annotated training data. Layout-to-image generation approaches have pro…
  • Blog:
  • Team: Wenhao Li,Zimeng Wu,Yu Wu 等
  1. VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detection

多模态/跨模态目标检测

  1. Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object Detection
  1. Distribution-Aligned Multimodal Fusion for Robust Object Detection
  1. RHCNet: Residual-Guided Hierarchical Calibration Network for Robust Underwater Object Detection
  1. Spike-driven Discrete Aggregation for Event-based Object Detection

行人搜索

  1. FSLoRA: Harmonizing Detection and Re-Identification via Freq-Spatial Low-Rank Adapter for One-Stage Person Search

专用场景检测

  1. Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
  • Paper: https://arxiv.org/abs/2511.18385
  • Code:
  • Keywords: Multimodal
  • Features: Automatic X-ray prohibited items detection is vital for security inspection and has been widely studied. Traditional methods rely on visual modality, often struggling with complex threats.
  • Blog:
  • Team: Chuang Peng,Renshuai Tao,Zhongwei Ren 等

其他

以下论文标题中出现 detect/detection 等字样,但不完全属于目标检测主线(如异常检测、伪造/生成内容检测、OOD 检测、变化检测、动作检测等)。按子方向分组列出。

异常检测

  1. A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection
  1. ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoning
  1. Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detection
  1. Anomaly-Related Residual Fields for Cross-domain Anomaly Detection
  1. AnomalyVFM – Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors
  1. Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection
  1. Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
  1. CHAL: Causal-guided Hierarchical Anomaly-aware Learning for Moving Infrared Small Target Detection
  1. Complementary Prototype Mapping for Efficient Multimodal Anomaly Detection
  1. Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection
  1. DLVP-CLIP: Enhancing Fine-Grained Zero-Shot Anomaly Detection via Dynamic Local Visual Prompting
  1. Dual-Prototype-Guided Multi-task Learning for Unsupervised Anomaly Detection and Classification
  1. FastRef: Fast Prototype Refinement for Few-shot Industrial Anomaly Detection
  1. FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
  1. Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity Learning
  1. From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection
  1. Geometry-Aligned and Anomaly-Aware Reconstruction for 3D Anomaly Detection
  1. GPFlow: Gaussian Prototype Probability Flow for Unsupervised Multi-Modal Anomaly Detection
  1. GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning
  1. Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
  1. Hunting Normality from Query Sample via Residual Learning for Generalist Anomaly Detection
  1. InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
  1. Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection
  1. LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly Detection
  1. Learning from Noisy Supervision: A Denoising-Debiasing Framework for Weakly Supervised Video Anomaly Detection
  1. MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
  1. MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly Detection
  1. Multi-Prototype Compactness and Boundary-Aware Synthesis for Unsupervised Anomaly Detection
  1. No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
  1. Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly Detection
  1. PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection
  1. RAID: Retrieval-Augmented Anomaly Detection
  1. RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation
  1. Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision
  1. SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modeling
  1. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection
  1. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection
  1. Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective
  1. UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression
  1. VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer
  1. Wavelet-Driven 3D Anomaly Detection under Pose-Agnostic and Sparse-View
  1. Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning

伪造/生成内容检测

  1. A Debiased Reconstruction-based Framework for Training-Free Detection of AI-Generated Images
  1. A Difference-in-Difference Approach to Detecting AI-Generated Images
  1. A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real World
  1. Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
  1. All in One: Unifying Deepfake Detection, Tampering Localization, and Source Tracing with a Robust Landmark-Identity Watermark
  1. AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
  1. Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake Detection
  1. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
  1. Cross-modal Representation Learning for Diffusion-generated Image Detection
  1. Decoupling Bias, Aligning Distributions: Synergistic Fairness Optimization for Deepfake Detection
  1. DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detection
  1. Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplification
  1. Detecting Compressed AI-Generated Images via Phase Spectrum Robustness
  1. DFD-HR: Generalizable Deepfake Detection via Hierarchical Routing Learning
  1. DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
  1. Diversity over Uniformity: Rethinking Representation in Generated Image Detection
  1. Enabling Supervised Learning of Generative Signatures for Generalized AI-Generated Images Detection
  1. FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models
  1. Investigating Self-Supervised Representations for Audio-Visual Deepfake Detection
  1. Layer Consistency Matters: Elegant Latent Transition Discrepancy for Generalizable Synthetic Image Detection
  1. Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
  1. Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
  1. Pixels Don’t Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
  1. PPM-CLIP: Probabilistic Prompt Modeling for Generalizable AI-Generated Image Detection
  1. ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation
  1. SAIDO: Generalizable Detection of AI-Generated Images via Scene-Aware and Importance-Guided Dynamic Optimization in Continual Learning
  1. Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
  1. Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
  1. Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
  1. TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
  1. Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection
  1. UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection
  1. Unleashing Vision-Language Semantics for Deepfake Video Detection
  1. VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation
  1. X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
  1. Your One-Stop Solution for AI-Generated Video Detection
  1. Zero-shot Detection of AI-Generated Image via RAW-RGB Alignment

OOD检测

  1. Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models
  1. ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
  1. Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal Transport
  1. Enhancing Out-of-Distribution Detection with Extended Logit Normalization
  1. Learning Latent Concepts for Detecting Out-of-Distribution Objects
  1. Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs
  1. Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis
  1. Neural Distribution Prior for LiDAR Out-of-Distribution Detection
  1. RankOOD - Class Ranking-based Out-of-Distribution Detection
  1. Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection
  1. The Invisible Gorilla Effect in Out-of-distribution Detection
  1. TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models
  1. UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language Modeling

变化检测

  1. Changes in Real Time: Online Scene Change Detection with Multi-View Fusion
  1. OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery
  1. RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change Detection
  1. SRGCD: Stability-Driven Region Growth Framework for 3D Change Detection
  1. UniChange: Unifying Change Detection with Multimodal Large Language Model

动作检测

  1. Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
  1. Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
  1. MoVie: Broaden Your Views with Human Motion for Action Detection
  1. RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
  1. Streamlined Open-Vocabulary Human-Object Interaction Detection
  1. TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

关键点/地标检测

  1. BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird’s-Eye View Images
  1. EV-CGNet: Co-visible Focused 3D-guided 2D Event Keypoint Detection Network
  1. From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

幻觉检测

  1. Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations
  1. Lyapunov Probes for Hallucination Detection in Large Foundation Models
  1. PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models
  1. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
  1. ZINA: Multimodal Fine-grained Hallucination Detection and Editing

讽刺/语义检测

  1. MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection

跟踪相关

  1. From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking

其他检测相关

  1. Adaptive Confidence Regularization for Multimodal Failure Detection
  1. ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
  1. AutoDebias: An Automated Framework for Detecting and Mitigating Backdoor Biases in Text-to-Image Models
  1. AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Models
  1. BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
  1. Breaking Spurious Correlations: Uncertainty-Driven Causal Transformers for AU Detection
  1. Bulk RNA-seq Guided Multi-modal Detection of Anomalous Regions in Human Cancer via Spatial Transcriptomics
  1. BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship Detection
  1. Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
  1. COPYLENS: Towards Copyrighted Characters Infringement Detection via Copyright-Aware Prompt Learning
  1. CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
  1. Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets
  1. DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video
  1. Detect Any AI-Counterfeited Text Image
  1. Detect Anything via Next Point Prediction
  1. DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive Imaging
  1. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection
  1. FedSDR: Federated Graph Learning with Structural Noise Detection and Reconstruction
  1. Geometry-driven OOD Detectors Are Class-Incremental Learners
  1. Ghost-FWL: A Large-Scale Full-Waveform LiDAR Dataset for Ghost Detection and Removal
  1. GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
  1. Homaloidal parametrization for detecting critical two-view configurations
  1. KLIP: Localized Distribution Shift Detection via KL-Divergence with Diffusion Priors in Inverse Problems
  1. Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
  1. Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection
  1. LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
  1. Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
  1. MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision
  1. MEMO: Human-like Crisp Edge Detection Using Masked Edge Prediction
  1. MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
  1. Neural Field-Based 3D Surface Reconstruction of Microstructures from Multi-Detector Signals in Scanning Electron Microscopy
  1. Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting
  1. OpenFS: Multi-Hand-Capable Fingerspelling Recognition with Implicit Signing-Hand Detection and Frame-Wise Letter-Conditioned Synthesis
  1. OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
  1. Physical Adversarial Clothing Evades Visible-Thermal Detectors via Non-Overlapping RGB-T Pattern
  1. Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
  1. Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction
  1. ReManNet: A Riemannian Manifold Network for Monocular 3D Lane Detection
  1. RPGFusion: 4D Radar Prior-Guided Multi-Modal Fusion for 3D Detection
  1. SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion
  1. Scene Reconstruction as Mapping Priors for 3D Detection
  1. Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspective
  1. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detection
  1. Similarity-Consistent Likelihood Diffusion enables Hidden Person Detection from Wall Reflections
  1. SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images
  1. Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
  1. Target-Aware Invertible Encoder with Reconstruction Guidance for Infrared Small Target Detection
  1. Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach
  1. Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods
  1. TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
  1. TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV Dramas
  1. UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detection
  1. Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary Detectors
  1. Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

总结

从本届接收论文来看,CVPR 2026 目标检测方向呈现以下趋势:

  1. 3D 目标检测体量最大:单目、多视角、BEV、LiDAR/Radar 融合与室内外统一检测持续活跃;雷达-相机融合、Gaussian Splatting 先验、token 压缩与不确定性估计是常见技术点。
  2. 开放词汇/开放世界检测成为主线之一:Open-Vocabulary Detection、Open-World Detection、未知类别发现、检索式检测(如 WeDetect)与热成像开放词汇检测等方向快速增长。
  3. 数据高效学习受重视:少样本、跨域少样本、增量检测、主动学习、在线数据筛选与弱监督设定显著增多,反映标注成本与持续部署需求。
  4. 实时高效架构回潮:YOLO 体系、Mamba/SSM 混合结构、轻量化模型与训练策略优化重新成为焦点。
  5. 场景专用化加深:遥感/旋转框、UAV、小目标、伪装/显著性、水下、X-ray 安检、Person Search 等方法继续细分。
  6. 检测概念外延明显:异常检测、深度伪造/生成内容检测、OOD 检测、变化检测等“泛检测”任务数量可观,但与经典目标检测主线有所区分。

总体而言,CVPR 2026 目标检测研究在通用检测框架演进之外,更强调开放词汇泛化、三维感知、数据高效学习与真实场景鲁棒落地。

参考资料

  1. CVPR 2026 Official Website
  2. CVPR 2026 Accepted Papers (Open Access)
  3. arXiv.org

(注:文档部分内容由 AI 生成;Code/Blog/单位信息以公开网页检索为准,如有遗漏欢迎补充指正。)

Logo

DAMO开发者矩阵,由阿里巴巴达摩院和中国互联网协会联合发起,致力于探讨最前沿的技术趋势与应用成果,搭建高质量的交流与分享平台,推动技术创新与产业应用链接,围绕“人工智能与新型计算”构建开放共享的开发者生态。

更多推荐