Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Shenghao, Yan, Junkai, Yang, Qize, Wei, Xihan, Xie, Xiaohua, Zheng, Wei-Shi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
Sim-DETR: Unlock DETR for Temporal Sentence Grounding
by: Tang, Jiajin, et al.
Published: (2025)
by: Tang, Jiajin, et al.
Published: (2025)
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
by: Peng, Yi-Xing, et al.
Published: (2025)
by: Peng, Yi-Xing, et al.
Published: (2025)
DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation
by: Yan, Junkai, et al.
Published: (2024)
by: Yan, Junkai, et al.
Published: (2024)
MS-DETR: Efficient DETR Training with Mixed Supervision
by: Zhao, Chuyang, et al.
Published: (2024)
by: Zhao, Chuyang, et al.
Published: (2024)
FrozenSeg: Harmonizing Frozen Foundation Models for Open-Vocabulary Segmentation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
MDS-DETR: DETR with Masked Duplicate Suppressor
by: Lee, Chanho, et al.
Published: (2026)
by: Lee, Chanho, et al.
Published: (2026)
ViSpeak: Visual Instruction Feedback in Streaming Videos
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
UAV-DETR: DETR for Anti-Drone Target Detection
by: Yang, Jun, et al.
Published: (2026)
by: Yang, Jun, et al.
Published: (2026)
DQ-DETR: DETR with Dynamic Query for Tiny Object Detection
by: Huang, Yi-Xin, et al.
Published: (2024)
by: Huang, Yi-Xin, et al.
Published: (2024)
CGF-DETR: Cross-Gated Fusion DETR for Enhanced Pneumonia Detection in Chest X-rays
by: Wu, Yefeng, et al.
Published: (2025)
by: Wu, Yefeng, et al.
Published: (2025)
D$^3$R-DETR: DETR with Dual-Domain Density Refinement for Tiny Object Detection in Aerial Images
by: Wen, Zixiao, et al.
Published: (2026)
by: Wen, Zixiao, et al.
Published: (2026)
RiO-DETR: DETR for Real-time Oriented Object Detection
by: Hu, Zhangchi, et al.
Published: (2026)
by: Hu, Zhangchi, et al.
Published: (2026)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
by: Li, Yuan-Ming, et al.
Published: (2025)
by: Li, Yuan-Ming, et al.
Published: (2025)
Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement
by: Hou, Xiuquan, et al.
Published: (2024)
by: Hou, Xiuquan, et al.
Published: (2024)
Understanding differences in applying DETR to natural and medical images
by: Xu, Yanqi, et al.
Published: (2024)
by: Xu, Yanqi, et al.
Published: (2024)
Caries DETR: Tooth Structure-aware Prior and Lesion-aware Dynamic Loss Refinement for DETR Based Caries Detection
by: Liu, Xuefen, et al.
Published: (2026)
by: Liu, Xuefen, et al.
Published: (2026)
Siamese-DETR for Generic Multi-Object Tracking
by: Liu, Qiankun, et al.
Published: (2023)
by: Liu, Qiankun, et al.
Published: (2023)
DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection
by: Wang, Siheng, et al.
Published: (2026)
by: Wang, Siheng, et al.
Published: (2026)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Dome-DETR: DETR with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object Detection
by: Hu, Zhangchi, et al.
Published: (2025)
by: Hu, Zhangchi, et al.
Published: (2025)
CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection
by: Chen, Qibo, et al.
Published: (2024)
by: Chen, Qibo, et al.
Published: (2024)
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
Le-DETR: Revisiting Real-Time Detection Transformer with Efficient Encoder Design
by: Huang, Jiannan, et al.
Published: (2026)
by: Huang, Jiannan, et al.
Published: (2026)
Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss
by: Cai, Zhi, et al.
Published: (2023)
by: Cai, Zhi, et al.
Published: (2023)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
OccupancyDETR: Using DETR for Mixed Dense-sparse 3D Occupancy Prediction
by: Jia, Yupeng, et al.
Published: (2023)
by: Jia, Yupeng, et al.
Published: (2023)
PillarDETR: YOLO-Backbone and RT-DETR Head for Real-Time 3D Object Detection
by: Kadvani, Smit, et al.
Published: (2026)
by: Kadvani, Smit, et al.
Published: (2026)
AO-DETR: Anti-Overlapping DETR for X-Ray Prohibited Items Detection
by: Li, Mingyuan, et al.
Published: (2024)
by: Li, Mingyuan, et al.
Published: (2024)
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
by: Orlova, Svetlana, et al.
Published: (2026)
by: Orlova, Svetlana, et al.
Published: (2026)
RT-DETR++ for UAV Object Detection
by: Shufang, Yuan
Published: (2025)
by: Shufang, Yuan
Published: (2025)
Utilizing dynamic sparsity on pretrained DETR
by: Sedghi, Reza, et al.
Published: (2025)
by: Sedghi, Reza, et al.
Published: (2025)
Pattern-Enhanced RT-DETR for Multi-Class Battery Detection
by: Zhong, Xu, et al.
Published: (2026)
by: Zhong, Xu, et al.
Published: (2026)
Increasing the Efficiency of DETR for Maritime High-Resolution Images
by: Yehuala, Tinsae, et al.
Published: (2026)
by: Yehuala, Tinsae, et al.
Published: (2026)
Source-Free Domain Adaptation with Frozen Multimodal Foundation Model
by: Tang, Song, et al.
Published: (2023)
by: Tang, Song, et al.
Published: (2023)
Bridge Past and Future: Overcoming Information Asymmetry in Incremental Object Detection
by: Mo, Qijie, et al.
Published: (2024)
by: Mo, Qijie, et al.
Published: (2024)
MonoDINO-DETR: Depth-Enhanced Monocular 3D Object Detection Using a Vision Foundation Model
by: Kim, Jihyeok, et al.
Published: (2025)
by: Kim, Jihyeok, et al.
Published: (2025)
Similar Items
-
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
by: Fu, Shenghao, et al.
Published: (2025) -
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
by: Fu, Shenghao, et al.
Published: (2025) -
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
by: Fu, Shenghao, et al.
Published: (2025) -
Sim-DETR: Unlock DETR for Temporal Sentence Grounding
by: Tang, Jiajin, et al.
Published: (2025) -
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
by: Peng, Yi-Xing, et al.
Published: (2025)