Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Guiping, Huang, Wenjian, Lan, Xiangyuan, Zhang, Jianguo, Jiang, Dongmei, Wang, Yaowei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
by: Cao, Guiping, et al.
Published: (2025)
by: Cao, Guiping, et al.
Published: (2025)
Open-Det: An Efficient Learning Framework for Open-Ended Detection
by: Cao, Guiping, et al.
Published: (2025)
by: Cao, Guiping, et al.
Published: (2025)
CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Efficient Adversarial Training via Criticality-Aware Fine-Tuning
by: Li, Wenyun, et al.
Published: (2026)
by: Li, Wenyun, et al.
Published: (2026)
EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
by: Xing, Yifei, et al.
Published: (2024)
by: Xing, Yifei, et al.
Published: (2024)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
Enhancing Knowledge Transfer in Hyperspectral Image Classification via Cross-scene Knowledge Integration
by: Huo, Lu, et al.
Published: (2025)
by: Huo, Lu, et al.
Published: (2025)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
h-calibration: Rethinking Classifier Recalibration with Probabilistic Error-Bounded Objective
by: Huang, Wenjian, et al.
Published: (2025)
by: Huang, Wenjian, et al.
Published: (2025)
Transferable Adversarial Face Attack with Text Controlled Attribute
by: Li, Wenyun, et al.
Published: (2024)
by: Li, Wenyun, et al.
Published: (2024)
Cross-Layer Feature Pyramid Transformer for Small Object Detection in Aerial Images
by: Du, Zewen, et al.
Published: (2024)
by: Du, Zewen, et al.
Published: (2024)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
Towards Visual Grounding: A Survey
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Deep Homography Estimation for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
Object Style Diffusion for Generalized Object Detection in Urban Scene
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place Recognition
by: Lu, Feng, et al.
Published: (2025)
by: Lu, Feng, et al.
Published: (2025)
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
by: Ansari, Muhammad Musab
Published: (2025)
by: Ansari, Muhammad Musab
Published: (2025)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
by: Wang, Jiabao, et al.
Published: (2023)
by: Wang, Jiabao, et al.
Published: (2023)
UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
by: Liu, Yating, et al.
Published: (2025)
by: Liu, Yating, et al.
Published: (2025)
Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
CV-Cities: Advancing Cross-View Geo-Localization in Global Cities
by: Huang, Gaoshuang, et al.
Published: (2024)
by: Huang, Gaoshuang, et al.
Published: (2024)
Cross-Level Sensor Fusion with Object Lists via Transformer for 3D Object Detection
by: Liu, Xiangzhong, et al.
Published: (2025)
by: Liu, Xiangzhong, et al.
Published: (2025)
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model
by: Nan, Zhixiong, et al.
Published: (2024)
by: Nan, Zhixiong, et al.
Published: (2024)
CollabOD: Collaborative Multi-Backbone with Cross-scale Vision for UAV Small Object Detection
by: Bai, Xuecheng, et al.
Published: (2026)
by: Bai, Xuecheng, et al.
Published: (2026)
Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
by: Lu, Yehao, et al.
Published: (2025)
by: Lu, Yehao, et al.
Published: (2025)
Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
by: Xie, Dingzhou, et al.
Published: (2025)
by: Xie, Dingzhou, et al.
Published: (2025)
Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction Detection
by: Jiang, Weibo, et al.
Published: (2024)
by: Jiang, Weibo, et al.
Published: (2024)
Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized Constraints
by: Xia, Jiahao, et al.
Published: (2025)
by: Xia, Jiahao, et al.
Published: (2025)
Small Object Detection for Birds with Swin Transformer
by: Huo, Da, et al.
Published: (2025)
by: Huo, Da, et al.
Published: (2025)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
SCTransNet: Spatial-channel Cross Transformer Network for Infrared Small Target Detection
by: Yuan, Shuai, et al.
Published: (2024)
by: Yuan, Shuai, et al.
Published: (2024)
Boosting Cross-Domain Point Classification via Distilling Relational Priors from 2D Transformers
by: Zou, Longkun, et al.
Published: (2024)
by: Zou, Longkun, et al.
Published: (2024)
ELMAR: Enhancing LiDAR Detection with 4D Radar Motion Awareness and Cross-modal Uncertainty
by: Peng, Xiangyuan, et al.
Published: (2025)
by: Peng, Xiangyuan, et al.
Published: (2025)
Multimodal Transformer Using Cross-Channel attention for Object Detection in Remote Sensing Images
by: Bahaduri, Bissmella, et al.
Published: (2023)
by: Bahaduri, Bissmella, et al.
Published: (2023)
CS-Mixer: A Cross-Scale Vision MLP Model with Spatial-Channel Mixing
by: Cui, Jonathan, et al.
Published: (2023)
by: Cui, Jonathan, et al.
Published: (2023)
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
by: Guo, Hao, et al.
Published: (2024)
by: Guo, Hao, et al.
Published: (2024)
Similar Items
-
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
by: Cao, Guiping, et al.
Published: (2025) -
Open-Det: An Efficient Learning Framework for Open-Ended Detection
by: Cao, Guiping, et al.
Published: (2025) -
CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024) -
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
by: Wang, Hao, et al.
Published: (2024) -
Efficient Adversarial Training via Criticality-Aware Fine-Tuning
by: Li, Wenyun, et al.
Published: (2026)