Saved in:
| Main Authors: | Ren, Tianhe, Chen, Yihao, Jiang, Qing, Zeng, Zhaoyang, Xiong, Yuda, Liu, Wenlong, Ma, Zhengyu, Shen, Junyi, Gao, Yuan, Jiang, Xiaoke, Chen, Xingyu, Song, Zhuheng, Zhang, Yuhong, Huang, Hongjie, Gao, Han, Liu, Shilong, Zhang, Hao, Li, Feng, Yu, Kent, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.14347 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
Detect Anything via Next Point Prediction
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
Referring to Any Person
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
by: Jiang, Qing, et al.
Published: (2024)
by: Jiang, Qing, et al.
Published: (2024)
T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
by: Jiang, Qing, et al.
Published: (2024)
by: Jiang, Qing, et al.
Published: (2024)
SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
by: Qu, Jinyuan, et al.
Published: (2025)
by: Qu, Jinyuan, et al.
Published: (2025)
HandOS: 3D Hand Reconstruction in One Stage
by: Chen, Xingyu, et al.
Published: (2024)
by: Chen, Xingyu, et al.
Published: (2024)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
by: Qu, Jinyuan, et al.
Published: (2024)
by: Qu, Jinyuan, et al.
Published: (2024)
TAPTR: Tracking Any Point with Transformers as Detection
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
by: Jiang, Dongsheng, et al.
Published: (2023)
by: Jiang, Dongsheng, et al.
Published: (2023)
Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
TAPTRv2: Attention-based Position Update Improves Tracking Any Point
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
by: Fu, Weifu, et al.
Published: (2026)
by: Fu, Weifu, et al.
Published: (2026)
DIVE: Taming DINO for Subject-Driven Video Editing
by: Huang, Yi, et al.
Published: (2024)
by: Huang, Yi, et al.
Published: (2024)
Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
by: Cao, Guiping, et al.
Published: (2025)
by: Cao, Guiping, et al.
Published: (2025)
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
by: Zeng, Haoxi, et al.
Published: (2026)
by: Zeng, Haoxi, et al.
Published: (2026)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
by: Ling, Yiran, et al.
Published: (2026)
by: Ling, Yiran, et al.
Published: (2026)
LeanGaussian: Breaking Pixel or Point Cloud Correspondence in Modeling 3D Gaussians
by: Wu, Jiamin, et al.
Published: (2024)
by: Wu, Jiamin, et al.
Published: (2024)
Economic Inequality Brings About More Inaction Over Climate Change: The Role of Perception, Discussion, and Responsibility
by: Changcheng Wang, et al.
Published: (2025)
by: Changcheng Wang, et al.
Published: (2025)
Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning
by: Ma, Yiyao, et al.
Published: (2026)
by: Ma, Yiyao, et al.
Published: (2026)
Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
by: Lu, Yehao, et al.
Published: (2025)
by: Lu, Yehao, et al.
Published: (2025)
Unified Entropy Optimization for Open-Set Test-Time Adaptation
by: Gao, Zhengqing, et al.
Published: (2024)
by: Gao, Zhengqing, et al.
Published: (2024)
Discrimination-free Insurance Pricing with Privatized Sensitive Attributes
by: Zhang, Tianhe, et al.
Published: (2025)
by: Zhang, Tianhe, et al.
Published: (2025)
Topological Scaffold‐Based Trispecific Recombinant Protein‐Drug Conjugates for Solid Tumor Eradication
by: Huiyi Jiang, et al.
Published: (2025)
by: Huiyi Jiang, et al.
Published: (2025)
TENG-BC: Unified Time-Evolving Natural Gradient for Neural PDE Solvers with General Boundary Conditions
by: Jiang, Hongjie, et al.
Published: (2026)
by: Jiang, Hongjie, et al.
Published: (2026)
Skeleton-Guided Spatial-Temporal Feature Learning for Video-Based Visible-Infrared Person Re-Identification
by: Jiang, Wenjia, et al.
Published: (2024)
by: Jiang, Wenjia, et al.
Published: (2024)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
by: Li, Aiden Yiliu, et al.
Published: (2025)
by: Li, Aiden Yiliu, et al.
Published: (2025)
Realtime Robust Shape Estimation of Deformable Linear Object
by: Zhang, Jiaming, et al.
Published: (2024)
by: Zhang, Jiaming, et al.
Published: (2024)
Mitigate the variation of energy band gap with electric field induced by quantum confinement Stark effect via a gradient quantum system for frequency‐stable laser diodes
by: Yuhong Wang, et al.
Published: (2025)
by: Yuhong Wang, et al.
Published: (2025)
Exploring Scalable Unified Modeling for General Low-Level Vision
by: Chen, Xiangyu, et al.
Published: (2025)
by: Chen, Xiangyu, et al.
Published: (2025)
Weak Dual Drazin Inverse and its Characterizations and Properties
by: Wang, Hongxing, et al.
Published: (2024)
by: Wang, Hongxing, et al.
Published: (2024)
Simplifying DINO via Coding Rate Regularization
by: Wu, Ziyang, et al.
Published: (2025)
by: Wu, Ziyang, et al.
Published: (2025)
Deploy DINO with Many-to-Many Association
by: Jiang, Haodong, et al.
Published: (2026)
by: Jiang, Haodong, et al.
Published: (2026)
Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving
by: Chen, Xuyang, et al.
Published: (2026)
by: Chen, Xuyang, et al.
Published: (2026)
Hypergraph Convolutional Network based Weakly Supervised Point Cloud Semantic Segmentation with Scene-Level Annotations
by: Lu, Zhuheng, et al.
Published: (2022)
by: Lu, Zhuheng, et al.
Published: (2022)
Fine-Grained DINO Tuning with Dual Supervision for Face Forgery Detection
by: Zhang, Tianxiang, et al.
Published: (2025)
by: Zhang, Tianxiang, et al.
Published: (2025)
Similar Items
-
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024) -
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023) -
Detect Anything via Next Point Prediction
by: Jiang, Qing, et al.
Published: (2025) -
Referring to Any Person
by: Jiang, Qing, et al.
Published: (2025) -
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
by: Jiang, Qing, et al.
Published: (2024)