An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xiangyu, Chen, Yicheng, Xu, Shilin, Li, Xiangtai, Wang, Xinjiang, Li, Yining, Huang, Haian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023)
by: Xu, Shilin, et al.
Published: (2023)
Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language
by: Chen, Yicheng, et al.
Published: (2024)
by: Chen, Yicheng, et al.
Published: (2024)
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
by: Yuan, Haobo, et al.
Published: (2024)
by: Yuan, Haobo, et al.
Published: (2024)
A Comprehensive Review of 3D Object Detection in Autonomous Driving: Technological Advances and Future Directions
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
by: Li, Jingliang, et al.
Published: (2026)
by: Li, Jingliang, et al.
Published: (2026)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
PanopticPartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation
by: Li, Xiangtai, et al.
Published: (2023)
by: Li, Xiangtai, et al.
Published: (2023)
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
by: Liu, Shuming, et al.
Published: (2025)
by: Liu, Shuming, et al.
Published: (2025)
OpenIllumination: A Multi-Illumination Dataset for Inverse Rendering Evaluation on Real Objects
by: Liu, Isabella, et al.
Published: (2023)
by: Liu, Isabella, et al.
Published: (2023)
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
A Unified Detection Pipeline for Robust Object Detection in Fisheye-Based Traffic Surveillance
by: Owor, Neema Jakisa, et al.
Published: (2025)
by: Owor, Neema Jakisa, et al.
Published: (2025)
RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation
by: Lu, Peng, et al.
Published: (2023)
by: Lu, Peng, et al.
Published: (2023)
Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection
by: Gao, Xiangyu, et al.
Published: (2025)
by: Gao, Xiangyu, et al.
Published: (2025)
Learning to Detect and Segment for Open Vocabulary Object Detection
by: Wang, Tao, et al.
Published: (2022)
by: Wang, Tao, et al.
Published: (2022)
Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning
by: Wang, Xinshun, et al.
Published: (2023)
by: Wang, Xinshun, et al.
Published: (2023)
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
by: Wei, Guoting, et al.
Published: (2026)
by: Wei, Guoting, et al.
Published: (2026)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
by: Xu, Shilin, et al.
Published: (2024)
by: Xu, Shilin, et al.
Published: (2024)
Towards Unified 3D Object Detection via Algorithm and Data Unification
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
Open World Object Detection: A Survey
by: Li, Yiming, et al.
Published: (2024)
by: Li, Yiming, et al.
Published: (2024)
Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
by: Huang, Yuzhi, et al.
Published: (2025)
by: Huang, Yuzhi, et al.
Published: (2025)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
by: Li, Xiangtai, et al.
Published: (2025)
by: Li, Xiangtai, et al.
Published: (2025)
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
by: Xu, Shilin, et al.
Published: (2025)
by: Xu, Shilin, et al.
Published: (2025)
Open-Vocabulary Object Detection with Meta Prompt Representation and Instance Contrastive Optimization
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
by: Wu, Yuchen, et al.
Published: (2026)
by: Wu, Yuchen, et al.
Published: (2026)
OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning
by: Hong, Lingyi, et al.
Published: (2024)
by: Hong, Lingyi, et al.
Published: (2024)
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
by: Xu, Shilin, et al.
Published: (2024)
by: Xu, Shilin, et al.
Published: (2024)
An Attribute-Enriched Dataset and Auto-Annotated Pipeline for Open Detection
by: Qi, Pengfei, et al.
Published: (2024)
by: Qi, Pengfei, et al.
Published: (2024)
4th PVUW MeViS 3rd Place Report: Sa2VA
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
Visible-Thermal Tiny Object Detection: A Benchmark Dataset and Baselines
by: Ying, Xinyi, et al.
Published: (2024)
by: Ying, Xinyi, et al.
Published: (2024)
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
Phrase Grounding-based Style Transfer for Single-Domain Generalized Object Detection
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Similar Items
-
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
by: Zhao, Xiangyu, et al.
Published: (2024) -
DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection
by: Xu, Shilin, et al.
Published: (2023) -
Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language
by: Chen, Yicheng, et al.
Published: (2024) -
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively
by: Yuan, Haobo, et al.
Published: (2024) -
A Comprehensive Review of 3D Object Detection in Autonomous Driving: Technological Advances and Future Directions
by: Wang, Yu, et al.
Published: (2024)