Towards Visual Grounding: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Linhui, Yang, Xiaoshan, Lan, Xiangyuan, Wang, Yaowei, Xu, Changsheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2023)
by: Xiao, Linhui, et al.
Published: (2023)
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
Pilot: Building the Federated Multimodal Instruction Tuning Framework
by: Xiong, Baochen, et al.
Published: (2025)
by: Xiong, Baochen, et al.
Published: (2025)
BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding
by: Li, Hongbing, et al.
Published: (2026)
by: Li, Hongbing, et al.
Published: (2026)
CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
by: Xu, Xiaoran, et al.
Published: (2026)
by: Xu, Xiaoran, et al.
Published: (2026)
SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place Recognition
by: Lu, Feng, et al.
Published: (2025)
by: Lu, Feng, et al.
Published: (2025)
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
by: Tao, Ming, et al.
Published: (2024)
by: Tao, Ming, et al.
Published: (2024)
A Comprehensive Review of Few-shot Action Recognition
by: Wanyan, Yuyang, et al.
Published: (2024)
by: Wanyan, Yuyang, et al.
Published: (2024)
Libra: Building Decoupled Vision System on Large Language Models
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
RGBT-Ground Benchmark: Visual Grounding Beyond RGB in Complex Real-World Scenarios
by: Zhao, Tianyi, et al.
Published: (2025)
by: Zhao, Tianyi, et al.
Published: (2025)
UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
by: Liu, Yating, et al.
Published: (2025)
by: Liu, Yating, et al.
Published: (2025)
Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
by: Cao, Guiping, et al.
Published: (2025)
by: Cao, Guiping, et al.
Published: (2025)
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
by: Cao, Guiping, et al.
Published: (2025)
by: Cao, Guiping, et al.
Published: (2025)
Efficient Adversarial Training via Criticality-Aware Fine-Tuning
by: Li, Wenyun, et al.
Published: (2026)
by: Li, Wenyun, et al.
Published: (2026)
Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC
by: Tao, Ming, et al.
Published: (2024)
by: Tao, Ming, et al.
Published: (2024)
Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
by: Wanyan, Yuyang, et al.
Published: (2025)
by: Wanyan, Yuyang, et al.
Published: (2025)
DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval
by: Liu, Yating, et al.
Published: (2025)
by: Liu, Yating, et al.
Published: (2025)
EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
by: Xing, Yifei, et al.
Published: (2024)
by: Xing, Yifei, et al.
Published: (2024)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer Era
by: Lu, Feng, et al.
Published: (2025)
by: Lu, Feng, et al.
Published: (2025)
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
by: Qiang, Chenhui, et al.
Published: (2025)
by: Qiang, Chenhui, et al.
Published: (2025)
TCP:Textual-based Class-aware Prompt tuning for Visual-Language Model
by: Yao, Hantao, et al.
Published: (2023)
by: Yao, Hantao, et al.
Published: (2023)
MixBCT: Towards Self-Adapting Backward-Compatible Training
by: Liang, Yu, et al.
Published: (2023)
by: Liang, Yu, et al.
Published: (2023)
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
by: Wang, Zhaozhi, et al.
Published: (2025)
by: Wang, Zhaozhi, et al.
Published: (2025)
ClickTrack: Towards Real-time Interactive Single Object Tracking
by: Wang, Kuiran, et al.
Published: (2024)
by: Wang, Kuiran, et al.
Published: (2024)
GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
by: Fan, Rong, et al.
Published: (2026)
by: Fan, Rong, et al.
Published: (2026)
Visual Mamba: A Survey and New Outlooks
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
Hierarchical Augmentation and Distillation for Class Incremental Audio-Visual Video Recognition
by: Zuo, Yukun, et al.
Published: (2024)
by: Zuo, Yukun, et al.
Published: (2024)
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
by: Liu, Daizong, et al.
Published: (2024)
by: Liu, Daizong, et al.
Published: (2024)
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging
by: Shao, Zibo, et al.
Published: (2026)
by: Shao, Zibo, et al.
Published: (2026)
UGround: Towards Unified Visual Grounding with Unrolled Transformers
by: Qian, Rui, et al.
Published: (2025)
by: Qian, Rui, et al.
Published: (2025)
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
by: Xu, Yunqiu, et al.
Published: (2024)
by: Xu, Yunqiu, et al.
Published: (2024)
Similar Items
-
HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2024) -
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2023) -
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
by: Xiao, Linhui, et al.
Published: (2024) -
Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024) -
Pilot: Building the Federated Multimodal Instruction Tuning Framework
by: Xiong, Baochen, et al.
Published: (2025)