Visual Position Prompt for MLLM based Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Wei, Sun, Yanpeng, Gu, Qinying, Li, Zechao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
by: Li, Bin, et al.
Published: (2022)
by: Li, Bin, et al.
Published: (2022)
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
by: Tang, Wei, et al.
Published: (2026)
by: Tang, Wei, et al.
Published: (2026)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
by: Chen, Yuan, et al.
Published: (2025)
by: Chen, Yuan, et al.
Published: (2025)
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
by: Liu, Junli, et al.
Published: (2025)
by: Liu, Junli, et al.
Published: (2025)
Robust MLLM Unlearning via Visual Knowledge Distillation
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
Exploring Visual Prompting: Robustness Inheritance and Beyond
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
Visual Prompt Selection for In-Context Learning Segmentation
by: Suo, Wei, et al.
Published: (2024)
by: Suo, Wei, et al.
Published: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
by: Chi, Donghwan, et al.
Published: (2025)
by: Chi, Donghwan, et al.
Published: (2025)
Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition
by: Choi, Jae Young, et al.
Published: (2026)
by: Choi, Jae Young, et al.
Published: (2026)
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
by: Bo, Weihao, et al.
Published: (2025)
by: Bo, Weihao, et al.
Published: (2025)
Visual Instance-aware Prompt Tuning
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
by: Lee, Daeun, et al.
Published: (2026)
by: Lee, Daeun, et al.
Published: (2026)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
by: Huang, Wenyuan, et al.
Published: (2025)
by: Huang, Wenyuan, et al.
Published: (2025)
Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization
by: Luo, Yongdong, et al.
Published: (2024)
by: Luo, Yongdong, et al.
Published: (2024)
Embedded Visual Prompt Tuning
by: Zu, Wenqiang, et al.
Published: (2024)
by: Zu, Wenqiang, et al.
Published: (2024)
SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection
by: Fu, Chenhao, et al.
Published: (2026)
by: Fu, Chenhao, et al.
Published: (2026)
Probing Visual Planning in Image Editing Models
by: Zhou, Zhimu, et al.
Published: (2026)
by: Zhou, Zhimu, et al.
Published: (2026)
Prompt-based Visual Alignment for Zero-shot Policy Transfer
by: Gao, Haihan, et al.
Published: (2024)
by: Gao, Haihan, et al.
Published: (2024)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
by: Shi, Jiaqi, et al.
Published: (2026)
by: Shi, Jiaqi, et al.
Published: (2026)
Towards Understanding Visual Grounding in Visual Language Models
by: Pantazopoulos, Georgios, et al.
Published: (2025)
by: Pantazopoulos, Georgios, et al.
Published: (2025)
Reasoning Matters for 3D Visual Grounding
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
by: Yang, Xiaobo, et al.
Published: (2025)
by: Yang, Xiaobo, et al.
Published: (2025)
Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment
by: Song, Zhixue, et al.
Published: (2026)
by: Song, Zhixue, et al.
Published: (2026)
Visual Fourier Prompt Tuning
by: Zeng, Runjia, et al.
Published: (2024)
by: Zeng, Runjia, et al.
Published: (2024)
Prompt When the Animal is: Temporal Animal Behavior Grounding with Positional Recovery Training
by: Yan, Sheng, et al.
Published: (2024)
by: Yan, Sheng, et al.
Published: (2024)
VRP-SAM: SAM with Visual Reference Prompt
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
by: Wu, Yixuan, et al.
Published: (2024)
by: Wu, Yixuan, et al.
Published: (2024)
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
by: Han, Kai, et al.
Published: (2024)
by: Han, Kai, et al.
Published: (2024)
P4Q: Learning to Prompt for Quantization in Visual-language Models
by: Sun, Huixin, et al.
Published: (2024)
by: Sun, Huixin, et al.
Published: (2024)
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
by: Li, Wenqiao, et al.
Published: (2025)
by: Li, Wenqiao, et al.
Published: (2025)
The Role of Entropy in Visual Grounding: Analysis and Optimization
by: Li, Shuo, et al.
Published: (2025)
by: Li, Shuo, et al.
Published: (2025)
VGR: Visual Grounded Reasoning
by: Wang, Jiacong, et al.
Published: (2025)
by: Wang, Jiacong, et al.
Published: (2025)
CVPT: Cross Visual Prompt Tuning
by: Huang, Lingyun, et al.
Published: (2024)
by: Huang, Lingyun, et al.
Published: (2024)
Selective Visual Prompting in Vision Mamba
by: Yao, Yifeng, et al.
Published: (2024)
by: Yao, Yifeng, et al.
Published: (2024)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
DOGR: Towards Versatile Visual Document Grounding and Referring
by: Zhou, Yinan, et al.
Published: (2024)
by: Zhou, Yinan, et al.
Published: (2024)
Similar Items
-
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025) -
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
by: Xu, Beining, et al.
Published: (2025) -
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
by: Li, Bin, et al.
Published: (2022) -
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
by: Tang, Wei, et al.
Published: (2026) -
IPCV: Information-Preserving Compression for MLLM Visual Encoders
by: Chen, Yuan, et al.
Published: (2025)