Visual Position Prompt for MLLM based Visual Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Wei, Sun, Yanpeng, Gu, Qinying, Li, Zechao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
von: Ding, Yihao, et al.
Veröffentlicht: (2025)
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
von: Xu, Beining, et al.
Veröffentlicht: (2025)
von: Xu, Beining, et al.
Veröffentlicht: (2025)
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
von: Li, Bin, et al.
Veröffentlicht: (2022)
von: Li, Bin, et al.
Veröffentlicht: (2022)
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
von: Tang, Wei, et al.
Veröffentlicht: (2026)
von: Tang, Wei, et al.
Veröffentlicht: (2026)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
von: Liu, Junli, et al.
Veröffentlicht: (2025)
von: Liu, Junli, et al.
Veröffentlicht: (2025)
Robust MLLM Unlearning via Visual Knowledge Distillation
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
Exploring Visual Prompting: Robustness Inheritance and Beyond
von: Li, Qi, et al.
Veröffentlicht: (2025)
von: Li, Qi, et al.
Veröffentlicht: (2025)
Visual Prompt Selection for In-Context Learning Segmentation
von: Suo, Wei, et al.
Veröffentlicht: (2024)
von: Suo, Wei, et al.
Veröffentlicht: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
von: Chi, Donghwan, et al.
Veröffentlicht: (2025)
Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition
von: Choi, Jae Young, et al.
Veröffentlicht: (2026)
von: Choi, Jae Young, et al.
Veröffentlicht: (2026)
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
von: Bo, Weihao, et al.
Veröffentlicht: (2025)
von: Bo, Weihao, et al.
Veröffentlicht: (2025)
Visual Instance-aware Prompt Tuning
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
von: Lee, Daeun, et al.
Veröffentlicht: (2026)
von: Lee, Daeun, et al.
Veröffentlicht: (2026)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
von: Huang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenyuan, et al.
Veröffentlicht: (2025)
Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
Embedded Visual Prompt Tuning
von: Zu, Wenqiang, et al.
Veröffentlicht: (2024)
von: Zu, Wenqiang, et al.
Veröffentlicht: (2024)
SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection
von: Fu, Chenhao, et al.
Veröffentlicht: (2026)
von: Fu, Chenhao, et al.
Veröffentlicht: (2026)
Probing Visual Planning in Image Editing Models
von: Zhou, Zhimu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhimu, et al.
Veröffentlicht: (2026)
Prompt-based Visual Alignment for Zero-shot Policy Transfer
von: Gao, Haihan, et al.
Veröffentlicht: (2024)
von: Gao, Haihan, et al.
Veröffentlicht: (2024)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
Towards Understanding Visual Grounding in Visual Language Models
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2025)
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2025)
Reasoning Matters for 3D Visual Grounding
von: Huang, Hsiang-Wei, et al.
Veröffentlicht: (2026)
von: Huang, Hsiang-Wei, et al.
Veröffentlicht: (2026)
Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
von: Yang, Xiaobo, et al.
Veröffentlicht: (2025)
von: Yang, Xiaobo, et al.
Veröffentlicht: (2025)
Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment
von: Song, Zhixue, et al.
Veröffentlicht: (2026)
von: Song, Zhixue, et al.
Veröffentlicht: (2026)
Visual Fourier Prompt Tuning
von: Zeng, Runjia, et al.
Veröffentlicht: (2024)
von: Zeng, Runjia, et al.
Veröffentlicht: (2024)
Prompt When the Animal is: Temporal Animal Behavior Grounding with Positional Recovery Training
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
VRP-SAM: SAM with Visual Reference Prompt
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
von: Wu, Yixuan, et al.
Veröffentlicht: (2024)
von: Wu, Yixuan, et al.
Veröffentlicht: (2024)
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
von: Han, Kai, et al.
Veröffentlicht: (2024)
von: Han, Kai, et al.
Veröffentlicht: (2024)
P4Q: Learning to Prompt for Quantization in Visual-language Models
von: Sun, Huixin, et al.
Veröffentlicht: (2024)
von: Sun, Huixin, et al.
Veröffentlicht: (2024)
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
von: Li, Wenqiao, et al.
Veröffentlicht: (2025)
von: Li, Wenqiao, et al.
Veröffentlicht: (2025)
The Role of Entropy in Visual Grounding: Analysis and Optimization
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
VGR: Visual Grounded Reasoning
von: Wang, Jiacong, et al.
Veröffentlicht: (2025)
von: Wang, Jiacong, et al.
Veröffentlicht: (2025)
CVPT: Cross Visual Prompt Tuning
von: Huang, Lingyun, et al.
Veröffentlicht: (2024)
von: Huang, Lingyun, et al.
Veröffentlicht: (2024)
Selective Visual Prompting in Vision Mamba
von: Yao, Yifeng, et al.
Veröffentlicht: (2024)
von: Yao, Yifeng, et al.
Veröffentlicht: (2024)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
DOGR: Towards Versatile Visual Document Grounding and Referring
von: Zhou, Yinan, et al.
Veröffentlicht: (2024)
von: Zhou, Yinan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
von: Ding, Yihao, et al.
Veröffentlicht: (2025) -
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
von: Xu, Beining, et al.
Veröffentlicht: (2025) -
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
von: Li, Bin, et al.
Veröffentlicht: (2022) -
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
von: Tang, Wei, et al.
Veröffentlicht: (2026) -
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025)