Direct Visual Grounding by Directing Attention of Visual Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Esmaeilkhani, Parsa, Latecki, Longin Jan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preserving Localized Patch Semantics in VLMs
by: Esmaeilkhani, Parsa, et al.
Published: (2026)
by: Esmaeilkhani, Parsa, et al.
Published: (2026)
Learning Object Focused Attention
by: Trivedy, Vivek, et al.
Published: (2025)
by: Trivedy, Vivek, et al.
Published: (2025)
Layout Stroke Imitation: A Layout Guided Handwriting Stroke Generation for Style Imitation with Diffusion Model
by: Hanif, Sidra, et al.
Published: (2025)
by: Hanif, Sidra, et al.
Published: (2025)
VAPO: Visibility-Aware Keypoint Localization for Efficient 6DoF Object Pose Estimation
by: Lian, Ruyi, et al.
Published: (2024)
by: Lian, Ruyi, et al.
Published: (2024)
Enhancing Renal Tumor Malignancy Prediction: Deep Learning with Automatic 3D CT Organ Focused Attention
by: Fan, Zhengkang, et al.
Published: (2026)
by: Fan, Zhengkang, et al.
Published: (2026)
FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding
by: Pan, Huitong, et al.
Published: (2024)
by: Pan, Huitong, et al.
Published: (2024)
Visual Grounding with Attention-Driven Constraint Balancing
by: Kang, Weitai, et al.
Published: (2024)
by: Kang, Weitai, et al.
Published: (2024)
Token Coordinated Prompt Attention is Needed for Visual Prompting
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
by: Narnaware, Vishal, et al.
Published: (2026)
by: Narnaware, Vishal, et al.
Published: (2026)
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
by: Liu, Daizong, et al.
Published: (2024)
by: Liu, Daizong, et al.
Published: (2024)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
by: Fan, Rong, et al.
Published: (2026)
by: Fan, Rong, et al.
Published: (2026)
GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning
by: Miao, Deshui, et al.
Published: (2026)
by: Miao, Deshui, et al.
Published: (2026)
TDANet: Target-Directed Attention Network For Object-Goal Visual Navigation With Zero-Shot Ability
by: Lian, Shiwei, et al.
Published: (2024)
by: Lian, Shiwei, et al.
Published: (2024)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
by: Li, Haotang, et al.
Published: (2026)
by: Li, Haotang, et al.
Published: (2026)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
by: Zhang, Qizhe, et al.
Published: (2024)
by: Zhang, Qizhe, et al.
Published: (2024)
Attention Grounded Enhancement for Visual Document Retrieval
by: Cui, Wanqing, et al.
Published: (2025)
by: Cui, Wanqing, et al.
Published: (2025)
Factorized Visual Tokenization and Generation
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
On the Role of Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2024)
by: Reich, Daniel, et al.
Published: (2024)
Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification
by: Jin, Xin, et al.
Published: (2026)
by: Jin, Xin, et al.
Published: (2026)
Visual Gyroscope: Combination of Deep Learning Features and Direct Alignment for Panoramic Stabilization
by: Berenguel-Baeta, Bruno, et al.
Published: (2024)
by: Berenguel-Baeta, Bruno, et al.
Published: (2024)
ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding
by: Peng, Qihang, et al.
Published: (2025)
by: Peng, Qihang, et al.
Published: (2025)
Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning
by: Chien, Tzu-Chun, et al.
Published: (2025)
by: Chien, Tzu-Chun, et al.
Published: (2025)
Directed-CP: Directed Collaborative Perception for Connected and Autonomous Vehicles via Proactive Attention
by: Tao, Yihang, et al.
Published: (2024)
by: Tao, Yihang, et al.
Published: (2024)
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
by: Lou, Haoran, et al.
Published: (2025)
by: Lou, Haoran, et al.
Published: (2025)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
by: Hu, Junshan, et al.
Published: (2025)
by: Hu, Junshan, et al.
Published: (2025)
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
by: Luo, Jiayun, et al.
Published: (2024)
by: Luo, Jiayun, et al.
Published: (2024)
ESVO2: Direct Visual-Inertial Odometry with Stereo Event Cameras
by: Niu, Junkai, et al.
Published: (2024)
by: Niu, Junkai, et al.
Published: (2024)
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
by: Hu, Zijing, et al.
Published: (2025)
by: Hu, Zijing, et al.
Published: (2025)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
Grounded Reinforcement Learning for Visual Reasoning
by: Sarch, Gabriel, et al.
Published: (2025)
by: Sarch, Gabriel, et al.
Published: (2025)
Visual Intention Grounding for Egocentric Assistants
by: Sun, Pengzhan, et al.
Published: (2025)
by: Sun, Pengzhan, et al.
Published: (2025)
Quantized Visual Geometry Grounded Transformer
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Towards Visual Grounding: A Survey
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
Context-Infused Visual Grounding for Art
by: Khan, Selina, et al.
Published: (2024)
by: Khan, Selina, et al.
Published: (2024)
Similar Items
-
Preserving Localized Patch Semantics in VLMs
by: Esmaeilkhani, Parsa, et al.
Published: (2026) -
Learning Object Focused Attention
by: Trivedy, Vivek, et al.
Published: (2025) -
Layout Stroke Imitation: A Layout Guided Handwriting Stroke Generation for Style Imitation with Diffusion Model
by: Hanif, Sidra, et al.
Published: (2025) -
VAPO: Visibility-Aware Keypoint Localization for Efficient 6DoF Object Pose Estimation
by: Lian, Ruyi, et al.
Published: (2024) -
Enhancing Renal Tumor Malignancy Prediction: Deep Learning with Automatic 3D CT Organ Focused Attention
by: Fan, Zhengkang, et al.
Published: (2026)