AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Junyang, Zhu, Tianyi, Tambe, Thierry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
by: Guo, Chaohong, et al.
Published: (2025)
by: Guo, Chaohong, et al.
Published: (2025)
Token Sequence Compression for Efficient Multimodal Computing
by: Omri, Yasmine, et al.
Published: (2025)
by: Omri, Yasmine, et al.
Published: (2025)
AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs
by: Gutiérrez, Juan, et al.
Published: (2026)
by: Gutiérrez, Juan, et al.
Published: (2026)
AttZoom: Attention Zoom for Better Visual Features
by: DeAlcala, Daniel, et al.
Published: (2025)
by: DeAlcala, Daniel, et al.
Published: (2025)
BiomedAP: A Vision-Informed Dual-Anchor Framework with Gated Cross-Modal Fusion for Robust Medical Vision-Language Adaptation
by: Tong, Huanyang, et al.
Published: (2026)
by: Tong, Huanyang, et al.
Published: (2026)
AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models
by: Qu, Zhen, et al.
Published: (2026)
by: Qu, Zhen, et al.
Published: (2026)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
by: Zhang, Qizhe, et al.
Published: (2024)
by: Zhang, Qizhe, et al.
Published: (2024)
ResAF-Net: An Anchor-Free Attention-Based Network for Tree Detection and Agricultural Mapping in Palestine
by: Al-Qasem, Rabee
Published: (2026)
by: Al-Qasem, Rabee
Published: (2026)
Beyond Fixed Anchors: Precisely Erasing Concepts with Sibling Exclusive Counterparts
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Bi-Anchor Interpolation Solver for Accelerating Generative Modeling
by: Chen, Hongxu, et al.
Published: (2026)
by: Chen, Hongxu, et al.
Published: (2026)
AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs
by: Basappa, Aahana, et al.
Published: (2026)
by: Basappa, Aahana, et al.
Published: (2026)
GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
by: Omri, Yasmine, et al.
Published: (2025)
by: Omri, Yasmine, et al.
Published: (2025)
PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning
by: Xi, Yingjie, et al.
Published: (2025)
by: Xi, Yingjie, et al.
Published: (2025)
MAP-Diff: Multi-Anchor Guided Diffusion for Progressive 3D Whole-Body Low-Dose PET Denoising
by: Jing, Peiyuan, et al.
Published: (2026)
by: Jing, Peiyuan, et al.
Published: (2026)
ACPO: Anchor-Constrained Perceptual Optimization for Diffusion Models with No-Reference Quality Guidance
by: Yang, Yang, et al.
Published: (2026)
by: Yang, Yang, et al.
Published: (2026)
AttFC: Attention Fully-Connected Layer for Large-Scale Face Recognition with One GPU
by: Zheng, Zhuowen, et al.
Published: (2025)
by: Zheng, Zhuowen, et al.
Published: (2025)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
by: Wang, Zun, et al.
Published: (2026)
by: Wang, Zun, et al.
Published: (2026)
Language-Driven Anchors for Zero-Shot Adversarial Robustness
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
by: Wang, Zun, et al.
Published: (2025)
by: Wang, Zun, et al.
Published: (2025)
SurgAtt-Tracker: Online Surgical Attention Tracking via Temporal Proposal Reranking and Motion-Aware Refinement
by: Zhou, Rulin, et al.
Published: (2026)
by: Zhou, Rulin, et al.
Published: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
Proxy-Anchor and EVT-Driven Continual Learning Method for Generalized Category Discovery
by: Fathalizadeh, Alireza, et al.
Published: (2025)
by: Fathalizadeh, Alireza, et al.
Published: (2025)
Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
by: Nadeem, Numair, et al.
Published: (2025)
by: Nadeem, Numair, et al.
Published: (2025)
PBADet: A One-Stage Anchor-Free Approach for Part-Body Association
by: Gao, Zhongpai, et al.
Published: (2024)
by: Gao, Zhongpai, et al.
Published: (2024)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
CMHANet: A Cross-Modal Hybrid Attention Network for Point Cloud Registration
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition
by: Shen, Yaomin, et al.
Published: (2025)
by: Shen, Yaomin, et al.
Published: (2025)
Towards Lossless Ultimate Vision Token Compression for VLMs
by: Zheng, Dehua, et al.
Published: (2025)
by: Zheng, Dehua, et al.
Published: (2025)
History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions
by: Salgado, Alberto G. Rodríguez
Published: (2026)
by: Salgado, Alberto G. Rodríguez
Published: (2026)
Learning with Instance-Dependent Noisy Labels by Anchor Hallucination and Hard Sample Label Correction
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
ContextGS: Compact 3D Gaussian Splatting with Anchor Level Context Model
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
Stateful Token Reduction for Long-Video Hybrid VLMs
by: Jiang, Jindong, et al.
Published: (2026)
by: Jiang, Jindong, et al.
Published: (2026)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation
by: Zhu, Jiayin, et al.
Published: (2025)
by: Zhu, Jiayin, et al.
Published: (2025)
Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
by: Yan, Hanqi, et al.
Published: (2025)
by: Yan, Hanqi, et al.
Published: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
Similar Items
-
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
by: Guo, Chaohong, et al.
Published: (2025) -
Token Sequence Compression for Efficient Multimodal Computing
by: Omri, Yasmine, et al.
Published: (2025) -
AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation
by: Zhang, Jian, et al.
Published: (2026) -
PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs
by: Gutiérrez, Juan, et al.
Published: (2026) -
AttZoom: Attention Zoom for Better Visual Features
by: DeAlcala, Daniel, et al.
Published: (2025)