Head Pursuit: Probing Attention Specialization in Multimodal Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Basile, Lorenzo, Maiorca, Valentino, Doimo, Diego, Locatello, Francesco, Cazzaniga, Alberto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ResiDual Transformer Alignment with Spectral Decomposition
von: Basile, Lorenzo, et al.
Veröffentlicht: (2024)
von: Basile, Lorenzo, et al.
Veröffentlicht: (2024)
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
von: Serra, Alessandro Pietro, et al.
Veröffentlicht: (2024)
von: Serra, Alessandro Pietro, et al.
Veröffentlicht: (2024)
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
von: Ortu, Francesco, et al.
Veröffentlicht: (2025)
von: Ortu, Francesco, et al.
Veröffentlicht: (2025)
The representation landscape of few-shot learning and fine-tuning in large language models
von: Doimo, Diego, et al.
Veröffentlicht: (2024)
von: Doimo, Diego, et al.
Veröffentlicht: (2024)
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
Contextual Knowledge Pursuit for Faithful Visual Synthesis
von: Luo, Jinqi, et al.
Veröffentlicht: (2023)
von: Luo, Jinqi, et al.
Veröffentlicht: (2023)
Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representations
von: Basile, Lorenzo, et al.
Veröffentlicht: (2024)
von: Basile, Lorenzo, et al.
Veröffentlicht: (2024)
Transformer with Controlled Attention for Synchronous Motion Captioning
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
Out-of-Distribution Detection with Relative Angles
von: Demirel, Berker, et al.
Veröffentlicht: (2024)
von: Demirel, Berker, et al.
Veröffentlicht: (2024)
CAT: Circular-Convolutional Attention for Sub-Quadratic Transformers
von: Yamada, Yoshihiro
Veröffentlicht: (2025)
von: Yamada, Yoshihiro
Veröffentlicht: (2025)
Text Role Classification in Scientific Charts Using Multimodal Transformers
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
Adaptive Slot Attention: Object Discovery with Dynamic Slot Number
von: Fan, Ke, et al.
Veröffentlicht: (2024)
von: Fan, Ke, et al.
Veröffentlicht: (2024)
Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
von: Lao, Dong, et al.
Veröffentlicht: (2023)
von: Lao, Dong, et al.
Veröffentlicht: (2023)
Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
von: Gabetni, Firas, et al.
Veröffentlicht: (2025)
von: Gabetni, Firas, et al.
Veröffentlicht: (2025)
Unifying Specialized Visual Encoders for Video Language Models
von: Chung, Jihoon, et al.
Veröffentlicht: (2025)
von: Chung, Jihoon, et al.
Veröffentlicht: (2025)
Reinforced Attention Learning
von: Li, Bangzheng, et al.
Veröffentlicht: (2026)
von: Li, Bangzheng, et al.
Veröffentlicht: (2026)
Match & Choose: Model Selection Framework for Fine-tuning Text-to-Image Diffusion Models
von: Lewandowski, Basile, et al.
Veröffentlicht: (2025)
von: Lewandowski, Basile, et al.
Veröffentlicht: (2025)
AiSciVision: A Framework for Specializing Large Multimodal Models in Scientific Image Classification
von: Hogan, Brendan, et al.
Veröffentlicht: (2024)
von: Hogan, Brendan, et al.
Veröffentlicht: (2024)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
Listen Then See: Video Alignment with Speaker Attention
von: Agrawal, Aviral, et al.
Veröffentlicht: (2024)
von: Agrawal, Aviral, et al.
Veröffentlicht: (2024)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
von: Fan, Xiaoran, et al.
Veröffentlicht: (2026)
von: Fan, Xiaoran, et al.
Veröffentlicht: (2026)
Context-Aware Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
von: Achtibat, Reduan, et al.
Veröffentlicht: (2024)
von: Achtibat, Reduan, et al.
Veröffentlicht: (2024)
Transformer-VQ: Linear-Time Transformers via Vector Quantization
von: Lingle, Lucas D.
Veröffentlicht: (2023)
von: Lingle, Lucas D.
Veröffentlicht: (2023)
General Transform: A Unified Framework for Adaptive Transform to Enhance Representations
von: Budiutama, Gekko, et al.
Veröffentlicht: (2025)
von: Budiutama, Gekko, et al.
Veröffentlicht: (2025)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
von: Luo, Jiayun, et al.
Veröffentlicht: (2024)
von: Luo, Jiayun, et al.
Veröffentlicht: (2024)
Hyperbolic Multimodal Representation Learning for Biological Taxonomies
von: Gong, ZeMing, et al.
Veröffentlicht: (2025)
von: Gong, ZeMing, et al.
Veröffentlicht: (2025)
Aya Vision: Advancing the Frontier of Multilingual Multimodality
von: Dash, Saurabh, et al.
Veröffentlicht: (2025)
von: Dash, Saurabh, et al.
Veröffentlicht: (2025)
A Practitioner's Guide to Continual Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
DreamLLM: Synergistic Multimodal Comprehension and Creation
von: Dong, Runpei, et al.
Veröffentlicht: (2023)
von: Dong, Runpei, et al.
Veröffentlicht: (2023)
A Survey on Transformer Compression
von: Tang, Yehui, et al.
Veröffentlicht: (2024)
von: Tang, Yehui, et al.
Veröffentlicht: (2024)
Intern-S1: A Scientific Multimodal Foundation Model
von: Bai, Lei, et al.
Veröffentlicht: (2025)
von: Bai, Lei, et al.
Veröffentlicht: (2025)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
Multimodal Arabic Captioning with Interpretable Visual Concept Integration
von: Elchafei, Passant, et al.
Veröffentlicht: (2025)
von: Elchafei, Passant, et al.
Veröffentlicht: (2025)
CROME: Cross-Modal Adapters for Efficient Multimodal LLM
von: Ebrahimi, Sayna, et al.
Veröffentlicht: (2024)
von: Ebrahimi, Sayna, et al.
Veröffentlicht: (2024)
Multimodal Language Models Cannot Spot Spatial Inconsistencies
von: Khangaonkar, Om, et al.
Veröffentlicht: (2026)
von: Khangaonkar, Om, et al.
Veröffentlicht: (2026)
How to Merge Your Multimodal Models Over Time?
von: Dziadzio, Sebastian, et al.
Veröffentlicht: (2024)
von: Dziadzio, Sebastian, et al.
Veröffentlicht: (2024)
LaVy: Vietnamese Multimodal Large Language Model
von: Tran, Chi, et al.
Veröffentlicht: (2024)
von: Tran, Chi, et al.
Veröffentlicht: (2024)
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Multimodal Latent Language Modeling with Next-Token Diffusion
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ResiDual Transformer Alignment with Spectral Decomposition
von: Basile, Lorenzo, et al.
Veröffentlicht: (2024) -
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
von: Serra, Alessandro Pietro, et al.
Veröffentlicht: (2024) -
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
von: Ortu, Francesco, et al.
Veröffentlicht: (2025) -
The representation landscape of few-shot learning and fine-tuning in large language models
von: Doimo, Diego, et al.
Veröffentlicht: (2024) -
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)