MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Weihao, Wu, Ning, Yang, Shiping, Ding, Wenbiao, Liang, Shining, Gong, Ming, Zhang, Dongmei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
von: Liu, Weihao, et al.
Veröffentlicht: (2024)
von: Liu, Weihao, et al.
Veröffentlicht: (2024)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
von: Yang, Shiping, et al.
Veröffentlicht: (2025)
von: Yang, Shiping, et al.
Veröffentlicht: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
von: Lu, Yi, et al.
Veröffentlicht: (2024)
von: Lu, Yi, et al.
Veröffentlicht: (2024)
Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
von: Chen, Nuo, et al.
Veröffentlicht: (2023)
von: Chen, Nuo, et al.
Veröffentlicht: (2023)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
von: Donhauser, Konstantin, et al.
Veröffentlicht: (2025)
von: Donhauser, Konstantin, et al.
Veröffentlicht: (2025)
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
von: Lin, Xihui, et al.
Veröffentlicht: (2024)
von: Lin, Xihui, et al.
Veröffentlicht: (2024)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
von: Guo, Yiju, et al.
Veröffentlicht: (2025)
von: Guo, Yiju, et al.
Veröffentlicht: (2025)
Lag-Relative Sparse Attention In Long Context Training
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
von: Yin, Shangjian, et al.
Veröffentlicht: (2025)
von: Yin, Shangjian, et al.
Veröffentlicht: (2025)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
Attention Mechanism and Heuristic Approach: Context-Aware File Ranking Using Multi-Head Self-Attention
von: Sharma, Pradeep Kumar, et al.
Veröffentlicht: (2026)
von: Sharma, Pradeep Kumar, et al.
Veröffentlicht: (2026)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
von: Kahardipraja, Patrick, et al.
Veröffentlicht: (2025)
von: Kahardipraja, Patrick, et al.
Veröffentlicht: (2025)
Accurate, fast, cheap: Choose three. Replacing Multi-Head-Attention with Bidirectional Recurrent Attention for Long-Form ASR
von: Ratajczak, Martin, et al.
Veröffentlicht: (2025)
von: Ratajczak, Martin, et al.
Veröffentlicht: (2025)
Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts
von: Zhu, Youxiang, et al.
Veröffentlicht: (2025)
von: Zhu, Youxiang, et al.
Veröffentlicht: (2025)
Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
Training-free Context-adaptive Attention for Efficient Long Context Modeling
von: You, Zeng, et al.
Veröffentlicht: (2025)
von: You, Zeng, et al.
Veröffentlicht: (2025)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
Long-Context Generalization with Sparse Attention
von: Vasylenko, Pavlo, et al.
Veröffentlicht: (2025)
von: Vasylenko, Pavlo, et al.
Veröffentlicht: (2025)
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
Long Context Pre-Training with Lighthouse Attention
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
von: Zhang, Wuwei, et al.
Veröffentlicht: (2025)
von: Zhang, Wuwei, et al.
Veröffentlicht: (2025)
Knocking-Heads Attention
von: Zhou, Zhanchao, et al.
Veröffentlicht: (2025)
von: Zhou, Zhanchao, et al.
Veröffentlicht: (2025)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
von: Ai, Xuan, et al.
Veröffentlicht: (2026)
von: Ai, Xuan, et al.
Veröffentlicht: (2026)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
von: Liu, Siran, et al.
Veröffentlicht: (2026)
von: Liu, Siran, et al.
Veröffentlicht: (2026)
MuLD: The Multitask Long Document Benchmark
von: Hudson, G Thomas, et al.
Veröffentlicht: (2022)
von: Hudson, G Thomas, et al.
Veröffentlicht: (2022)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
von: Shyam, Vasudev, et al.
Veröffentlicht: (2024)
von: Shyam, Vasudev, et al.
Veröffentlicht: (2024)
Multipole Attention for Efficient Long Context Reasoning
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
von: Hua, Kai, et al.
Veröffentlicht: (2025)
von: Hua, Kai, et al.
Veröffentlicht: (2025)
Reducing Distraction in Long-Context Language Models by Focused Learning
von: Wu, Zijun, et al.
Veröffentlicht: (2024)
von: Wu, Zijun, et al.
Veröffentlicht: (2024)
Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
von: Long, Lingkun, et al.
Veröffentlicht: (2026)
von: Long, Lingkun, et al.
Veröffentlicht: (2026)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
von: Fu, Zichuan, et al.
Veröffentlicht: (2026)
von: Fu, Zichuan, et al.
Veröffentlicht: (2026)
Efficient OpAmp Adaptation for Zoom Attention to Golden Contexts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
Analyzing Multi-Head Attention on Trojan BERT Models
von: Wang, Jingwei
Veröffentlicht: (2024)
von: Wang, Jingwei
Veröffentlicht: (2024)
Latent Multi-Head Attention for Small Language Models
von: Mehta, Sushant, et al.
Veröffentlicht: (2025)
von: Mehta, Sushant, et al.
Veröffentlicht: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
von: Xiao, Da, et al.
Veröffentlicht: (2024)
von: Xiao, Da, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
von: Liu, Weihao, et al.
Veröffentlicht: (2024) -
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
von: Yang, Shiping, et al.
Veröffentlicht: (2025) -
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
von: Lu, Yi, et al.
Veröffentlicht: (2024) -
Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
von: Chen, Nuo, et al.
Veröffentlicht: (2023) -
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
von: Donhauser, Konstantin, et al.
Veröffentlicht: (2025)