Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Wuwei, Yin, Fangcong, Yen, Howard, Chen, Danqi, Ye, Xi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models
by: Ye, Xi, et al.
Published: (2026)
by: Ye, Xi, et al.
Published: (2026)
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
by: Ye, Xi, et al.
Published: (2025)
by: Ye, Xi, et al.
Published: (2025)
Understanding Synthetic Context Extension via Retrieval Heads
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
Long-Context Language Modeling with Parallel Context Encoding
by: Yen, Howard, et al.
Published: (2024)
by: Yen, Howard, et al.
Published: (2024)
Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
by: Lee, Yoonsang, et al.
Published: (2026)
by: Lee, Yoonsang, et al.
Published: (2026)
How to Train Long-Context Language Models (Effectively)
by: Gao, Tianyu, et al.
Published: (2024)
by: Gao, Tianyu, et al.
Published: (2024)
LoFiT: Localized Fine-tuning on LLM Representations
by: Yin, Fangcong, et al.
Published: (2024)
by: Yin, Fangcong, et al.
Published: (2024)
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
by: Yen, Howard, et al.
Published: (2025)
by: Yen, Howard, et al.
Published: (2025)
DQ-LoRe: Dual Queries with Low Rank Approximation Re-ranking for In-Context Learning
by: Xiong, Jing, et al.
Published: (2023)
by: Xiong, Jing, et al.
Published: (2023)
HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
by: Yen, Howard, et al.
Published: (2024)
by: Yen, Howard, et al.
Published: (2024)
Grounding Long-Context Reasoning with Contextual Normalization for Retrieval-Augmented Generation
by: Chen, Jiamin, et al.
Published: (2025)
by: Chen, Jiamin, et al.
Published: (2025)
Retrieval, Reasoning, Re-ranking: A Context-Enriched Framework for Knowledge Graph Completion
by: Li, Muzhi, et al.
Published: (2024)
by: Li, Muzhi, et al.
Published: (2024)
Language Models that Think, Chat Better
by: Bhaskar, Adithya, et al.
Published: (2025)
by: Bhaskar, Adithya, et al.
Published: (2025)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
by: Liu, Zhuorui, et al.
Published: (2025)
by: Liu, Zhuorui, et al.
Published: (2025)
Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models
by: Tian, Yuxing, et al.
Published: (2026)
by: Tian, Yuxing, et al.
Published: (2026)
Retrieval Head Mechanistically Explains Long-Context Factuality
by: Wu, Wenhao, et al.
Published: (2024)
by: Wu, Wenhao, et al.
Published: (2024)
Unstructured Evidence Attribution for Long Context Query Focused Summarization
by: Wright, Dustin, et al.
Published: (2025)
by: Wright, Dustin, et al.
Published: (2025)
Learning Composable Chains-of-Thought
by: Yin, Fangcong, et al.
Published: (2025)
by: Yin, Fangcong, et al.
Published: (2025)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
by: Liu, Weihao, et al.
Published: (2025)
by: Liu, Weihao, et al.
Published: (2025)
MRAG-Suite: A Diagnostic Evaluation Platform for Visual Retrieval-Augmented Generation
by: Ji, Yuelyu, et al.
Published: (2025)
by: Ji, Yuelyu, et al.
Published: (2025)
Improving Retrieval in Sponsored Search by Leveraging Query Context Signals
by: Mohankumar, Akash Kumar, et al.
Published: (2024)
by: Mohankumar, Akash Kumar, et al.
Published: (2024)
Evaluating Multilingual Long-Context Models for Retrieval and Reasoning
by: Agrawal, Ameeta, et al.
Published: (2024)
by: Agrawal, Ameeta, et al.
Published: (2024)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
by: Ma, Youmi, et al.
Published: (2026)
by: Ma, Youmi, et al.
Published: (2026)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
Reasoning-Aware Query-Focused Summarization over Multi-Table Data
by: Lin, Xiaochuan, et al.
Published: (2024)
by: Lin, Xiaochuan, et al.
Published: (2024)
QUITO: Accelerating Long-Context Reasoning through Query-Guided Context Compression
by: Wang, Wenshan, et al.
Published: (2024)
by: Wang, Wenshan, et al.
Published: (2024)
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
by: Bhaskar, Adithya, et al.
Published: (2025)
by: Bhaskar, Adithya, et al.
Published: (2025)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
by: Xiao, Qingfa, et al.
Published: (2025)
by: Xiao, Qingfa, et al.
Published: (2025)
A Reasoning-Focused Legal Retrieval Benchmark
by: Zheng, Lucia, et al.
Published: (2025)
by: Zheng, Lucia, et al.
Published: (2025)
Precise Information Control in Long-Form Text Generation
by: He, Jacqueline, et al.
Published: (2025)
by: He, Jacqueline, et al.
Published: (2025)
VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding
by: Pei, Rongcan, et al.
Published: (2026)
by: Pei, Rongcan, et al.
Published: (2026)
Attendre: Wait To Attend By Retrieval With Evicted Queries in Memory-Based Transformers for Long Context Processing
by: Yang, Zi, et al.
Published: (2024)
by: Yang, Zi, et al.
Published: (2024)
Query Suggestion for Retrieval-Augmented Generation via Dynamic In-Context Learning
by: Spaeh, Fabian, et al.
Published: (2026)
by: Spaeh, Fabian, et al.
Published: (2026)
ImpRAG: Retrieval-Augmented Generation with Implicit Queries
by: Zhang, Wenzheng, et al.
Published: (2025)
by: Zhang, Wenzheng, et al.
Published: (2025)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
by: Wang, Songtao, et al.
Published: (2026)
by: Wang, Songtao, et al.
Published: (2026)
Query-focused and Memory-aware Reranker for Long Context Processing
by: Li, Yuqing, et al.
Published: (2026)
by: Li, Yuqing, et al.
Published: (2026)
Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
by: Ye, Xiaoju, et al.
Published: (2025)
by: Ye, Xiaoju, et al.
Published: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
by: Long, Lingkun, et al.
Published: (2026)
by: Long, Lingkun, et al.
Published: (2026)
ReAttn: Improving Attention-based Re-ranking via Attention Re-weighting
by: Tian, Yuxing, et al.
Published: (2026)
by: Tian, Yuxing, et al.
Published: (2026)
Similar Items
-
DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models
by: Ye, Xi, et al.
Published: (2026) -
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
by: Ye, Xi, et al.
Published: (2025) -
Understanding Synthetic Context Extension via Retrieval Heads
by: Zhao, Xinyu, et al.
Published: (2024) -
Long-Context Language Modeling with Parallel Context Encoding
by: Yen, Howard, et al.
Published: (2024) -
Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
by: Lee, Yoonsang, et al.
Published: (2026)