Retrieval Head Mechanistically Explains Long-Context Factuality
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Wenhao, Wang, Yizhong, Xiao, Guangxuan, Peng, Hao, Fu, Yao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
Long Context Alignment with Short Instructions and Synthesized Positions
von: Wu, Wenhao, et al.
Veröffentlicht: (2024)
von: Wu, Wenhao, et al.
Veröffentlicht: (2024)
Context-Efficient Retrieval with Factual Decomposition
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
LongEmbed: Extending Embedding Models for Long Context Retrieval
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
von: Wang, Shuxun, et al.
Veröffentlicht: (2025)
von: Wang, Shuxun, et al.
Veröffentlicht: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
von: Liu, Zhuorui, et al.
Veröffentlicht: (2025)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
von: Ma, Youmi, et al.
Veröffentlicht: (2026)
von: Ma, Youmi, et al.
Veröffentlicht: (2026)
Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
von: Zhang, Wuwei, et al.
Veröffentlicht: (2025)
von: Zhang, Wuwei, et al.
Veröffentlicht: (2025)
VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding
von: Pei, Rongcan, et al.
Veröffentlicht: (2026)
von: Pei, Rongcan, et al.
Veröffentlicht: (2026)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
von: Liu, Runheng, et al.
Veröffentlicht: (2024)
Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation
von: Ren, Ruiyang, et al.
Veröffentlicht: (2023)
von: Ren, Ruiyang, et al.
Veröffentlicht: (2023)
Understanding Synthetic Context Extension via Retrieval Heads
von: Zhao, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xinyu, et al.
Veröffentlicht: (2024)
HGOT: Hierarchical Graph of Thoughts for Retrieval-Augmented In-Context Learning in Factuality Evaluation
von: Fang, Yihao, et al.
Veröffentlicht: (2024)
von: Fang, Yihao, et al.
Veröffentlicht: (2024)
Long$^2$RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall
von: Qi, Zehan, et al.
Veröffentlicht: (2024)
von: Qi, Zehan, et al.
Veröffentlicht: (2024)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
von: Lu, Yi, et al.
Veröffentlicht: (2024)
von: Lu, Yi, et al.
Veröffentlicht: (2024)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
Co-occurrence is not Factual Association in Language Models
von: Zhang, Xiao, et al.
Veröffentlicht: (2024)
von: Zhang, Xiao, et al.
Veröffentlicht: (2024)
Knowledgeable In-Context Tuning: Exploring and Exploiting Factual Knowledge for In-Context Learning
von: Wang, Jianing, et al.
Veröffentlicht: (2023)
von: Wang, Jianing, et al.
Veröffentlicht: (2023)
FoRAG: Factuality-optimized Retrieval Augmented Generation for Web-enhanced Long-form Question Answering
von: Cai, Tianchi, et al.
Veröffentlicht: (2024)
von: Cai, Tianchi, et al.
Veröffentlicht: (2024)
Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
von: Ye, Xiaoju, et al.
Veröffentlicht: (2025)
von: Ye, Xiaoju, et al.
Veröffentlicht: (2025)
Inference Scaling for Long-Context Retrieval Augmented Generation
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
Factuality and Transparency Are All RAG Needs! Self-Explaining Contrastive Evidence Re-ranking
von: Vargas, Francielle, et al.
Veröffentlicht: (2025)
von: Vargas, Francielle, et al.
Veröffentlicht: (2025)
Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding
von: Li, Yuqing, et al.
Veröffentlicht: (2025)
von: Li, Yuqing, et al.
Veröffentlicht: (2025)
CTkvr: KV Cache Retrieval for Long-Context LLMs via Centroid then Token Indexing
von: Lu, Kuan, et al.
Veröffentlicht: (2025)
von: Lu, Kuan, et al.
Veröffentlicht: (2025)
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
von: Ge, Suyu, et al.
Veröffentlicht: (2024)
von: Ge, Suyu, et al.
Veröffentlicht: (2024)
Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality
von: Luo, Wen, et al.
Veröffentlicht: (2026)
von: Luo, Wen, et al.
Veröffentlicht: (2026)
Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation
von: Samarinas, Chris, et al.
Veröffentlicht: (2025)
von: Samarinas, Chris, et al.
Veröffentlicht: (2025)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2024)
Agent-as-Judge for Factual Summarization of Long Narratives
von: Jeong, Yeonseok, et al.
Veröffentlicht: (2025)
von: Jeong, Yeonseok, et al.
Veröffentlicht: (2025)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
von: Liu, Siran, et al.
Veröffentlicht: (2026)
von: Liu, Siran, et al.
Veröffentlicht: (2026)
Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
von: Fu, Yu, et al.
Veröffentlicht: (2024)
von: Fu, Yu, et al.
Veröffentlicht: (2024)
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis
von: Fu, Yao
Veröffentlicht: (2024)
von: Fu, Yao
Veröffentlicht: (2024)
WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems
von: Yu, Jiangnan, et al.
Veröffentlicht: (2026)
von: Yu, Jiangnan, et al.
Veröffentlicht: (2026)
Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness
von: Bi, Baolong, et al.
Veröffentlicht: (2024)
von: Bi, Baolong, et al.
Veröffentlicht: (2024)
Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
von: Fei, Weizhi, et al.
Veröffentlicht: (2025)
BTR: Binary Token Representations for Efficient Retrieval Augmented Language Models
von: Cao, Qingqing, et al.
Veröffentlicht: (2023)
von: Cao, Qingqing, et al.
Veröffentlicht: (2023)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
von: Lin, Xihui, et al.
Veröffentlicht: (2024)
von: Lin, Xihui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024) -
Long Context Alignment with Short Instructions and Synthesized Positions
von: Wu, Wenhao, et al.
Veröffentlicht: (2024) -
Context-Efficient Retrieval with Factual Decomposition
von: Li, Yanhong, et al.
Veröffentlicht: (2025) -
LongEmbed: Extending Embedding Models for Long Context Retrieval
von: Zhu, Dawei, et al.
Veröffentlicht: (2024) -
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)