Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yong, Li, Heng, Huang, Yanwen, Cheng, Ning, Guo, Yang, Zhu, Yun, Wang, Yanmeng, Wang, Shaojun, Xiao, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
by: Huang, Yanwen, et al.
Published: (2025)
by: Huang, Yanwen, et al.
Published: (2025)
Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models
by: Liu, Kainan, et al.
Published: (2026)
by: Liu, Kainan, et al.
Published: (2026)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
by: Li, Zhuochun, et al.
Published: (2026)
by: Li, Zhuochun, et al.
Published: (2026)
Recurrent Context Compression: Efficiently Expanding the Context Window of LLM
by: Huang, Chensen, et al.
Published: (2024)
by: Huang, Chensen, et al.
Published: (2024)
GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compression
by: Liu, Kainan, et al.
Published: (2024)
by: Liu, Kainan, et al.
Published: (2024)
SSPO: Subsentence-level Policy Optimization
by: Yang, Kun, et al.
Published: (2025)
by: Yang, Kun, et al.
Published: (2025)
Attn-GS: Attention-Guided Context Compression for Efficient Personalized LLMs
by: Zeng, Shenglai, et al.
Published: (2026)
by: Zeng, Shenglai, et al.
Published: (2026)
Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Pipelined Decoder for Efficient Context-Aware Text Generation
by: Huang, Zixian, et al.
Published: (2025)
by: Huang, Zixian, et al.
Published: (2025)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
by: Zhang, Jianfei, et al.
Published: (2025)
by: Zhang, Jianfei, et al.
Published: (2025)
VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
by: Ma, Da, et al.
Published: (2024)
by: Ma, Da, et al.
Published: (2024)
Evaluating Zero-Shot Long-Context LLM Compression
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Position IDs Matter: An Enhanced Position Layout for Efficient Context Compression in Large Language Models
by: Zhao, Runsong, et al.
Published: (2024)
by: Zhao, Runsong, et al.
Published: (2024)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
by: Wu, Zimeng, et al.
Published: (2026)
by: Wu, Zimeng, et al.
Published: (2026)
Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing
by: Zheng, Tong, et al.
Published: (2026)
by: Zheng, Tong, et al.
Published: (2026)
Training-free Context-adaptive Attention for Efficient Long Context Modeling
by: You, Zeng, et al.
Published: (2025)
by: You, Zeng, et al.
Published: (2025)
Submodular Context Partitioning and Compression for In-Context Learning
by: Zheng, Shaoyi, et al.
Published: (2025)
by: Zheng, Shaoyi, et al.
Published: (2025)
Learning to Adapt to Low-Resource Paraphrase Generation
by: Li, Zhigen, et al.
Published: (2024)
by: Li, Zhigen, et al.
Published: (2024)
QUITO: Accelerating Long-Context Reasoning through Query-Guided Context Compression
by: Wang, Wenshan, et al.
Published: (2024)
by: Wang, Wenshan, et al.
Published: (2024)
Correlation-Aware Select and Merge Attention for Efficient Fine-Tuning and Context Length Extension
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression
by: Trukhina, Natalia, et al.
Published: (2026)
by: Trukhina, Natalia, et al.
Published: (2026)
Make Your LLM Fully Utilize the Context
by: An, Shengnan, et al.
Published: (2024)
by: An, Shengnan, et al.
Published: (2024)
Glyph: Scaling Context Windows via Visual-Text Compression
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
ACON: Optimizing Context Compression for Long-horizon LLM Agents
by: Kang, Minki, et al.
Published: (2025)
by: Kang, Minki, et al.
Published: (2025)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
by: Cheng, Yun, et al.
Published: (2026)
by: Cheng, Yun, et al.
Published: (2026)
Fix the Structural Bottleneck: Context Compression via Explicit Information Transmission
by: Ye, Jiangnan, et al.
Published: (2026)
by: Ye, Jiangnan, et al.
Published: (2026)
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
Effective In-Context Example Selection through Data Compression
by: Sun, Zhongxiang, et al.
Published: (2024)
by: Sun, Zhongxiang, et al.
Published: (2024)
Efficient Context Scaling with LongCat ZigZag Attention
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies
by: Chen, Zhiyang, et al.
Published: (2026)
by: Chen, Zhiyang, et al.
Published: (2026)
Structured Packing in LLM Training Improves Long Context Utilization
by: Staniszewski, Konrad, et al.
Published: (2023)
by: Staniszewski, Konrad, et al.
Published: (2023)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
by: Ma, Qingsen, et al.
Published: (2026)
by: Ma, Qingsen, et al.
Published: (2026)
A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
by: Ren, Jincheng, et al.
Published: (2026)
by: Ren, Jincheng, et al.
Published: (2026)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
by: Lin, Gang, et al.
Published: (2026)
by: Lin, Gang, et al.
Published: (2026)
CompLLM: Compression for Long Context Q&A
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
Similar Items
-
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
by: Huang, Yanwen, et al.
Published: (2025) -
Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models
by: Liu, Kainan, et al.
Published: (2026) -
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
by: Li, Zhuochun, et al.
Published: (2026) -
Recurrent Context Compression: Efficiently Expanding the Context Window of LLM
by: Huang, Chensen, et al.
Published: (2024) -
GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compression
by: Liu, Kainan, et al.
Published: (2024)