UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Lang, Li, Shuxuan, Li, Zhuohao, Liu, Shi, Zhao, Zhilin, Zheng, Wei-Shi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
by: Sun, Hanshi, et al.
Published: (2024)
by: Sun, Hanshi, et al.
Published: (2024)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
by: Shang, Xianpeng, et al.
Published: (2026)
by: Shang, Xianpeng, et al.
Published: (2026)
Towards Better Understanding of In-Context Learning Ability from In-Context Uncertainty Quantification
by: Liu, Shang, et al.
Published: (2024)
by: Liu, Shang, et al.
Published: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
by: Li, Weizhuo, et al.
Published: (2024)
by: Li, Weizhuo, et al.
Published: (2024)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
by: Ye, Jiancai, et al.
Published: (2026)
by: Ye, Jiancai, et al.
Published: (2026)
Core Context Aware Transformers for Long Context Language Modeling
by: Chen, Yaofo, et al.
Published: (2024)
by: Chen, Yaofo, et al.
Published: (2024)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
Intrinsic Entropy of Context Length Scaling in LLMs
by: Shi, Jingzhe, et al.
Published: (2025)
by: Shi, Jingzhe, et al.
Published: (2025)
Characterizing Prompt Compression Methods for Long Context Inference
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
by: Li, Zeju, et al.
Published: (2026)
by: Li, Zeju, et al.
Published: (2026)
Uncertainty Quantification for In-Context Learning of Large Language Models
by: Ling, Chen, et al.
Published: (2024)
by: Ling, Chen, et al.
Published: (2024)
LongEmbed: Extending Embedding Models for Long Context Retrieval
by: Zhu, Dawei, et al.
Published: (2024)
by: Zhu, Dawei, et al.
Published: (2024)
A Comprehensive Survey on Long Context Language Modeling
by: Liu, Jiaheng, et al.
Published: (2025)
by: Liu, Jiaheng, et al.
Published: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
by: Li, Kunxi, et al.
Published: (2025)
by: Li, Kunxi, et al.
Published: (2025)
LongSafety: Enhance Safety for Long-Context LLMs
by: Huang, Mianqiu, et al.
Published: (2024)
by: Huang, Mianqiu, et al.
Published: (2024)
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
by: Li, Chengye, et al.
Published: (2025)
by: Li, Chengye, et al.
Published: (2025)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
by: Donhauser, Konstantin, et al.
Published: (2025)
by: Donhauser, Konstantin, et al.
Published: (2025)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
by: Zhu, Qianchao, et al.
Published: (2024)
by: Zhu, Qianchao, et al.
Published: (2024)
Calibrate to Discriminate: Improve In-Context Learning with Label-Free Comparative Inference
by: Cheng, Wei, et al.
Published: (2024)
by: Cheng, Wei, et al.
Published: (2024)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
by: Li, Dongfang, et al.
Published: (2026)
by: Li, Dongfang, et al.
Published: (2026)
Can Language Models Compose Skills In-Context?
by: Liu, Zidong, et al.
Published: (2025)
by: Liu, Zidong, et al.
Published: (2025)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
by: Behnam, Payman, et al.
Published: (2025)
by: Behnam, Payman, et al.
Published: (2025)
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
The Impossibility Triangle of Long-Context Modeling
by: Zhou, Yan
Published: (2026)
by: Zhou, Yan
Published: (2026)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
by: Ma, Qingsen, et al.
Published: (2026)
by: Ma, Qingsen, et al.
Published: (2026)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
by: Wu, Zimeng, et al.
Published: (2026)
by: Wu, Zimeng, et al.
Published: (2026)
Out-of-Domain Intent Detection Considering Multi-Turn Dialogue Contexts
by: Lang, Hao, et al.
Published: (2023)
by: Lang, Hao, et al.
Published: (2023)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
by: Qiu, Quantong, et al.
Published: (2026)
by: Qiu, Quantong, et al.
Published: (2026)
MoBA: Mixture of Block Attention for Long-Context LLMs
by: Lu, Enzhe, et al.
Published: (2025)
by: Lu, Enzhe, et al.
Published: (2025)
Multipole Attention for Efficient Long Context Reasoning
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
MOOSComp: Improving Lightweight Long-Context Compressor via Mitigating Over-Smoothing and Incorporating Outlier Scores
by: Zhou, Fengwei, et al.
Published: (2025)
by: Zhou, Fengwei, et al.
Published: (2025)
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
by: Ping, Bowen, et al.
Published: (2026)
by: Ping, Bowen, et al.
Published: (2026)
LongAlign: A Recipe for Long Context Alignment of Large Language Models
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
by: Dzikanyanga, Gradwell, et al.
Published: (2026)
by: Dzikanyanga, Gradwell, et al.
Published: (2026)
Similar Items
-
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025) -
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
by: Sun, Hanshi, et al.
Published: (2024) -
Training-Inference Consistent Segmented Execution for Long-Context LLMs
by: Shang, Xianpeng, et al.
Published: (2026) -
Towards Better Understanding of In-Context Learning Ability from In-Context Uncertainty Quantification
by: Liu, Shang, et al.
Published: (2024) -
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)