CHESS: Context-aware Hierarchical Efficient Semantic Selection for Long-Context LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fei, Chao, Li, Guozhong, Liu, Chenxi, Kalnis, Panos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMComp: A Language Modeling Paradigm for Error-Bounded Scientific Data Compression (Technical Report)
von: Li, Guozhong, et al.
Veröffentlicht: (2025)
von: Li, Guozhong, et al.
Veröffentlicht: (2025)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
von: He, Junhui, et al.
Veröffentlicht: (2024)
von: He, Junhui, et al.
Veröffentlicht: (2024)
ART: Attention Run-time Termination for Efficient Large Language Model Decoding
von: Qiu, Chen, et al.
Veröffentlicht: (2026)
von: Qiu, Chen, et al.
Veröffentlicht: (2026)
LLMs Meet Cross-Modal Time Series Analytics: Overview and Directions
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
Comprehending Spatio-temporal Data via Cinematic Storytelling using Large Language Models
von: Shang, Panos Kalnis. Shuo, et al.
Veröffentlicht: (2025)
von: Shang, Panos Kalnis. Shuo, et al.
Veröffentlicht: (2025)
Task-Oriented GNNs Training on Large Knowledge Graphs for Accurate and Efficient Modeling
von: Abdallah, Hussein, et al.
Veröffentlicht: (2024)
von: Abdallah, Hussein, et al.
Veröffentlicht: (2024)
Leveraging LLM-GNN Integration for Open-World Question Answering over Knowledge Graphs
von: Abdallah, Hussein, et al.
Veröffentlicht: (2026)
von: Abdallah, Hussein, et al.
Veröffentlicht: (2026)
From GPS Points to Travel Patterns: Flexible and Semantic Trajectory Generation with LLMs
von: Zhou, Silin, et al.
Veröffentlicht: (2026)
von: Zhou, Silin, et al.
Veröffentlicht: (2026)
ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
von: Li, Junjie, et al.
Veröffentlicht: (2026)
von: Li, Junjie, et al.
Veröffentlicht: (2026)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
Hierarchical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLM
von: Yao, Yongqiang, et al.
Veröffentlicht: (2025)
von: Yao, Yongqiang, et al.
Veröffentlicht: (2025)
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
von: Zhou, Ruijie, et al.
Veröffentlicht: (2026)
von: Zhou, Ruijie, et al.
Veröffentlicht: (2026)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
Rover: Context-aware Conflict Resolution with LLM
von: Zhang, Qingyu, et al.
Veröffentlicht: (2026)
von: Zhang, Qingyu, et al.
Veröffentlicht: (2026)
CHESS: Contextual Harnessing for Efficient SQL Synthesis
von: Talaei, Shayan, et al.
Veröffentlicht: (2024)
von: Talaei, Shayan, et al.
Veröffentlicht: (2024)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
RED: Effective Trajectory Representation Learning with Comprehensive Information
von: Zhou, Silin, et al.
Veröffentlicht: (2024)
von: Zhou, Silin, et al.
Veröffentlicht: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
von: Mai, Tho, et al.
Veröffentlicht: (2026)
von: Mai, Tho, et al.
Veröffentlicht: (2026)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
Context Memorization for Efficient Long Context Generation
von: Okoshi, Yasuyuki, et al.
Veröffentlicht: (2026)
von: Okoshi, Yasuyuki, et al.
Veröffentlicht: (2026)
Efficient Low Rank Attention for Long-Context Inference in Large Language Models
von: Li, Tenghui, et al.
Veröffentlicht: (2025)
von: Li, Tenghui, et al.
Veröffentlicht: (2025)
KV Admission: Learning What to Write for Efficient Long-Context Inference
von: Huang, Yen-Chieh, et al.
Veröffentlicht: (2025)
von: Huang, Yen-Chieh, et al.
Veröffentlicht: (2025)
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-$k$
von: Taguchi, Chihiro, et al.
Veröffentlicht: (2025)
von: Taguchi, Chihiro, et al.
Veröffentlicht: (2025)
Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs
von: Song, Woomin, et al.
Veröffentlicht: (2024)
von: Song, Woomin, et al.
Veröffentlicht: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
von: Yao, Feiyu, et al.
Veröffentlicht: (2026)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
von: Li, Zhouyang, et al.
Veröffentlicht: (2025)
von: Li, Zhouyang, et al.
Veröffentlicht: (2025)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
von: Kolasani, Sai, et al.
Veröffentlicht: (2025)
von: Kolasani, Sai, et al.
Veröffentlicht: (2025)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
von: Monteiro, João, et al.
Veröffentlicht: (2024)
von: Monteiro, João, et al.
Veröffentlicht: (2024)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
von: Zhang, Hailin, et al.
Veröffentlicht: (2024)
von: Zhang, Hailin, et al.
Veröffentlicht: (2024)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
von: Lin, Gang, et al.
Veröffentlicht: (2026)
von: Lin, Gang, et al.
Veröffentlicht: (2026)
Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context Prefilling
von: Chen, Yujie, et al.
Veröffentlicht: (2026)
von: Chen, Yujie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LLMComp: A Language Modeling Paradigm for Error-Bounded Scientific Data Compression (Technical Report)
von: Li, Guozhong, et al.
Veröffentlicht: (2025) -
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
von: He, Junhui, et al.
Veröffentlicht: (2024) -
ART: Attention Run-time Termination for Efficient Large Language Model Decoding
von: Qiu, Chen, et al.
Veröffentlicht: (2026) -
LLMs Meet Cross-Modal Time Series Analytics: Overview and Directions
von: Liu, Chenxi, et al.
Veröffentlicht: (2025) -
Comprehending Spatio-temporal Data via Cinematic Storytelling using Large Language Models
von: Shang, Panos Kalnis. Shuo, et al.
Veröffentlicht: (2025)