ContextPilot: Fast Long-Context Inference via Context Reuse
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yinsicheng, Huang, Yeqi, Cheng, Liang, Deng, Cheng, Sun, Xuan, Mai, Luo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
KV Admission: Learning What to Write for Efficient Long-Context Inference
by: Huang, Yen-Chieh, et al.
Published: (2025)
by: Huang, Yen-Chieh, et al.
Published: (2025)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
by: Deng, Weishu, et al.
Published: (2025)
by: Deng, Weishu, et al.
Published: (2025)
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
by: Chen, Yaoqi, et al.
Published: (2025)
by: Chen, Yaoqi, et al.
Published: (2025)
UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference
by: Zhou, Lang, et al.
Published: (2026)
by: Zhou, Lang, et al.
Published: (2026)
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
by: Gao, Yifei, et al.
Published: (2026)
by: Gao, Yifei, et al.
Published: (2026)
Radar: Fast Long-Context Decoding for Any Transformer
by: Hao, Yongchang, et al.
Published: (2025)
by: Hao, Yongchang, et al.
Published: (2025)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
by: Shang, Xianpeng, et al.
Published: (2026)
by: Shang, Xianpeng, et al.
Published: (2026)
Learning to Extract Context for Context-Aware LLM Inference
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism
by: Bu, Tao, et al.
Published: (2025)
by: Bu, Tao, et al.
Published: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
Revisiting In-Context Learning with Long Context Language Models
by: Baek, Jinheon, et al.
Published: (2024)
by: Baek, Jinheon, et al.
Published: (2024)
DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
by: Jiang, Chenyu, et al.
Published: (2025)
by: Jiang, Chenyu, et al.
Published: (2025)
Characterizing Prompt Compression Methods for Long Context Inference
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
by: Jiang, Huiqiang, et al.
Published: (2023)
by: Jiang, Huiqiang, et al.
Published: (2023)
Long Context In-Context Compression by Getting to the Gist of Gisting
by: Petrov, Aleksandar, et al.
Published: (2025)
by: Petrov, Aleksandar, et al.
Published: (2025)
Mosaic: Unlocking Long-Context Inference for Diffusion LLMs via Global Memory Planning and Dynamic Peak Taming
by: Zheng, Liang, et al.
Published: (2026)
by: Zheng, Liang, et al.
Published: (2026)
CoMem: Context Management with A Decoupled Long-Context Model
by: Zhang, Yuwei, et al.
Published: (2026)
by: Zhang, Yuwei, et al.
Published: (2026)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Technical Debt in In-Context Learning: Diminishing Efficiency in Long Context
by: Joo, Taejong, et al.
Published: (2025)
by: Joo, Taejong, et al.
Published: (2025)
Core Context Aware Transformers for Long Context Language Modeling
by: Chen, Yaofo, et al.
Published: (2024)
by: Chen, Yaofo, et al.
Published: (2024)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
by: Huang, Yanwen, et al.
Published: (2025)
by: Huang, Yanwen, et al.
Published: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
by: Dzikanyanga, Gradwell, et al.
Published: (2026)
by: Dzikanyanga, Gradwell, et al.
Published: (2026)
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Scaling Long-Horizon LLM Agent via Context-Folding
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Calibrate to Discriminate: Improve In-Context Learning with Label-Free Comparative Inference
by: Cheng, Wei, et al.
Published: (2024)
by: Cheng, Wei, et al.
Published: (2024)
Online Statistical Inference in Decision-Making with Matrix Context
by: Han, Qiyu, et al.
Published: (2022)
by: Han, Qiyu, et al.
Published: (2022)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
by: Qi, Yanlin, et al.
Published: (2026)
by: Qi, Yanlin, et al.
Published: (2026)
In-Context Semi-Supervised Learning
by: Fan, Jiashuo, et al.
Published: (2025)
by: Fan, Jiashuo, et al.
Published: (2025)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
by: Zhou, Ruijie, et al.
Published: (2026)
by: Zhou, Ruijie, et al.
Published: (2026)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
In-Context Positive-Unlabeled Learning
by: Liu, Siyan, et al.
Published: (2026)
by: Liu, Siyan, et al.
Published: (2026)
ContextFlow: Context-Aware Flow Matching For Trajectory Inference From Spatial Omics Data
by: Rathod, Santanu Subhash, et al.
Published: (2025)
by: Rathod, Santanu Subhash, et al.
Published: (2025)
Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference
by: Liskavets, Barys, et al.
Published: (2024)
by: Liskavets, Barys, et al.
Published: (2024)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
by: Sun, Hanshi, et al.
Published: (2024)
by: Sun, Hanshi, et al.
Published: (2024)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Similar Items
-
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024) -
KV Admission: Learning What to Write for Efficient Long-Context Inference
by: Huang, Yen-Chieh, et al.
Published: (2025) -
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025) -
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
by: Zhang, Junyang, et al.
Published: (2025) -
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
by: Deng, Weishu, et al.
Published: (2025)