SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Jiaming, Pan, Jiayi, Wang, Hanzhen, Zhou, Yongkang, Ye, Jiancai, Wang, Yu, Dai, Guohao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
by: Wang, Hanzhen, et al.
Published: (2025)
by: Wang, Hanzhen, et al.
Published: (2025)
SpecDiff: Accelerating Diffusion Model Inference with Self-Speculation
by: Pan, Jiayi, et al.
Published: (2025)
by: Pan, Jiayi, et al.
Published: (2025)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
by: Xu, Jiaming, et al.
Published: (2025)
by: Xu, Jiaming, et al.
Published: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
by: Ye, Jiancai, et al.
Published: (2026)
by: Ye, Jiancai, et al.
Published: (2026)
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
by: Wang, Luning, et al.
Published: (2024)
by: Wang, Luning, et al.
Published: (2024)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
by: Yang, Penghui, et al.
Published: (2025)
by: Yang, Penghui, et al.
Published: (2025)
PRISM: Efficient Long-Range Reasoning With Short-Context LLMs
by: Jayalath, Dulhan, et al.
Published: (2024)
by: Jayalath, Dulhan, et al.
Published: (2024)
SpeCa: Accelerating Diffusion Transformers with Speculative Feature Caching
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
Reasoning BO: Enhancing Bayesian Optimization with Long-Context Reasoning Power of LLMs
by: Yang, Zhuo, et al.
Published: (2025)
by: Yang, Zhuo, et al.
Published: (2025)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
BalanceGS: Algorithm-System Co-design for Efficient 3D Gaussian Splatting Training on GPU
by: Wu, Junyi, et al.
Published: (2025)
by: Wu, Junyi, et al.
Published: (2025)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
by: Jia, Junlong, et al.
Published: (2025)
by: Jia, Junlong, et al.
Published: (2025)
SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs
by: Jie, Shibo, et al.
Published: (2025)
by: Jie, Shibo, et al.
Published: (2025)
Are Long-LLMs A Necessity For Long-Context Tasks?
by: Qian, Hongjin, et al.
Published: (2024)
by: Qian, Hongjin, et al.
Published: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning
by: Liu, Dong, et al.
Published: (2026)
by: Liu, Dong, et al.
Published: (2026)
QUITO: Accelerating Long-Context Reasoning through Query-Guided Context Compression
by: Wang, Wenshan, et al.
Published: (2024)
by: Wang, Wenshan, et al.
Published: (2024)
LongSafety: Enhance Safety for Long-Context LLMs
by: Huang, Mianqiu, et al.
Published: (2024)
by: Huang, Mianqiu, et al.
Published: (2024)
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
by: Cui, Hao, et al.
Published: (2025)
by: Cui, Hao, et al.
Published: (2025)
KCR: Resolving Long-Context Knowledge Conflicts via Reasoning in LLMs
by: Zheng, Xianda, et al.
Published: (2025)
by: Zheng, Xianda, et al.
Published: (2025)
A Decomposition Perspective to Long-context Reasoning for LLMs
by: Xiao, Yanling, et al.
Published: (2026)
by: Xiao, Yanling, et al.
Published: (2026)
GraphIC: A Graph-Based In-Context Example Retrieval Model for Multi-Step Reasoning
by: Fu, Jiale, et al.
Published: (2024)
by: Fu, Jiale, et al.
Published: (2024)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
by: Peng, Dan, et al.
Published: (2025)
by: Peng, Dan, et al.
Published: (2025)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
ATACompressor: Adaptive Task-Aware Compression for Efficient Long-Context Processing in LLMs
by: Li, Xuancheng, et al.
Published: (2026)
by: Li, Xuancheng, et al.
Published: (2026)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
by: Kuratov, Yuri, et al.
Published: (2024)
by: Kuratov, Yuri, et al.
Published: (2024)
Context Memorization for Efficient Long Context Generation
by: Okoshi, Yasuyuki, et al.
Published: (2026)
by: Okoshi, Yasuyuki, et al.
Published: (2026)
Can Language Models Enable In-Context Database?
by: Pan, Yu, et al.
Published: (2024)
by: Pan, Yu, et al.
Published: (2024)
Artificial Hippocampus Networks for Efficient Long-Context Modeling
by: Fang, Yunhao, et al.
Published: (2025)
by: Fang, Yunhao, et al.
Published: (2025)
I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models
by: Mi, Zhenxing, et al.
Published: (2025)
by: Mi, Zhenxing, et al.
Published: (2025)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
by: Chaudhary, Siddharth, et al.
Published: (2025)
by: Chaudhary, Siddharth, et al.
Published: (2025)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoning
by: Zhuang, Ren, et al.
Published: (2026)
by: Zhuang, Ren, et al.
Published: (2026)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
by: Gao, Yifei, et al.
Published: (2026)
by: Gao, Yifei, et al.
Published: (2026)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
by: Li, Zeju, et al.
Published: (2026)
by: Li, Zeju, et al.
Published: (2026)
Similar Items
-
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
by: Wang, Hanzhen, et al.
Published: (2025) -
SpecDiff: Accelerating Diffusion Model Inference with Self-Speculation
by: Pan, Jiayi, et al.
Published: (2025) -
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
by: Xu, Jiaming, et al.
Published: (2025) -
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
by: Ye, Jiancai, et al.
Published: (2026) -
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
by: Wang, Luning, et al.
Published: (2024)