ParallelComp: Parallel Long-Context Compressor for Length Extrapolation
Fuente:
arXiv
Saved in:
| Main Authors: | Xiong, Jing, Shen, Jianghan, Zheng, Chuanyang, Wan, Zhongwei, Zhao, Chenyang, Yang, Chiwun, Ye, Fanghua, Yang, Hongxia, Kong, Lingpeng, Wong, Ngai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective
by: Xiong, Jing, et al.
Published: (2024)
by: Xiong, Jing, et al.
Published: (2024)
UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
ATTS: Asynchronous Test-Time Scaling via Conformal Prediction
by: Xiong, Jing, et al.
Published: (2025)
by: Xiong, Jing, et al.
Published: (2025)
CodeComp: Structural KV Cache Compression for Agentic Coding
by: Chen, Qiujiang, et al.
Published: (2026)
by: Chen, Qiujiang, et al.
Published: (2026)
DoPE: Denoising Rotary Position Embedding
by: Xiong, Jing, et al.
Published: (2025)
by: Xiong, Jing, et al.
Published: (2025)
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
LongEmotion: Measuring Emotional Intelligence of Large Language Models in Long-Context Interaction
by: Liu, Weichu, et al.
Published: (2025)
by: Liu, Weichu, et al.
Published: (2025)
SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving
by: Xu, Wendong, et al.
Published: (2025)
by: Xu, Wendong, et al.
Published: (2025)
One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
by: Addanki, Raghav, et al.
Published: (2023)
by: Addanki, Raghav, et al.
Published: (2023)
DAPE: Data-Adaptive Positional Encoding for Length Extrapolation
by: Zheng, Chuanyang, et al.
Published: (2024)
by: Zheng, Chuanyang, et al.
Published: (2024)
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
by: Yang, Chiwun
Published: (2025)
by: Yang, Chiwun
Published: (2025)
MMFormalizer: Multimodal Autoformalization in the Wild
by: Xiong, Jing, et al.
Published: (2026)
by: Xiong, Jing, et al.
Published: (2026)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
DAPE V2: Process Attention Score as Feature Map for Length Extrapolation
by: Zheng, Chuanyang, et al.
Published: (2024)
by: Zheng, Chuanyang, et al.
Published: (2024)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
by: Xu, Wendong, et al.
Published: (2025)
by: Xu, Wendong, et al.
Published: (2025)
Context-aware Biases for Length Extrapolation
by: Veisi, Ali, et al.
Published: (2025)
by: Veisi, Ali, et al.
Published: (2025)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
by: Wang, Shiju, et al.
Published: (2025)
by: Wang, Shiju, et al.
Published: (2025)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
Long-Context Language Modeling with Parallel Context Encoding
by: Yen, Howard, et al.
Published: (2024)
by: Yen, Howard, et al.
Published: (2024)
Towards Infinite-Long Prefix in Transformer
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Self-Infilling Code Generation
by: Zheng, Lin, et al.
Published: (2023)
by: Zheng, Lin, et al.
Published: (2023)
OVD: On-policy Verbal Distillation
by: Xiong, Jing, et al.
Published: (2026)
by: Xiong, Jing, et al.
Published: (2026)
Autoregressive Models in Vision: A Survey
by: Xiong, Jing, et al.
Published: (2024)
by: Xiong, Jing, et al.
Published: (2024)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
by: He, Zhenyu, et al.
Published: (2024)
by: He, Zhenyu, et al.
Published: (2024)
DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
by: Jiang, Chenyu, et al.
Published: (2025)
by: Jiang, Chenyu, et al.
Published: (2025)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
by: Yao, Feiyu, et al.
Published: (2026)
by: Yao, Feiyu, et al.
Published: (2026)
An Evaluation of Context Length Extrapolation in Long Code via Positional Embeddings and Efficient Attention
by: Ghosh, Madhusudan, et al.
Published: (2026)
by: Ghosh, Madhusudan, et al.
Published: (2026)
Why Does the Effective Context Length of LLMs Fall Short?
by: An, Chenxin, et al.
Published: (2024)
by: An, Chenxin, et al.
Published: (2024)
An Adaptive Parallel Arc-Length Method
by: Verhelst, H. M., et al.
Published: (2023)
by: Verhelst, H. M., et al.
Published: (2023)
Parallelizing MCMC Across the Sequence Length
by: Zoltowski, David M., et al.
Published: (2025)
by: Zoltowski, David M., et al.
Published: (2025)
Context Parallelism for Scalable Million-Token Inference
by: Yang, Amy, et al.
Published: (2024)
by: Yang, Amy, et al.
Published: (2024)
Parallel-in-Time Nonlinear Optimal Control via GPU-native Sequential Convex Programming
by: Zou, Yilin, et al.
Published: (2026)
by: Zou, Yilin, et al.
Published: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
Minute-Long Videos with Dual Parallelisms
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
MSz: An Efficient Parallel Algorithm for Correcting Morse-Smale Segmentations in Error-Bounded Lossy Compressors
by: Li, Yuxiao, et al.
Published: (2024)
by: Li, Yuxiao, et al.
Published: (2024)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
by: Gu, Diandian, et al.
Published: (2024)
by: Gu, Diandian, et al.
Published: (2024)
Balancing Pipeline Parallelism with Vocabulary Parallelism
by: Yeung, Man Tsung, et al.
Published: (2024)
by: Yeung, Man Tsung, et al.
Published: (2024)
ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
by: Kong, Quan, et al.
Published: (2026)
by: Kong, Quan, et al.
Published: (2026)
DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor Search
by: Yang, Lei, et al.
Published: (2024)
by: Yang, Lei, et al.
Published: (2024)
Similar Items
-
UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective
by: Xiong, Jing, et al.
Published: (2024) -
UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation
by: Li, Zixuan, et al.
Published: (2024) -
ATTS: Asynchronous Test-Time Scaling via Conformal Prediction
by: Xiong, Jing, et al.
Published: (2025) -
CodeComp: Structural KV Cache Compression for Agentic Coding
by: Chen, Qiujiang, et al.
Published: (2026) -
DoPE: Denoising Rotary Position Embedding
by: Xiong, Jing, et al.
Published: (2025)