MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Huiqiang, Li, Yucheng, Zhang, Chengruidong, Wu, Qianhui, Luo, Xufang, Ahn, Surin, Han, Zhenhua, Abdi, Amir H., Li, Dongsheng, Lin, Chin-Yew, Yang, Yuqing, Qiu, Lili |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
SCBench: A KV Cache-Centric Analysis of Long-Context Methods
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2023)
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2023)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
Mitigate Position Bias in Large Language Models via Scaling a Single Dimension
von: Yu, Yijiong, et al.
Veröffentlicht: (2024)
von: Yu, Yijiong, et al.
Veröffentlicht: (2024)
Accelerating Prefilling via Decoding-time Contribution Sparsity
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
On Memory Construction and Retrieval for Personalized Conversational Agents
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2025)
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
von: Zhang, Yike, et al.
Veröffentlicht: (2025)
von: Zhang, Yike, et al.
Veröffentlicht: (2025)
LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2024)
von: Pan, Zhuoshi, et al.
Veröffentlicht: (2024)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
von: Liu, Anmin, et al.
Veröffentlicht: (2026)
von: Liu, Anmin, et al.
Veröffentlicht: (2026)
Unified Medical Image Pre-training in Language-Guided Common Semantic Space
von: He, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: He, Xiaoxuan, et al.
Veröffentlicht: (2023)
Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
von: Lin, Chaofan, et al.
Veröffentlicht: (2024)
von: Lin, Chaofan, et al.
Veröffentlicht: (2024)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
Position Engineering: Boosting Large Language Models through Positional Information Manipulation
von: He, Zhiyuan, et al.
Veröffentlicht: (2024)
von: He, Zhiyuan, et al.
Veröffentlicht: (2024)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
Region-Adaptive Sampling for Diffusion Transformers
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
Chain-of-Model Learning for Language Model
von: Song, Kaitao, et al.
Veröffentlicht: (2025)
von: Song, Kaitao, et al.
Veröffentlicht: (2025)
Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
von: Xiao, Emily, et al.
Veröffentlicht: (2025)
von: Xiao, Emily, et al.
Veröffentlicht: (2025)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
A Large-scale Medical Visual Task Adaptation Benchmark
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
von: Luo, Xufang, et al.
Veröffentlicht: (2025)
von: Luo, Xufang, et al.
Veröffentlicht: (2025)
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
Designing Network Algorithms via Large Language Models
von: He, Zhiyuan, et al.
Veröffentlicht: (2024)
von: He, Zhiyuan, et al.
Veröffentlicht: (2024)
Rates of convergence in long time asymptotics of an alignment model with symmetry breaking
von: Surin, Alexandre
Veröffentlicht: (2025)
von: Surin, Alexandre
Veröffentlicht: (2025)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
von: Anshumann, et al.
Veröffentlicht: (2025)
von: Anshumann, et al.
Veröffentlicht: (2025)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
Improving Sparse Autoencoder with Dynamic Attention
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
Lag-Relative Sparse Attention In Long Context Training
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
Rectified Sparse Attention
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
P‐3.9: The Influence of Parallax and Shape Type Factors on the Perception of AR Equipment in Dark Environment
von: Huiqiang Xia, et al.
Veröffentlicht: (2024)
von: Huiqiang Xia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
von: Li, Yucheng, et al.
Veröffentlicht: (2025) -
SCBench: A KV Cache-Centric Analysis of Long-Context Methods
von: Li, Yucheng, et al.
Veröffentlicht: (2024) -
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
von: Jiang, Huiqiang, et al.
Veröffentlicht: (2023) -
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
von: Li, Wenxuan, et al.
Veröffentlicht: (2025) -
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
von: Li, Yucheng, et al.
Veröffentlicht: (2025)