CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Lv, Junlin, Feng, Yuan, Xie, Xike, Jia, Xin, Peng, Qirong, Xie, Guiming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
por: Peng, Dan, et al.
Publicado: (2025)
por: Peng, Dan, et al.
Publicado: (2025)
Block-Attention for Efficient Prefilling
por: Ma, Dongyang, et al.
Publicado: (2024)
por: Ma, Dongyang, et al.
Publicado: (2024)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
por: Feng, Yuan, et al.
Publicado: (2024)
por: Feng, Yuan, et al.
Publicado: (2024)
FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
por: Fan, Qihang, et al.
Publicado: (2026)
por: Fan, Qihang, et al.
Publicado: (2026)
PreFT: Prefill-only finetuning for efficient inference
por: Lanpouthakoun, Andrew, et al.
Publicado: (2026)
por: Lanpouthakoun, Andrew, et al.
Publicado: (2026)
Prefill-Guided Thinking for zero-shot detection of AI-generated images
por: Kachwala, Zoher, et al.
Publicado: (2025)
por: Kachwala, Zoher, et al.
Publicado: (2025)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
por: Kasnavieh, Hossein Hosseini, et al.
Publicado: (2026)
por: Kasnavieh, Hossein Hosseini, et al.
Publicado: (2026)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
por: Qiao, Aurick, et al.
Publicado: (2024)
por: Qiao, Aurick, et al.
Publicado: (2024)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
por: Struppek, Lukas, et al.
Publicado: (2026)
por: Struppek, Lukas, et al.
Publicado: (2026)
Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
por: Zhao, Siyan, et al.
Publicado: (2024)
por: Zhao, Siyan, et al.
Publicado: (2024)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
por: Gautam, Aayush, et al.
Publicado: (2026)
por: Gautam, Aayush, et al.
Publicado: (2026)
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
por: Zhang, Hao, et al.
Publicado: (2025)
por: Zhang, Hao, et al.
Publicado: (2025)
UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification
por: Fan, Qihang, et al.
Publicado: (2026)
por: Fan, Qihang, et al.
Publicado: (2026)
Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing
por: Liu, Ziyang
Publicado: (2026)
por: Liu, Ziyang
Publicado: (2026)
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
por: Liu, Jingyu, et al.
Publicado: (2025)
por: Liu, Jingyu, et al.
Publicado: (2025)
Path Pooling: Training-Free Structure Enhancement for Efficient Knowledge Graph Retrieval-Augmented Generation
por: Wang, Hairu, et al.
Publicado: (2025)
por: Wang, Hairu, et al.
Publicado: (2025)
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
por: An, Tai, et al.
Publicado: (2025)
por: An, Tai, et al.
Publicado: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
por: Lai, Xin, et al.
Publicado: (2024)
por: Lai, Xin, et al.
Publicado: (2024)
QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
por: Jones, Dalton, et al.
Publicado: (2026)
por: Jones, Dalton, et al.
Publicado: (2026)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
por: Dotsinski, Asen, et al.
Publicado: (2026)
por: Dotsinski, Asen, et al.
Publicado: (2026)
LLM Serving Optimization with Variable Prefill and Decode Lengths
por: Wang, Meixuan, et al.
Publicado: (2025)
por: Wang, Meixuan, et al.
Publicado: (2025)
POP: Prefill-Only Pruning for Efficient Large Model Inference
por: He, Junhui, et al.
Publicado: (2026)
por: He, Junhui, et al.
Publicado: (2026)
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
por: McDanel, Bradley, et al.
Publicado: (2026)
por: McDanel, Bradley, et al.
Publicado: (2026)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
por: Song, Dinghong, et al.
Publicado: (2025)
por: Song, Dinghong, et al.
Publicado: (2025)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
por: Tang, Xiaojuan, et al.
Publicado: (2025)
por: Tang, Xiaojuan, et al.
Publicado: (2025)
VSPrefill: Vertical-Slash Sparse Attention with Lightweight Indexing for Long-Context Prefilling
por: Guanzhong, Chen
Publicado: (2026)
por: Guanzhong, Chen
Publicado: (2026)
LAPS: A Length-Aware-Prefill LLM Serving System
por: She, Jianshu, et al.
Publicado: (2026)
por: She, Jianshu, et al.
Publicado: (2026)
Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models
por: Shinoda, Kazutoshi, et al.
Publicado: (2025)
por: Shinoda, Kazutoshi, et al.
Publicado: (2025)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
por: Ji, Xiaodong, et al.
Publicado: (2025)
por: Ji, Xiaodong, et al.
Publicado: (2025)
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
por: Xie, Shuo, et al.
Publicado: (2024)
por: Xie, Shuo, et al.
Publicado: (2024)
LLM Router: Rethinking Routing with Prefill Activations
por: Varshney, Tanay, et al.
Publicado: (2026)
por: Varshney, Tanay, et al.
Publicado: (2026)
CritiSense: Critical Digital Literacy and Resilience Against Misinformation
por: Alam, Firoj, et al.
Publicado: (2026)
por: Alam, Firoj, et al.
Publicado: (2026)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
por: Jo, Dongwon, et al.
Publicado: (2025)
por: Jo, Dongwon, et al.
Publicado: (2025)
HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
por: Zhang, Jun, et al.
Publicado: (2025)
por: Zhang, Jun, et al.
Publicado: (2025)
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
por: Feng, Yuan, et al.
Publicado: (2025)
por: Feng, Yuan, et al.
Publicado: (2025)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
por: He, Di, et al.
Publicado: (2025)
por: He, Di, et al.
Publicado: (2025)
DuetGraph: Coarse-to-Fine Knowledge Graph Reasoning with Dual-Pathway Global-Local Fusion
por: Li, Jin, et al.
Publicado: (2025)
por: Li, Jin, et al.
Publicado: (2025)
SamGoG: A Sampling-Based Graph-of-Graphs Framework for Imbalanced Graph Classification
por: Wang, Shangyou, et al.
Publicado: (2025)
por: Wang, Shangyou, et al.
Publicado: (2025)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
por: Hsieh, Chia-chi, et al.
Publicado: (2026)
por: Hsieh, Chia-chi, et al.
Publicado: (2026)
Accelerating Prefilling via Decoding-time Contribution Sparsity
por: He, Zhiyuan, et al.
Publicado: (2025)
por: He, Zhiyuan, et al.
Publicado: (2025)
Ejemplares similares
-
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
por: Peng, Dan, et al.
Publicado: (2025) -
Block-Attention for Efficient Prefilling
por: Ma, Dongyang, et al.
Publicado: (2024) -
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
por: Feng, Yuan, et al.
Publicado: (2024) -
FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
por: Fan, Qihang, et al.
Publicado: (2026) -
PreFT: Prefill-only finetuning for efficient inference
por: Lanpouthakoun, Andrew, et al.
Publicado: (2026)