CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Junhui, Wu, Shangyu, Wen, Weidong, Xue, Chun Jason, Li, Qingan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
di: Sattarifard, Amirmohsen, et al.
Pubblicazione: (2025)
di: Sattarifard, Amirmohsen, et al.
Pubblicazione: (2025)
EvoP: Robust LLM Inference via Evolutionary Pruning
di: Wu, Shangyu, et al.
Pubblicazione: (2025)
di: Wu, Shangyu, et al.
Pubblicazione: (2025)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
di: Yang, Yifei, et al.
Pubblicazione: (2024)
di: Yang, Yifei, et al.
Pubblicazione: (2024)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
di: Li, Xing, et al.
Pubblicazione: (2025)
di: Li, Xing, et al.
Pubblicazione: (2025)
SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs
di: Czakó, Patrik, et al.
Pubblicazione: (2025)
di: Czakó, Patrik, et al.
Pubblicazione: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering
di: Zhou, Yefan, et al.
Pubblicazione: (2026)
di: Zhou, Yefan, et al.
Pubblicazione: (2026)
POP: Prefill-Only Pruning for Efficient Large Model Inference
di: He, Junhui, et al.
Pubblicazione: (2026)
di: He, Junhui, et al.
Pubblicazione: (2026)
Language Model Prompt Selection via Simulation Optimization
di: Zhang, Haoting, et al.
Pubblicazione: (2024)
di: Zhang, Haoting, et al.
Pubblicazione: (2024)
Tree Matching Networks for Natural Language Inference: Parameter-Efficient Semantic Understanding via Dependency Parse Trees
di: Lunder, Jason
Pubblicazione: (2025)
di: Lunder, Jason
Pubblicazione: (2025)
Less is More: Improving LLM Alignment via Preference Data Selection
di: Deng, Xun, et al.
Pubblicazione: (2025)
di: Deng, Xun, et al.
Pubblicazione: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
di: Dzikanyanga, Gradwell, et al.
Pubblicazione: (2026)
di: Dzikanyanga, Gradwell, et al.
Pubblicazione: (2026)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
di: Wu, Wei, et al.
Pubblicazione: (2024)
di: Wu, Wei, et al.
Pubblicazione: (2024)
Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
di: Bohne, Jason, et al.
Pubblicazione: (2025)
di: Bohne, Jason, et al.
Pubblicazione: (2025)
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
di: Sim, Woo Seob, et al.
Pubblicazione: (2026)
di: Sim, Woo Seob, et al.
Pubblicazione: (2026)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
di: Guo, Song, et al.
Pubblicazione: (2024)
di: Guo, Song, et al.
Pubblicazione: (2024)
CHESS: Contextual Harnessing for Efficient SQL Synthesis
di: Talaei, Shayan, et al.
Pubblicazione: (2024)
di: Talaei, Shayan, et al.
Pubblicazione: (2024)
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
di: Zhong, Tianle, et al.
Pubblicazione: (2026)
di: Zhong, Tianle, et al.
Pubblicazione: (2026)
On the Compressibility of Quantized Large Language Models
di: Mao, Yu, et al.
Pubblicazione: (2024)
di: Mao, Yu, et al.
Pubblicazione: (2024)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
di: Huang, Kexin, et al.
Pubblicazione: (2025)
di: Huang, Kexin, et al.
Pubblicazione: (2025)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
di: Qiao, Aurick, et al.
Pubblicazione: (2024)
di: Qiao, Aurick, et al.
Pubblicazione: (2024)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
di: Zhang, Xichen, et al.
Pubblicazione: (2025)
di: Zhang, Xichen, et al.
Pubblicazione: (2025)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
di: Chen, Zhipeng, et al.
Pubblicazione: (2026)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
di: Lu, Liming, et al.
Pubblicazione: (2026)
di: Lu, Liming, et al.
Pubblicazione: (2026)
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
di: Wang, Yanshu, et al.
Pubblicazione: (2024)
di: Wang, Yanshu, et al.
Pubblicazione: (2024)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
Shapley-Value-Based Graph Sparsification for GNN Inference
di: Akkas, Selahattin, et al.
Pubblicazione: (2025)
di: Akkas, Selahattin, et al.
Pubblicazione: (2025)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
di: Huang, Audrey, et al.
Pubblicazione: (2024)
di: Huang, Audrey, et al.
Pubblicazione: (2024)
One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging
di: Luo, Yingfeng, et al.
Pubblicazione: (2025)
di: Luo, Yingfeng, et al.
Pubblicazione: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
di: Wu, Junkang, et al.
Pubblicazione: (2024)
di: Wu, Junkang, et al.
Pubblicazione: (2024)
Fibration Policy Optimization
di: Li, Chang, et al.
Pubblicazione: (2026)
di: Li, Chang, et al.
Pubblicazione: (2026)
Filter-then-Weight: Online Data Selection and Reweighting for LLM Fine-Tuning
di: Wang, Fangxin, et al.
Pubblicazione: (2026)
di: Wang, Fangxin, et al.
Pubblicazione: (2026)
Membership Inference Attacks on LLM-based Recommender Systems
di: He, Jiajie, et al.
Pubblicazione: (2025)
di: He, Jiajie, et al.
Pubblicazione: (2025)
LLM-Select: Feature Selection with Large Language Models
di: Jeong, Daniel P., et al.
Pubblicazione: (2024)
di: Jeong, Daniel P., et al.
Pubblicazione: (2024)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
di: Ju, Yiming, et al.
Pubblicazione: (2024)
di: Ju, Yiming, et al.
Pubblicazione: (2024)
ChameleonLLM: Batch-Aware Dynamic Low-Rank Adaptation via Inference-Time Clusters
di: Yuksel, Kamer Ali, et al.
Pubblicazione: (2025)
di: Yuksel, Kamer Ali, et al.
Pubblicazione: (2025)
Communication Compression for Tensor Parallel LLM Inference
di: Hansen-Palmus, Jan, et al.
Pubblicazione: (2024)
di: Hansen-Palmus, Jan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
di: Taniguchi, Rei, et al.
Pubblicazione: (2026) -
GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
di: Sattarifard, Amirmohsen, et al.
Pubblicazione: (2025) -
EvoP: Robust LLM Inference via Evolutionary Pruning
di: Wu, Shangyu, et al.
Pubblicazione: (2025) -
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
di: Yang, Yifei, et al.
Pubblicazione: (2024) -
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
di: Li, Xing, et al.
Pubblicazione: (2025)