Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Xuan, Dong, Peiyan, Lu, Lei, Kong, Zhenglun, Li, Zhengang, Lin, Ming, Wu, Chao, Wang, Yanzhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Squat: Quant Small Language Models on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2024)
von: Shen, Xuan, et al.
Veröffentlicht: (2024)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
von: Han, Chao, et al.
Veröffentlicht: (2025)
von: Han, Chao, et al.
Veröffentlicht: (2025)
Pruning Foundation Models for High Accuracy without Retraining
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
Structured Agent Distillation for Large Language Model
von: Liu, Jun, et al.
Veröffentlicht: (2025)
von: Liu, Jun, et al.
Veröffentlicht: (2025)
Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment
von: Liu, Jun, et al.
Veröffentlicht: (2024)
von: Liu, Jun, et al.
Veröffentlicht: (2024)
Search for Efficient Large Language Models
von: Shen, Xuan, et al.
Veröffentlicht: (2024)
von: Shen, Xuan, et al.
Veröffentlicht: (2024)
RCR-Router: Efficient Role-Aware Context Routing for Multi-Agent LLM Systems with Structured Memory
von: Liu, Jun, et al.
Veröffentlicht: (2025)
von: Liu, Jun, et al.
Veröffentlicht: (2025)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
von: Lu, Haiquan, et al.
Veröffentlicht: (2026)
von: Lu, Haiquan, et al.
Veröffentlicht: (2026)
Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
Rethinking Token Reduction for State Space Models
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation
von: Kong, Zhenglun, et al.
Veröffentlicht: (2025)
von: Kong, Zhenglun, et al.
Veröffentlicht: (2025)
GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion
von: Xie, Qizhuo, et al.
Veröffentlicht: (2026)
von: Xie, Qizhuo, et al.
Veröffentlicht: (2026)
SliderQuant: Accurate Post-Training Quantization for LLMs
von: Wang, Shigeng, et al.
Veröffentlicht: (2026)
von: Wang, Shigeng, et al.
Veröffentlicht: (2026)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025)
von: Chen, Han, et al.
Veröffentlicht: (2025)
RoRA: Efficient Fine-Tuning of LLM with Reliability Optimization for Rank Adaptation
von: Liu, Jun, et al.
Veröffentlicht: (2025)
von: Liu, Jun, et al.
Veröffentlicht: (2025)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
von: Li, Ke, et al.
Veröffentlicht: (2026)
von: Li, Ke, et al.
Veröffentlicht: (2026)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
All-in-One Tuning and Structural Pruning for Domain-Specific LLMs
von: Lu, Lei, et al.
Veröffentlicht: (2024)
von: Lu, Lei, et al.
Veröffentlicht: (2024)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
von: Liang, Yesheng, et al.
Veröffentlicht: (2025)
von: Liang, Yesheng, et al.
Veröffentlicht: (2025)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
ICPC: In-context Prompt Compression with Faster Inference
von: Yu, Ziyang, et al.
Veröffentlicht: (2025)
von: Yu, Ziyang, et al.
Veröffentlicht: (2025)
QuantClaw: Precision Where It Matters for OpenClaw
von: Zhang, Manyi, et al.
Veröffentlicht: (2026)
von: Zhang, Manyi, et al.
Veröffentlicht: (2026)
When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making
von: Liu, Jun, et al.
Veröffentlicht: (2026)
von: Liu, Jun, et al.
Veröffentlicht: (2026)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
Mission Impossible: Feedback-Guided Dynamic Interactive Planning for Improving Reasoning on LLMs
von: Yan, Dong, et al.
Veröffentlicht: (2025)
von: Yan, Dong, et al.
Veröffentlicht: (2025)
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios
von: Li, Shuhuai, et al.
Veröffentlicht: (2026)
von: Li, Shuhuai, et al.
Veröffentlicht: (2026)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
von: Xiao, Yuxin, et al.
Veröffentlicht: (2024)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2024)
LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
von: Ma, Guangyuan, et al.
Veröffentlicht: (2025)
von: Ma, Guangyuan, et al.
Veröffentlicht: (2025)
FlatQuant: Flatness Matters for LLM Quantization
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Squat: Quant Small Language Models on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2024) -
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
von: Han, Chao, et al.
Veröffentlicht: (2025) -
Pruning Foundation Models for High Accuracy without Retraining
von: Zhao, Pu, et al.
Veröffentlicht: (2024) -
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
von: Li, Zhengang, et al.
Veröffentlicht: (2024) -
Structured Agent Distillation for Large Language Model
von: Liu, Jun, et al.
Veröffentlicht: (2025)