Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
Fuente:
arXiv
Saved in:
| Main Authors: | Dumitru, Razvan-Gabriel, Yadav, Vikas, Maheshwary, Rishabh, Clotan, Paul-Ioan, Madhusudhan, Sathwik Tejaswi, Surdeanu, Mihai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
Exploring Model Invariance with Discrete Search for Ultra-Low-Bit Quantization
by: Wen, Yuqiao, et al.
Published: (2025)
by: Wen, Yuqiao, et al.
Published: (2025)
BitSkip: An Empirical Analysis of Quantization and Early Exit Composition in Transformers
by: Bhuvaneswaran, Ramshankar, et al.
Published: (2025)
by: Bhuvaneswaran, Ramshankar, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
by: Rath, Plawan Kumar, et al.
Published: (2026)
by: Rath, Plawan Kumar, et al.
Published: (2026)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
by: Tiwari, Rishabh, et al.
Published: (2026)
by: Tiwari, Rishabh, et al.
Published: (2026)
Super Apriel: One Checkpoint, Many Speeds
by: Labs, SLAM, et al.
Published: (2026)
by: Labs, SLAM, et al.
Published: (2026)
Enhancing Ultra-Low-Bit Quantization of Large Language Models Through Saliency-Aware Partial Retraining
by: Cao, Deyu, et al.
Published: (2025)
by: Cao, Deyu, et al.
Published: (2025)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Sliced-Wasserstein Distribution Alignment Loss Improves the Ultra-Low-Bit Quantization of Large Language Models
by: Cao, Deyu, et al.
Published: (2026)
by: Cao, Deyu, et al.
Published: (2026)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
S-HR-VQVAE: Sequential Hierarchical Residual Learning Vector Quantized Variational Autoencoder for Video Prediction
by: Adiban, Mohammad, et al.
Published: (2023)
by: Adiban, Mohammad, et al.
Published: (2023)
Quantization-Robust LLM Unlearning via Low-Rank Adaptation
by: Abitante, João Vitor Boer, et al.
Published: (2026)
by: Abitante, João Vitor Boer, et al.
Published: (2026)
Beyond Mimicry: Preference Coherence in LLMs
by: Mikaelson, Luhan, et al.
Published: (2025)
by: Mikaelson, Luhan, et al.
Published: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
by: Gwak, Jiho, et al.
Published: (2025)
by: Gwak, Jiho, et al.
Published: (2025)
Multi-Agent Pathfinding with Non-Unit Integer Edge Costs via Enhanced Conflict-Based Search and Graph Discretization
by: Fan, Hongkai, et al.
Published: (2026)
by: Fan, Hongkai, et al.
Published: (2026)
Fane at SemEval-2025 Task 10: Zero-Shot Entity Framing with Large Language Models
by: Fane, Enfa, et al.
Published: (2025)
by: Fane, Enfa, et al.
Published: (2025)
Fuzzy Norm-Explicit Product Quantization for Recommender Systems
by: Jamalifard, Mohammadreza, et al.
Published: (2024)
by: Jamalifard, Mohammadreza, et al.
Published: (2024)
SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
by: Liu, Han, et al.
Published: (2026)
by: Liu, Han, et al.
Published: (2026)
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
Quantized FCA: Efficient Zero-Shot Texture Anomaly Detection
by: Ardelean, Andrei-Timotei, et al.
Published: (2025)
by: Ardelean, Andrei-Timotei, et al.
Published: (2025)
4OPS: Structural Difficulty Modeling in Integer Arithmetic Puzzles
by: Zeytuncu, Yunus E.
Published: (2026)
by: Zeytuncu, Yunus E.
Published: (2026)
Procedural Game Level Design with Deep Reinforcement Learning
by: Özkan, Miraç Buğra
Published: (2025)
by: Özkan, Miraç Buğra
Published: (2025)
The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
by: Yang, Jisoo, et al.
Published: (2026)
by: Yang, Jisoo, et al.
Published: (2026)
LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB
by: Bekmyradov, Vekil, et al.
Published: (2026)
by: Bekmyradov, Vekil, et al.
Published: (2026)
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
by: Lee, Christine, et al.
Published: (2025)
by: Lee, Christine, et al.
Published: (2025)
Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers
by: Zhao, Yiming
Published: (2026)
by: Zhao, Yiming
Published: (2026)
Federated Multi-Agent Mapping for Planetary Exploration
by: Szatmari, Tiberiu-Ioan, et al.
Published: (2024)
by: Szatmari, Tiberiu-Ioan, et al.
Published: (2024)
Compensate Quantization Errors+: Quantized Models Are Inquisitive Learners
by: Gao, Yifei, et al.
Published: (2024)
by: Gao, Yifei, et al.
Published: (2024)
Equip Pre-ranking with Target Attention by Residual Quantization
by: Li, Yutong, et al.
Published: (2025)
by: Li, Yutong, et al.
Published: (2025)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
by: Costa, Rimom
Published: (2025)
by: Costa, Rimom
Published: (2025)
Deep Policy Iteration with Integer Programming for Inventory Management
by: Harsha, Pavithra, et al.
Published: (2021)
by: Harsha, Pavithra, et al.
Published: (2021)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
by: Khanna, Danush, et al.
Published: (2025)
by: Khanna, Danush, et al.
Published: (2025)
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
by: Dong, Jia-Kai, et al.
Published: (2025)
by: Dong, Jia-Kai, et al.
Published: (2025)
Distributed Multi-Layer Editing for Rule-Level Knowledge in Large Language Models
by: Wang, Yating, et al.
Published: (2026)
by: Wang, Yating, et al.
Published: (2026)
Composing Linear Layers from Irreducibles
by: Pence, Travis, et al.
Published: (2025)
by: Pence, Travis, et al.
Published: (2025)
Similar Items
-
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024) -
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025) -
Enhancing Transformer RNNs with Multiple Temporal Perspectives
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024) -
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025) -
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
by: Patarlapalli, Sai Babu, et al.
Published: (2026)