A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kong, Jason, Pandey, Nilesh Prasad, Ponzina, Flavio, Rosing, Tajana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MicroHD: An Accuracy-Driven Optimization of Hyperdimensional Computing Algorithms for TinyML systems
von: Ponzina, Flavio, et al.
Veröffentlicht: (2024)
von: Ponzina, Flavio, et al.
Veröffentlicht: (2024)
QMC: Efficient SLM Edge Inference via Outlier-Aware Quantization and Emergent Memories Co-Design
von: Pandey, Nilesh Prasad, et al.
Veröffentlicht: (2026)
von: Pandey, Nilesh Prasad, et al.
Veröffentlicht: (2026)
FaTRQ: Tiered Residual Quantization for LLM Vector Search in Far-Memory-Aware ANNS Systems
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
DPQ-HD: Post-Training Compression for Ultra-Low Power Hyperdimensional Computing
von: Pandey, Nilesh Prasad, et al.
Veröffentlicht: (2025)
von: Pandey, Nilesh Prasad, et al.
Veröffentlicht: (2025)
E-QUARTIC: Energy Efficient Edge Ensemble of Convolutional Neural Networks for Resource-Optimized Learning
von: Zhang, Le, et al.
Veröffentlicht: (2024)
von: Zhang, Le, et al.
Veröffentlicht: (2024)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
von: Hu, Lanxiang, et al.
Veröffentlicht: (2024)
von: Hu, Lanxiang, et al.
Veröffentlicht: (2024)
Mem-Rec: Memory Efficient Recommendation System using Alternative Representation
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2023)
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2023)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
von: Federici, Marco, et al.
Veröffentlicht: (2025)
von: Federici, Marco, et al.
Veröffentlicht: (2025)
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
von: Zhao, Daniel, et al.
Veröffentlicht: (2025)
von: Zhao, Daniel, et al.
Veröffentlicht: (2025)
Mixed-Precision Quantization for Language Models: Techniques and Prospects
von: Rakka, Mariam, et al.
Veröffentlicht: (2025)
von: Rakka, Mariam, et al.
Veröffentlicht: (2025)
Oh! We Freeze: Improving Quantized Knowledge Distillation via Signal Propagation Analysis for Large Language Models
von: Bhardwaj, Kartikeya, et al.
Veröffentlicht: (2024)
von: Bhardwaj, Kartikeya, et al.
Veröffentlicht: (2024)
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
von: Liu, Wenyuan, et al.
Veröffentlicht: (2025)
von: Liu, Wenyuan, et al.
Veröffentlicht: (2025)
FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
von: Lee, You Hak, et al.
Veröffentlicht: (2025)
von: Lee, You Hak, et al.
Veröffentlicht: (2025)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
Divide and Learn: Multi-Objective Combinatorial Optimization at Scale
von: Singh, Esha, et al.
Veröffentlicht: (2026)
von: Singh, Esha, et al.
Veröffentlicht: (2026)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
von: Bouzouad, Meriem, et al.
Veröffentlicht: (2026)
von: Bouzouad, Meriem, et al.
Veröffentlicht: (2026)
Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
von: Yao, Wei, et al.
Veröffentlicht: (2025)
von: Yao, Wei, et al.
Veröffentlicht: (2025)
Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
von: Zheng, Xinzhe, et al.
Veröffentlicht: (2025)
von: Zheng, Xinzhe, et al.
Veröffentlicht: (2025)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
von: Lee, Sangjun, et al.
Veröffentlicht: (2025)
von: Lee, Sangjun, et al.
Veröffentlicht: (2025)
SpANNS: Optimizing Approximate Nearest Neighbor Search for Sparse Vectors Using Near Memory Processing
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
von: Gautam, Arpit Singh, et al.
Veröffentlicht: (2026)
von: Gautam, Arpit Singh, et al.
Veröffentlicht: (2026)
Global-Lens Transformers: Adaptive Token Mixing for Dynamic Link Prediction
von: Zou, Tao, et al.
Veröffentlicht: (2025)
von: Zou, Tao, et al.
Veröffentlicht: (2025)
Decision MetaMamba: Enhancing Selective SSM in Offline RL with Heterogeneous Sequence Mixing
von: Kim, Wall, et al.
Veröffentlicht: (2026)
von: Kim, Wall, et al.
Veröffentlicht: (2026)
Decision MetaMamba: Enhancing Selective SSM in Offline RL with Heterogeneous Sequence Mixing
von: Kim, Wall, et al.
Veröffentlicht: (2024)
von: Kim, Wall, et al.
Veröffentlicht: (2024)
Forward Only Learning for Orthogonal Neural Networks of any Depth
von: Caillon, Paul, et al.
Veröffentlicht: (2025)
von: Caillon, Paul, et al.
Veröffentlicht: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
von: Li, Xing, et al.
Veröffentlicht: (2025)
von: Li, Xing, et al.
Veröffentlicht: (2025)
Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
von: Varshney, Ayush K., et al.
Veröffentlicht: (2026)
von: Varshney, Ayush K., et al.
Veröffentlicht: (2026)
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
von: Zhang, Shu-Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Shu-Hao, et al.
Veröffentlicht: (2026)
Forward Target Propagation: A Forward-Only Approach to Global Error Credit Assignment via Local Losses
von: As-Saquib, Nazmus Saadat, et al.
Veröffentlicht: (2025)
von: As-Saquib, Nazmus Saadat, et al.
Veröffentlicht: (2025)
EVA-0: Test-Time Model Evolution with Only Two Forward Passes per Sample
von: Chen, Guohao, et al.
Veröffentlicht: (2026)
von: Chen, Guohao, et al.
Veröffentlicht: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
Intrinsic Structure as a Proxy for Saliency: SVD-Based Weight Preservation for Mixed-Precision Quantization in Large Language Models
von: Landge, Shashank, et al.
Veröffentlicht: (2025)
von: Landge, Shashank, et al.
Veröffentlicht: (2025)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
von: Lee, Seoungsub, et al.
Veröffentlicht: (2026)
von: Lee, Seoungsub, et al.
Veröffentlicht: (2026)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
von: Guan, Ziyi, et al.
Veröffentlicht: (2024)
von: Guan, Ziyi, et al.
Veröffentlicht: (2024)
TS-OOD: Evaluating Time-Series Out-of-Distribution Detection and Prospective Directions for Progress
von: Gungor, Onat, et al.
Veröffentlicht: (2025)
von: Gungor, Onat, et al.
Veröffentlicht: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
SynEHRgy: Synthesizing Mixed-Type Structured Electronic Health Records using Decoder-Only Transformers
von: Karami, Hojjat, et al.
Veröffentlicht: (2024)
von: Karami, Hojjat, et al.
Veröffentlicht: (2024)
Every Bit Counts: A Theoretical Study of Precision-Expressivity Tradeoffs in Quantized Transformers
von: Chakrabarti, Sayak, et al.
Veröffentlicht: (2026)
von: Chakrabarti, Sayak, et al.
Veröffentlicht: (2026)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026)
von: Zhao, Qingyue, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MicroHD: An Accuracy-Driven Optimization of Hyperdimensional Computing Algorithms for TinyML systems
von: Ponzina, Flavio, et al.
Veröffentlicht: (2024) -
QMC: Efficient SLM Edge Inference via Outlier-Aware Quantization and Emergent Memories Co-Design
von: Pandey, Nilesh Prasad, et al.
Veröffentlicht: (2026) -
FaTRQ: Tiered Residual Quantization for LLM Vector Search in Far-Memory-Aware ANNS Systems
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026) -
DPQ-HD: Post-Training Compression for Ultra-Low Power Hyperdimensional Computing
von: Pandey, Nilesh Prasad, et al.
Veröffentlicht: (2025) -
E-QUARTIC: Energy Efficient Edge Ensemble of Convolutional Neural Networks for Resource-Optimized Learning
von: Zhang, Le, et al.
Veröffentlicht: (2024)