AMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
Fuente:
arXiv
Guardado en:
| Autores principales: | Lv, Mengtao, Zhu, Ruiqi, Wang, Xinyu, Li, Yun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
por: Tang, Xuan, et al.
Publicado: (2025)
por: Tang, Xuan, et al.
Publicado: (2025)
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
por: Lo, Yun-Chen, et al.
Publicado: (2024)
por: Lo, Yun-Chen, et al.
Publicado: (2024)
Achieving binary weight and activation for LLMs using Post-Training Quantization
por: Song, Siqing, et al.
Publicado: (2025)
por: Song, Siqing, et al.
Publicado: (2025)
QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
por: Zhao, Zhixiong, et al.
Publicado: (2025)
por: Zhao, Zhixiong, et al.
Publicado: (2025)
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
por: Yun, Juyoung, et al.
Publicado: (2023)
por: Yun, Juyoung, et al.
Publicado: (2023)
Towards Superior Quantization Accuracy: A Layer-sensitive Approach
por: Zhang, Feng, et al.
Publicado: (2025)
por: Zhang, Feng, et al.
Publicado: (2025)
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
por: An, Selim, et al.
Publicado: (2026)
por: An, Selim, et al.
Publicado: (2026)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
por: Zhao, Zhixiong, et al.
Publicado: (2025)
por: Zhao, Zhixiong, et al.
Publicado: (2025)
VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation
por: Zhai, Linwei, et al.
Publicado: (2026)
por: Zhai, Linwei, et al.
Publicado: (2026)
AMSbench: A Comprehensive Benchmark for Evaluating MLLM Capabilities in AMS Circuits
por: Shi, Yichen, et al.
Publicado: (2025)
por: Shi, Yichen, et al.
Publicado: (2025)
ASER: Activation Smoothing and Error Reconstruction for Large Language Model Quantization
por: Zhao, Weibo, et al.
Publicado: (2024)
por: Zhao, Weibo, et al.
Publicado: (2024)
Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized Environments
por: Qu, Yun, et al.
Publicado: (2025)
por: Qu, Yun, et al.
Publicado: (2025)
TAH-QUANT: Effective Activation Quantization in Pipeline Parallelism over Slow Network
por: He, Guangxin, et al.
Publicado: (2025)
por: He, Guangxin, et al.
Publicado: (2025)
Saliency-Aware Regularized Quantization Calibration for Large Language Models
por: Zhao, Yanlong, et al.
Publicado: (2026)
por: Zhao, Yanlong, et al.
Publicado: (2026)
AIS: Adaptive Importance Sampling for Quantized RL
por: Zhou, Jiajun, et al.
Publicado: (2026)
por: Zhou, Jiajun, et al.
Publicado: (2026)
Enhancing Model Privacy in Federated Learning with Random Masking and Quantization
por: Xu, Zhibo, et al.
Publicado: (2025)
por: Xu, Zhibo, et al.
Publicado: (2025)
Adaptive Policy Backbone via Shared Network
por: Park, Bumgeun, et al.
Publicado: (2025)
por: Park, Bumgeun, et al.
Publicado: (2025)
BeamVQ: Beam Search with Vector Quantization to Mitigate Data Scarcity in Physical Spatiotemporal Forecasting
por: Wang, Weiyan, et al.
Publicado: (2025)
por: Wang, Weiyan, et al.
Publicado: (2025)
Revealing the Attention Floating Mechanism in Masked Diffusion Models
por: Dai, Xin, et al.
Publicado: (2026)
por: Dai, Xin, et al.
Publicado: (2026)
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
por: Wang, Guoan, et al.
Publicado: (2026)
por: Wang, Guoan, et al.
Publicado: (2026)
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering
por: Cai, Yuzhu, et al.
Publicado: (2026)
por: Cai, Yuzhu, et al.
Publicado: (2026)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
por: Li, Cheng, et al.
Publicado: (2025)
por: Li, Cheng, et al.
Publicado: (2025)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
por: Bouzouad, Meriem, et al.
Publicado: (2026)
por: Bouzouad, Meriem, et al.
Publicado: (2026)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
por: Gernigon, Cédric, et al.
Publicado: (2024)
por: Gernigon, Cédric, et al.
Publicado: (2024)
HiFloat4 Format for Language Model Inference
por: Luo, Yuanyong, et al.
Publicado: (2026)
por: Luo, Yuanyong, et al.
Publicado: (2026)
Adapformer: Adaptive Channel Management for Multivariate Time Series Forecasting
por: Luo, Yuchen, et al.
Publicado: (2025)
por: Luo, Yuchen, et al.
Publicado: (2025)
HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization
por: Zagitov, Artur, et al.
Publicado: (2026)
por: Zagitov, Artur, et al.
Publicado: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
por: Wang, Yixuan, et al.
Publicado: (2025)
por: Wang, Yixuan, et al.
Publicado: (2025)
Theory-optimal Quantization Based on Flatness
por: Huang, Xiusheng, et al.
Publicado: (2026)
por: Huang, Xiusheng, et al.
Publicado: (2026)
Why Do Some Inputs Break Low-Bit LLM Quantization?
por: Chang, Ting-Yun, et al.
Publicado: (2025)
por: Chang, Ting-Yun, et al.
Publicado: (2025)
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
por: Chiang, Hung-Yueh, et al.
Publicado: (2025)
por: Chiang, Hung-Yueh, et al.
Publicado: (2025)
Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing
por: Filippova, Anastasiia, et al.
Publicado: (2026)
por: Filippova, Anastasiia, et al.
Publicado: (2026)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
Quantum-Classical Hybrid Quantized Neural Network
por: Li, Wenxin, et al.
Publicado: (2025)
por: Li, Wenxin, et al.
Publicado: (2025)
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
por: Zhang, Jinhao, et al.
Publicado: (2025)
por: Zhang, Jinhao, et al.
Publicado: (2025)
Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
por: Varshney, Ayush K., et al.
Publicado: (2026)
por: Varshney, Ayush K., et al.
Publicado: (2026)
APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding
por: Yang, Xinyu, et al.
Publicado: (2025)
por: Yang, Xinyu, et al.
Publicado: (2025)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
por: Zhou, Sifan, et al.
Publicado: (2025)
por: Zhou, Sifan, et al.
Publicado: (2025)
Adaptive Neighborhood-Constrained Q Learning for Offline Reinforcement Learning
por: Mao, Yixiu, et al.
Publicado: (2025)
por: Mao, Yixiu, et al.
Publicado: (2025)
Graph Negative Feedback Bias Correction Framework for Adaptive Heterophily Modeling
por: Lv, Jiaqi, et al.
Publicado: (2026)
por: Lv, Jiaqi, et al.
Publicado: (2026)
Ejemplares similares
-
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
por: Tang, Xuan, et al.
Publicado: (2025) -
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
por: Lo, Yun-Chen, et al.
Publicado: (2024) -
Achieving binary weight and activation for LLMs using Post-Training Quantization
por: Song, Siqing, et al.
Publicado: (2025) -
QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
por: Zhao, Zhixiong, et al.
Publicado: (2025) -
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
por: Yun, Juyoung, et al.
Publicado: (2023)