FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Jaemin, Um, Hongjun, Kim, Sungkyun, Park, Yongjun, Seo, Jiwon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
por: Kim, Jaemin, et al.
Publicado: (2026)
por: Kim, Jaemin, et al.
Publicado: (2026)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
por: Oh, Hyungjun, et al.
Publicado: (2024)
por: Oh, Hyungjun, et al.
Publicado: (2024)
PINT: Physics-Informed Neural Time Series Models with Applications to Long-term Inference on WeatherBench 2m-Temperature Data
por: Park, Keonvin, et al.
Publicado: (2025)
por: Park, Keonvin, et al.
Publicado: (2025)
Coupling-Robust Accuracy in Multiphysics Physics Informed Neural Networks via Kronecker-Preconditioned Optimization
por: Park, Youngjae, et al.
Publicado: (2026)
por: Park, Youngjae, et al.
Publicado: (2026)
Design Optimization of Nuclear Fusion Reactor through Deep Reinforcement Learning
por: Kim, Jinsu, et al.
Publicado: (2024)
por: Kim, Jinsu, et al.
Publicado: (2024)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
por: Lee, Seoungsub, et al.
Publicado: (2026)
por: Lee, Seoungsub, et al.
Publicado: (2026)
Efficient Mixed Precision Quantization in Graph Neural Networks
por: Moustafa, Samir, et al.
Publicado: (2025)
por: Moustafa, Samir, et al.
Publicado: (2025)
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization
por: Kim, Daeun, et al.
Publicado: (2025)
por: Kim, Daeun, et al.
Publicado: (2025)
Joint Pruning and Channel-wise Mixed-Precision Quantization for Efficient Deep Neural Networks
por: Motetti, Beatrice Alessandra, et al.
Publicado: (2024)
por: Motetti, Beatrice Alessandra, et al.
Publicado: (2024)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
por: Kurtic, Eldar, et al.
Publicado: (2024)
por: Kurtic, Eldar, et al.
Publicado: (2024)
Fairness-Accuracy Trade-Offs: A Causal Perspective
por: Plecko, Drago, et al.
Publicado: (2024)
por: Plecko, Drago, et al.
Publicado: (2024)
InfoQ: Mixed-Precision Quantization via Global Information Flow
por: Akbulut, Mehmet Emre, et al.
Publicado: (2025)
por: Akbulut, Mehmet Emre, et al.
Publicado: (2025)
Accuracy-Efficiency Trade-Offs in Spiking Neural Networks: A Lempel-Ziv Complexity Perspective on Learning Rules
por: Rudnicka, Zofia, et al.
Publicado: (2025)
por: Rudnicka, Zofia, et al.
Publicado: (2025)
MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
por: Kim, Han-Byul, et al.
Publicado: (2023)
por: Kim, Han-Byul, et al.
Publicado: (2023)
CoopQ: Cooperative Game Inspired Layerwise Mixed Precision Quantization for LLMs
por: Zhao, Junchen, et al.
Publicado: (2025)
por: Zhao, Junchen, et al.
Publicado: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
por: Huang, Wei, et al.
Publicado: (2023)
por: Huang, Wei, et al.
Publicado: (2023)
Neural Variance-aware Dueling Bandits with Deep Representation and Shallow Exploration
por: Oh, Youngmin, et al.
Publicado: (2025)
por: Oh, Youngmin, et al.
Publicado: (2025)
Quantization of Spiking Neural Networks Beyond Accuracy
por: Smith, Evan Gibson, et al.
Publicado: (2026)
por: Smith, Evan Gibson, et al.
Publicado: (2026)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
por: Bouzouad, Meriem, et al.
Publicado: (2026)
por: Bouzouad, Meriem, et al.
Publicado: (2026)
Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
por: Kim, Sungkyun, et al.
Publicado: (2025)
por: Kim, Sungkyun, et al.
Publicado: (2025)
OMPQ: Orthogonal Mixed Precision Quantization
por: Ma, Yuexiao, et al.
Publicado: (2021)
por: Ma, Yuexiao, et al.
Publicado: (2021)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
por: Saxena, Utkarsh, et al.
Publicado: (2024)
por: Saxena, Utkarsh, et al.
Publicado: (2024)
Adaptive Guidance for Retrieval-Augmented Masked Diffusion Models
por: Kim, Jaemin, et al.
Publicado: (2026)
por: Kim, Jaemin, et al.
Publicado: (2026)
Dual Precision Deep Neural Network
por: Park, Jae Hyun, et al.
Publicado: (2020)
por: Park, Jae Hyun, et al.
Publicado: (2020)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
por: Yang, June Yong, et al.
Publicado: (2024)
por: Yang, June Yong, et al.
Publicado: (2024)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
The More the Merrier? Navigating Accuracy vs. Energy Efficiency Design Trade-Offs in Ensemble Learning Systems
por: Omar, Rafiullah, et al.
Publicado: (2024)
por: Omar, Rafiullah, et al.
Publicado: (2024)
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation
por: Park, Jaehyun, et al.
Publicado: (2024)
por: Park, Jaehyun, et al.
Publicado: (2024)
Enhancing Accuracy and Parameter-Efficiency of Neural Representations for Network Parameterization
por: Choi, Hongjun, et al.
Publicado: (2024)
por: Choi, Hongjun, et al.
Publicado: (2024)
Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based GNNs for Fluid Dynamics Simulations
por: Seo, Sangwoo, et al.
Publicado: (2025)
por: Seo, Sangwoo, et al.
Publicado: (2025)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
por: Park, Gunho, et al.
Publicado: (2025)
por: Park, Gunho, et al.
Publicado: (2025)
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
por: Gafni, Tomer, et al.
Publicado: (2025)
por: Gafni, Tomer, et al.
Publicado: (2025)
Multi-Resolution End-to-End Deep Neural Network for Optimizing Latency-Accuracy Tradeoff in Autonomous Driving
por: Weng, Qitao, et al.
Publicado: (2026)
por: Weng, Qitao, et al.
Publicado: (2026)
Efficient and Effective Methods for Mixed Precision Neural Network Quantization for Faster, Energy-efficient Inference
por: Bablani, Deepika, et al.
Publicado: (2023)
por: Bablani, Deepika, et al.
Publicado: (2023)
AMED: Automatic Mixed-Precision Quantization for Edge Devices
por: Kimhi, Moshe, et al.
Publicado: (2022)
por: Kimhi, Moshe, et al.
Publicado: (2022)
MoPEQ: Mixture of Mixed Precision Quantized Experts
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
Trade-Offs of Diagonal Fisher Information Matrix Estimators
por: Soen, Alexander, et al.
Publicado: (2024)
por: Soen, Alexander, et al.
Publicado: (2024)
Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
por: Varshney, Ayush K., et al.
Publicado: (2026)
por: Varshney, Ayush K., et al.
Publicado: (2026)
Precision Adaptive Imputation Network : An Unified Technique for Mixed Datasets
por: Joshi, Harsh, et al.
Publicado: (2025)
por: Joshi, Harsh, et al.
Publicado: (2025)
Eau De $Q$-Network: Adaptive Distillation of Neural Networks in Deep Reinforcement Learning
por: Vincent, Théo, et al.
Publicado: (2025)
por: Vincent, Théo, et al.
Publicado: (2025)
Ejemplares similares
-
Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
por: Kim, Jaemin, et al.
Publicado: (2026) -
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
por: Oh, Hyungjun, et al.
Publicado: (2024) -
PINT: Physics-Informed Neural Time Series Models with Applications to Long-term Inference on WeatherBench 2m-Temperature Data
por: Park, Keonvin, et al.
Publicado: (2025) -
Coupling-Robust Accuracy in Multiphysics Physics Informed Neural Networks via Kronecker-Preconditioned Optimization
por: Park, Youngjae, et al.
Publicado: (2026) -
Design Optimization of Nuclear Fusion Reactor through Deep Reinforcement Learning
por: Kim, Jinsu, et al.
Publicado: (2024)