QuantuneV2: Compiler-Based Local Metric-Driven Mixed Precision Quantization for Practical Embedded AI Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jeongseok, Lee, Jemin, Kwon, Yongin, Kim, Daeyoung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixed Non-linear Quantization for Vision Transformers
by: Kim, Gihwan, et al.
Published: (2024)
by: Kim, Gihwan, et al.
Published: (2024)
QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
by: Oh, Sehyeon, et al.
Published: (2026)
by: Oh, Sehyeon, et al.
Published: (2026)
Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to Giant
by: Lee, Jemin, et al.
Published: (2024)
by: Lee, Jemin, et al.
Published: (2024)
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
by: Lee, Jemin, et al.
Published: (2023)
by: Lee, Jemin, et al.
Published: (2023)
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
by: Kim, Gihwan, et al.
Published: (2025)
by: Kim, Gihwan, et al.
Published: (2025)
Intent-Driven UAM Rescheduling
by: Kim, Jeongseok, et al.
Published: (2025)
by: Kim, Jeongseok, et al.
Published: (2025)
Dialogue Possibilities between a Human Supervisor and UAM Air Traffic Management: Route Alteration
by: Kim, Jeongseok, et al.
Published: (2023)
by: Kim, Jeongseok, et al.
Published: (2023)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
by: Lee, Seoungsub, et al.
Published: (2026)
by: Lee, Seoungsub, et al.
Published: (2026)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
LungCRCT: Causal Representation based Lung CT Processing for Lung Cancer Treatment
by: Kim, Daeyoung
Published: (2026)
by: Kim, Daeyoung
Published: (2026)
LightHCG: a Lightweight yet powerful HSIC Disentanglement based Causal Glaucoma Detection Model framework
by: Kim, Daeyoung
Published: (2025)
by: Kim, Daeyoung
Published: (2025)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
A Metric Driven Approach to Mixed Precision Training
by: Rasquinha, Mitchelle, et al.
Published: (2024)
by: Rasquinha, Mitchelle, et al.
Published: (2024)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
by: Kim, Taeho, et al.
Published: (2024)
by: Kim, Taeho, et al.
Published: (2024)
MuLMINet: Multi-Layer Multi-Input Transformer Network with Weighted Loss
by: Seong, Minwoo, et al.
Published: (2023)
by: Seong, Minwoo, et al.
Published: (2023)
PARAN: Persona-Augmented Review ANswering system on Food Delivery Review Dataset
by: Park, Moonsoo, et al.
Published: (2025)
by: Park, Moonsoo, et al.
Published: (2025)
Attribute Based Interpretable Evaluation Metrics for Generative Models
by: Kim, Dongkyun, et al.
Published: (2023)
by: Kim, Dongkyun, et al.
Published: (2023)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
Mixed-Precision Quantization for Language Models: Techniques and Prospects
by: Rakka, Mariam, et al.
Published: (2025)
by: Rakka, Mariam, et al.
Published: (2025)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
Learning Quadrupedal Locomotion for a Heavy Hydraulic Robot Using an Actuator Model
by: Lee, Minho, et al.
Published: (2026)
by: Lee, Minho, et al.
Published: (2026)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
by: Bouzouad, Meriem, et al.
Published: (2026)
by: Bouzouad, Meriem, et al.
Published: (2026)
Channel-Wise Mixed-Precision Quantization for Large Language Models
by: Chen, Zihan, et al.
Published: (2024)
by: Chen, Zihan, et al.
Published: (2024)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
by: Lee, Joonhyung, et al.
Published: (2024)
by: Lee, Joonhyung, et al.
Published: (2024)
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
by: Liu, Wenyuan, et al.
Published: (2025)
by: Liu, Wenyuan, et al.
Published: (2025)
Not Only Rewards But Also Constraints: Applications on Legged Robot Locomotion
by: Kim, Yunho, et al.
Published: (2023)
by: Kim, Yunho, et al.
Published: (2023)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
Local Temporal Feature Enhanced Transformer with ROI-rank Based Masking for Diagnosis of ADHD
by: Kim, Byunggun, et al.
Published: (2025)
by: Kim, Byunggun, et al.
Published: (2025)
Legged Robot State Estimation With Invariant Extended Kalman Filter Using Neural Measurement Network
by: Youm, Donghoon, et al.
Published: (2024)
by: Youm, Donghoon, et al.
Published: (2024)
Learning Rapid Turning, Aerial Reorientation, and Balancing using Manipulator as a Tail
by: Yang, Insung, et al.
Published: (2024)
by: Yang, Insung, et al.
Published: (2024)
ML$^2$Tuner: Efficient Code Tuning via Multi-Level Machine Learning Models
by: Cha, JooHyoung, et al.
Published: (2024)
by: Cha, JooHyoung, et al.
Published: (2024)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
by: Xu, Haoning, et al.
Published: (2025)
by: Xu, Haoning, et al.
Published: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
Learning Semantic Traversability with Egocentric Video and Automated Annotation Strategy
by: Kim, Yunho, et al.
Published: (2024)
by: Kim, Yunho, et al.
Published: (2024)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
by: Gautam, Arpit Singh, et al.
Published: (2026)
by: Gautam, Arpit Singh, et al.
Published: (2026)
Visual Preference Inference: An Image Sequence-Based Preference Reasoning in Tabletop Object Manipulation
by: Lee, Joonhyung, et al.
Published: (2024)
by: Lee, Joonhyung, et al.
Published: (2024)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Similar Items
-
Mixed Non-linear Quantization for Vision Transformers
by: Kim, Gihwan, et al.
Published: (2024) -
QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
by: Oh, Sehyeon, et al.
Published: (2026) -
Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to Giant
by: Lee, Jemin, et al.
Published: (2024) -
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
by: Lee, Jemin, et al.
Published: (2023) -
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
by: Kim, Gihwan, et al.
Published: (2025)