QuantuneV2: Compiler-Based Local Metric-Driven Mixed Precision Quantization for Practical Embedded AI Applications
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Jeongseok, Lee, Jemin, Kwon, Yongin, Kim, Daeyoung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mixed Non-linear Quantization for Vision Transformers
di: Kim, Gihwan, et al.
Pubblicazione: (2024)
di: Kim, Gihwan, et al.
Pubblicazione: (2024)
QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
di: Oh, Sehyeon, et al.
Pubblicazione: (2026)
di: Oh, Sehyeon, et al.
Pubblicazione: (2026)
Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to Giant
di: Lee, Jemin, et al.
Pubblicazione: (2024)
di: Lee, Jemin, et al.
Pubblicazione: (2024)
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
di: Lee, Jemin, et al.
Pubblicazione: (2023)
di: Lee, Jemin, et al.
Pubblicazione: (2023)
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
di: Kim, Gihwan, et al.
Pubblicazione: (2025)
di: Kim, Gihwan, et al.
Pubblicazione: (2025)
Intent-Driven UAM Rescheduling
di: Kim, Jeongseok, et al.
Pubblicazione: (2025)
di: Kim, Jeongseok, et al.
Pubblicazione: (2025)
Dialogue Possibilities between a Human Supervisor and UAM Air Traffic Management: Route Alteration
di: Kim, Jeongseok, et al.
Pubblicazione: (2023)
di: Kim, Jeongseok, et al.
Pubblicazione: (2023)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
di: Lee, Seoungsub, et al.
Pubblicazione: (2026)
di: Lee, Seoungsub, et al.
Pubblicazione: (2026)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
di: Yang, June Yong, et al.
Pubblicazione: (2024)
di: Yang, June Yong, et al.
Pubblicazione: (2024)
LungCRCT: Causal Representation based Lung CT Processing for Lung Cancer Treatment
di: Kim, Daeyoung
Pubblicazione: (2026)
di: Kim, Daeyoung
Pubblicazione: (2026)
LightHCG: a Lightweight yet powerful HSIC Disentanglement based Causal Glaucoma Detection Model framework
di: Kim, Daeyoung
Pubblicazione: (2025)
di: Kim, Daeyoung
Pubblicazione: (2025)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
A Metric Driven Approach to Mixed Precision Training
di: Rasquinha, Mitchelle, et al.
Pubblicazione: (2024)
di: Rasquinha, Mitchelle, et al.
Pubblicazione: (2024)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
di: Kim, Taeho, et al.
Pubblicazione: (2024)
di: Kim, Taeho, et al.
Pubblicazione: (2024)
MuLMINet: Multi-Layer Multi-Input Transformer Network with Weighted Loss
di: Seong, Minwoo, et al.
Pubblicazione: (2023)
di: Seong, Minwoo, et al.
Pubblicazione: (2023)
PARAN: Persona-Augmented Review ANswering system on Food Delivery Review Dataset
di: Park, Moonsoo, et al.
Pubblicazione: (2025)
di: Park, Moonsoo, et al.
Pubblicazione: (2025)
Attribute Based Interpretable Evaluation Metrics for Generative Models
di: Kim, Dongkyun, et al.
Pubblicazione: (2023)
di: Kim, Dongkyun, et al.
Pubblicazione: (2023)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
di: Hyun, Jeongseok, et al.
Pubblicazione: (2024)
di: Hyun, Jeongseok, et al.
Pubblicazione: (2024)
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
di: Zhao, Tianchen, et al.
Pubblicazione: (2024)
di: Zhao, Tianchen, et al.
Pubblicazione: (2024)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
Mixed-Precision Quantization for Language Models: Techniques and Prospects
di: Rakka, Mariam, et al.
Pubblicazione: (2025)
di: Rakka, Mariam, et al.
Pubblicazione: (2025)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
di: Federici, Marco, et al.
Pubblicazione: (2025)
di: Federici, Marco, et al.
Pubblicazione: (2025)
Learning Quadrupedal Locomotion for a Heavy Hydraulic Robot Using an Actuator Model
di: Lee, Minho, et al.
Pubblicazione: (2026)
di: Lee, Minho, et al.
Pubblicazione: (2026)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
di: Bouzouad, Meriem, et al.
Pubblicazione: (2026)
di: Bouzouad, Meriem, et al.
Pubblicazione: (2026)
Channel-Wise Mixed-Precision Quantization for Large Language Models
di: Chen, Zihan, et al.
Pubblicazione: (2024)
di: Chen, Zihan, et al.
Pubblicazione: (2024)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
di: Liu, Wenyuan, et al.
Pubblicazione: (2025)
di: Liu, Wenyuan, et al.
Pubblicazione: (2025)
Not Only Rewards But Also Constraints: Applications on Legged Robot Locomotion
di: Kim, Yunho, et al.
Pubblicazione: (2023)
di: Kim, Yunho, et al.
Pubblicazione: (2023)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
di: Zhang, Tao, et al.
Pubblicazione: (2025)
di: Zhang, Tao, et al.
Pubblicazione: (2025)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
Local Temporal Feature Enhanced Transformer with ROI-rank Based Masking for Diagnosis of ADHD
di: Kim, Byunggun, et al.
Pubblicazione: (2025)
di: Kim, Byunggun, et al.
Pubblicazione: (2025)
Legged Robot State Estimation With Invariant Extended Kalman Filter Using Neural Measurement Network
di: Youm, Donghoon, et al.
Pubblicazione: (2024)
di: Youm, Donghoon, et al.
Pubblicazione: (2024)
Learning Rapid Turning, Aerial Reorientation, and Balancing using Manipulator as a Tail
di: Yang, Insung, et al.
Pubblicazione: (2024)
di: Yang, Insung, et al.
Pubblicazione: (2024)
ML$^2$Tuner: Efficient Code Tuning via Multi-Level Machine Learning Models
di: Cha, JooHyoung, et al.
Pubblicazione: (2024)
di: Cha, JooHyoung, et al.
Pubblicazione: (2024)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
di: Xu, Haoning, et al.
Pubblicazione: (2025)
di: Xu, Haoning, et al.
Pubblicazione: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
di: Huang, Wei, et al.
Pubblicazione: (2023)
di: Huang, Wei, et al.
Pubblicazione: (2023)
Learning Semantic Traversability with Egocentric Video and Automated Annotation Strategy
di: Kim, Yunho, et al.
Pubblicazione: (2024)
di: Kim, Yunho, et al.
Pubblicazione: (2024)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
di: Gautam, Arpit Singh, et al.
Pubblicazione: (2026)
di: Gautam, Arpit Singh, et al.
Pubblicazione: (2026)
Visual Preference Inference: An Image Sequence-Based Preference Reasoning in Tabletop Object Manipulation
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Mixed Non-linear Quantization for Vision Transformers
di: Kim, Gihwan, et al.
Pubblicazione: (2024) -
QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
di: Oh, Sehyeon, et al.
Pubblicazione: (2026) -
Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to Giant
di: Lee, Jemin, et al.
Pubblicazione: (2024) -
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
di: Lee, Jemin, et al.
Pubblicazione: (2023) -
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
di: Kim, Gihwan, et al.
Pubblicazione: (2025)