Characterization and Mitigation of Training Instabilities in Microscaling Formats
Fuente:
arXiv
Salvato in:
| Autori principali: | Su, Huangyuan, Kwun, Mujin, Gil, Stephanie, Kakade, Sham, Anand, Nikhil |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LOTION: Smoothing the Optimization Landscape for Quantized Training
di: Kwun, Mujin, et al.
Pubblicazione: (2025)
di: Kwun, Mujin, et al.
Pubblicazione: (2025)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
di: Lee, Jungi, et al.
Pubblicazione: (2025)
di: Lee, Jungi, et al.
Pubblicazione: (2025)
Is Finer Better? The Limits of Microscaling Formats in Large Language Models
di: Fasoli, Andrea, et al.
Pubblicazione: (2026)
di: Fasoli, Andrea, et al.
Pubblicazione: (2026)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
di: Koo, Jahyun, et al.
Pubblicazione: (2024)
di: Koo, Jahyun, et al.
Pubblicazione: (2024)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
di: Cheng, Jianyi, et al.
Pubblicazione: (2023)
di: Cheng, Jianyi, et al.
Pubblicazione: (2023)
ElfCore: A 28nm Neural Processor Enabling Dynamic Structured Sparse Training and Online Self-Supervised Learning with Activity-Dependent Weight Update
di: Su, Zhe, et al.
Pubblicazione: (2025)
di: Su, Zhe, et al.
Pubblicazione: (2025)
Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
di: Cuyckens, Stef, et al.
Pubblicazione: (2025)
di: Cuyckens, Stef, et al.
Pubblicazione: (2025)
Refining Datapath for Microscaling ViTs
di: Xiao, Can, et al.
Pubblicazione: (2025)
di: Xiao, Can, et al.
Pubblicazione: (2025)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
di: Wipfli, Max, et al.
Pubblicazione: (2026)
di: Wipfli, Max, et al.
Pubblicazione: (2026)
M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
di: Hu, Weiming, et al.
Pubblicazione: (2026)
di: Hu, Weiming, et al.
Pubblicazione: (2026)
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
di: Park, Dahoon, et al.
Pubblicazione: (2026)
di: Park, Dahoon, et al.
Pubblicazione: (2026)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
di: Bhattacharya, Swastik, et al.
Pubblicazione: (2025)
di: Bhattacharya, Swastik, et al.
Pubblicazione: (2025)
Leveraging Stochastic Depth Training for Adaptive Inference
di: Korol, Guilherme, et al.
Pubblicazione: (2025)
di: Korol, Guilherme, et al.
Pubblicazione: (2025)
Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
di: Banasik, Spencer
Pubblicazione: (2025)
di: Banasik, Spencer
Pubblicazione: (2025)
AxMoE: Characterizing the Impact of Approximate Multipliers on Mixture-of-Experts DNN Architectures
di: Shende, Omkar B, et al.
Pubblicazione: (2026)
di: Shende, Omkar B, et al.
Pubblicazione: (2026)
Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
di: Baruah, Trinayan, et al.
Pubblicazione: (2025)
di: Baruah, Trinayan, et al.
Pubblicazione: (2025)
Orion: Characterizing and Programming Apple's Neural Engine for LLM Training and Inference
di: Kumaresan, Ramchand
Pubblicazione: (2026)
di: Kumaresan, Ramchand
Pubblicazione: (2026)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
di: Chen, Yuzong, et al.
Pubblicazione: (2025)
di: Chen, Yuzong, et al.
Pubblicazione: (2025)
Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
di: Mao, Gang, et al.
Pubblicazione: (2025)
di: Mao, Gang, et al.
Pubblicazione: (2025)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2024)
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2024)
Performance Analysis of DNN Inference/Training with Convolution and non-Convolution Operations
di: Esmaeilzadeh, Hadi, et al.
Pubblicazione: (2023)
di: Esmaeilzadeh, Hadi, et al.
Pubblicazione: (2023)
Cell Library Characterization for Composite Current Source Models Based on Gaussian Process Regression and Active Learning
di: Bai, Tao, et al.
Pubblicazione: (2025)
di: Bai, Tao, et al.
Pubblicazione: (2025)
SNIP: An Adaptive Mixed Precision Framework for Subbyte Large Language Model Training
di: Pan, Yunjie, et al.
Pubblicazione: (2026)
di: Pan, Yunjie, et al.
Pubblicazione: (2026)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
di: Nikolić, Miloš, et al.
Pubblicazione: (2022)
di: Nikolić, Miloš, et al.
Pubblicazione: (2022)
DAISM: Digital Approximate In-SRAM Multiplier-based Accelerator for DNN Training and Inference
di: Sonnino, Lorenzo, et al.
Pubblicazione: (2023)
di: Sonnino, Lorenzo, et al.
Pubblicazione: (2023)
Characterizing and Understanding HGNN Training on GPUs
di: Han, Dengke, et al.
Pubblicazione: (2024)
di: Han, Dengke, et al.
Pubblicazione: (2024)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
di: Mueller, Lion, et al.
Pubblicazione: (2025)
di: Mueller, Lion, et al.
Pubblicazione: (2025)
Decentor-V: Lightweight ML Training on Low-Power RISC-V Edge Devices
di: Ribeiro, Marcelo, et al.
Pubblicazione: (2025)
di: Ribeiro, Marcelo, et al.
Pubblicazione: (2025)
Exploring the Limitations of Kolmogorov-Arnold Networks in Classification: Insights to Software Training and Hardware Implementation
di: Tran, Van Duy, et al.
Pubblicazione: (2024)
di: Tran, Van Duy, et al.
Pubblicazione: (2024)
SetupKit: Efficient Multi-Corner Setup/Hold Time Characterization Using Bias-Enhanced Interpolation and Active Learning
di: Zhou, Junzhuo, et al.
Pubblicazione: (2025)
di: Zhou, Junzhuo, et al.
Pubblicazione: (2025)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
di: Killian, Earl
Pubblicazione: (2026)
di: Killian, Earl
Pubblicazione: (2026)
TsetlinWiSARD: On-Chip Training of Weightless Neural Networks using Tsetlin Automata on FPGAs
di: Duan, Shengyu, et al.
Pubblicazione: (2026)
di: Duan, Shengyu, et al.
Pubblicazione: (2026)
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
di: Wang, Run, et al.
Pubblicazione: (2026)
di: Wang, Run, et al.
Pubblicazione: (2026)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
di: Wu, Qizhe, et al.
Pubblicazione: (2024)
di: Wu, Qizhe, et al.
Pubblicazione: (2024)
FRED: Flexible REduction-Distribution Interconnect and Communication Implementation for Wafer-Scale Distributed Training of DNN Models
di: Rashidi, Saeed, et al.
Pubblicazione: (2024)
di: Rashidi, Saeed, et al.
Pubblicazione: (2024)
Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System
di: Jang, Hongsun, et al.
Pubblicazione: (2024)
di: Jang, Hongsun, et al.
Pubblicazione: (2024)
FiCABU: A Fisher-Based, Context-Adaptive Machine Unlearning Processor for Edge AI
di: Cho, Eun-Su, et al.
Pubblicazione: (2025)
di: Cho, Eun-Su, et al.
Pubblicazione: (2025)
PowerGenie: Analytically-Guided Evolutionary Discovery of Superior Reconfigurable Power Converters
di: Gao, Jian, et al.
Pubblicazione: (2026)
di: Gao, Jian, et al.
Pubblicazione: (2026)
Documenti analoghi
-
LOTION: Smoothing the Optimization Landscape for Quantized Training
di: Kwun, Mujin, et al.
Pubblicazione: (2025) -
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
di: Lee, Jungi, et al.
Pubblicazione: (2025) -
Is Finer Better? The Limits of Microscaling Formats in Large Language Models
di: Fasoli, Andrea, et al.
Pubblicazione: (2026) -
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
di: Koo, Jahyun, et al.
Pubblicazione: (2024) -
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)