Is Finer Better? The Limits of Microscaling Formats in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fasoli, Andrea, Kar, Monodeep, Liu, Chi-Chun, Venkataramani, Swagath, Srinivasan, Viji, Chang, Leland, Wang, Naigang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation
by: Chen, Mu-Chi, et al.
Published: (2026)
by: Chen, Mu-Chi, et al.
Published: (2026)
OpenEye: A Scalable Open-Source Hardware Accelerator for DNNs
by: Lebold, Denis, et al.
Published: (2026)
by: Lebold, Denis, et al.
Published: (2026)
SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation
by: Chen, Mu-Chi, et al.
Published: (2026)
by: Chen, Mu-Chi, et al.
Published: (2026)
CPU Simulation with Ranked Set Sampling and Repeated Subsampling
by: Ekman, Magnus
Published: (2026)
by: Ekman, Magnus
Published: (2026)
CPU Simulation Using Two-Phase Stratified Sampling
by: Ekman, Magnus
Published: (2026)
by: Ekman, Magnus
Published: (2026)
CORE: Constraint-Aware One-Step Reinforcement Learning for Simulation-Guided Neural Network Accelerator Design
by: Xiao, Yifeng, et al.
Published: (2025)
by: Xiao, Yifeng, et al.
Published: (2025)
Biological Intuition on Digital Hardware: An RTL Implementation of Poisson-Encoded SNNs for Static Image Classification
by: Das, Debabrata, et al.
Published: (2026)
by: Das, Debabrata, et al.
Published: (2026)
Photonic AI: A Hybrid Diffractive Holographic Neural System for Passive Optical Real-Time Image Classification
by: Hiremath, Prakul Sunil
Published: (2026)
by: Hiremath, Prakul Sunil
Published: (2026)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
by: Lee, Jungi, et al.
Published: (2025)
by: Lee, Jungi, et al.
Published: (2025)
Machine Learning for Energy-Performance-aware Scheduling
by: Hu, Zheyuan, et al.
Published: (2026)
by: Hu, Zheyuan, et al.
Published: (2026)
CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
by: Goswami, Kaustav, et al.
Published: (2026)
by: Goswami, Kaustav, et al.
Published: (2026)
parti-gem5: gem5's Timing Mode Parallelised
by: Cubero-Cascante, José, et al.
Published: (2023)
by: Cubero-Cascante, José, et al.
Published: (2023)
GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition
by: Li, Peijing, et al.
Published: (2025)
by: Li, Peijing, et al.
Published: (2025)
Nonvolatile Charge-Domain Attention with HZO Ferroelectric Capacitors: A Simulation-Based Device-to-System Evaluation
by: Abouagour, Faris
Published: (2026)
by: Abouagour, Faris
Published: (2026)
FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
by: Parameshwara, Arya, et al.
Published: (2025)
by: Parameshwara, Arya, et al.
Published: (2025)
ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration
by: Eryilmaz, Cagri
Published: (2026)
by: Eryilmaz, Cagri
Published: (2026)
Accelerating Recommendation System Training by Leveraging Popular Choices
by: Adnan, Muhammad, et al.
Published: (2021)
by: Adnan, Muhammad, et al.
Published: (2021)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
by: Cheng, Jianyi, et al.
Published: (2023)
by: Cheng, Jianyi, et al.
Published: (2023)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
EM-Aware Physical Synthesis: Neural Inductor Modeling and Intelligent Placement & Routing for RF Circuits
by: Huang, Yilun, et al.
Published: (2026)
by: Huang, Yilun, et al.
Published: (2026)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
by: Su, Huangyuan, et al.
Published: (2025)
by: Su, Huangyuan, et al.
Published: (2025)
Refining Datapath for Microscaling ViTs
by: Xiao, Can, et al.
Published: (2025)
by: Xiao, Can, et al.
Published: (2025)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
by: Wipfli, Max, et al.
Published: (2026)
by: Wipfli, Max, et al.
Published: (2026)
Transaction Level Hierarchy Guided and Functional Coverage Driven Deductive Formal Verification
by: Strauch, Tobias
Published: (2025)
by: Strauch, Tobias
Published: (2025)
M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
by: Hu, Weiming, et al.
Published: (2026)
by: Hu, Weiming, et al.
Published: (2026)
NotSoTiny: A Large, Living Benchmark for RTL Code Generation
by: Ghorab, Razine Moundir, et al.
Published: (2025)
by: Ghorab, Razine Moundir, et al.
Published: (2025)
Deep Learning-Based Early-Stage IR-Drop Estimation via CNN Surrogate Modeling
by: Bhadana, Ritesh
Published: (2026)
by: Bhadana, Ritesh
Published: (2026)
Chameleon: A MatMul-Free Temporal Convolutional Network Accelerator for End-to-End Few-Shot and Continual Learning from Sequential Data
by: Blanken, Douwe den, et al.
Published: (2025)
by: Blanken, Douwe den, et al.
Published: (2025)
Differentiable Logic Synthesis: Spectral Coefficient Selection via Sinkhorn-Constrained Composition
by: Pavlov, Gorgi
Published: (2026)
by: Pavlov, Gorgi
Published: (2026)
TuRTLe: A Unified Evaluation of LLMs for RTL Generation
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
by: Park, Dahoon, et al.
Published: (2026)
by: Park, Dahoon, et al.
Published: (2026)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
by: Karfakis, George, et al.
Published: (2026)
by: Karfakis, George, et al.
Published: (2026)
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
by: İslamoğlu, Gamze, et al.
Published: (2025)
by: İslamoğlu, Gamze, et al.
Published: (2025)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
by: Koo, Jahyun, et al.
Published: (2024)
by: Koo, Jahyun, et al.
Published: (2024)
SR-NCL: an Area-/Energy-Efficient Resilient NCL Architecture Based on Selective Redundancy
by: Ziad, Hasnain A., et al.
Published: (2025)
by: Ziad, Hasnain A., et al.
Published: (2025)
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
by: Aspros, Maxime Henri, et al.
Published: (2025)
by: Aspros, Maxime Henri, et al.
Published: (2025)
SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization
by: Xie, Rui, et al.
Published: (2024)
by: Xie, Rui, et al.
Published: (2024)
The Monte Carlo Method and New Device and Architectural Techniques for Accelerating It
by: Petangoda, Janith, et al.
Published: (2025)
by: Petangoda, Janith, et al.
Published: (2025)
Similar Items
-
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation
by: Chen, Mu-Chi, et al.
Published: (2026) -
OpenEye: A Scalable Open-Source Hardware Accelerator for DNNs
by: Lebold, Denis, et al.
Published: (2026) -
SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation
by: Chen, Mu-Chi, et al.
Published: (2026) -
CPU Simulation with Ranked Set Sampling and Repeated Subsampling
by: Ekman, Magnus
Published: (2026) -
CPU Simulation Using Two-Phase Stratified Sampling
by: Ekman, Magnus
Published: (2026)