M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Weiming, Zhang, Zihan, Zhang, Haoyan, Zhang, Chen, Guo, Cong, Feng, Yu, Hu, Tianchi, Li, Guanglin, Hu, Guipeng, Wang, Junsong, Leng, Jingwen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type
por: Hu, Weiming, et al.
Publicado: (2025)
por: Hu, Weiming, et al.
Publicado: (2025)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
por: Cheng, Jianyi, et al.
Publicado: (2023)
por: Cheng, Jianyi, et al.
Publicado: (2023)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
por: Wipfli, Max, et al.
Publicado: (2026)
por: Wipfli, Max, et al.
Publicado: (2026)
Ascend HiFloat8 Format for Deep Learning
por: Luo, Yuanyong, et al.
Publicado: (2024)
por: Luo, Yuanyong, et al.
Publicado: (2024)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
por: Su, Huangyuan, et al.
Publicado: (2025)
por: Su, Huangyuan, et al.
Publicado: (2025)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
por: Lee, Jungi, et al.
Publicado: (2025)
por: Lee, Jungi, et al.
Publicado: (2025)
HiFloat4 Format for Language Model Inference
por: Luo, Yuanyong, et al.
Publicado: (2026)
por: Luo, Yuanyong, et al.
Publicado: (2026)
Refining Datapath for Microscaling ViTs
por: Xiao, Can, et al.
Publicado: (2025)
por: Xiao, Can, et al.
Publicado: (2025)
KANtize: Exploring Low-bit Quantization of Kolmogorov-Arnold Networks for Efficient Inference
por: Errabii, Sohaib, et al.
Publicado: (2026)
por: Errabii, Sohaib, et al.
Publicado: (2026)
Cicero: Addressing Algorithmic and Architectural Bottlenecks in Neural Rendering by Radiance Warping and Memory Optimizations
por: Feng, Yu, et al.
Publicado: (2024)
por: Feng, Yu, et al.
Publicado: (2024)
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
por: Cuyckens, Stef, et al.
Publicado: (2025)
por: Cuyckens, Stef, et al.
Publicado: (2025)
SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
por: Li, Xingyang, et al.
Publicado: (2025)
por: Li, Xingyang, et al.
Publicado: (2025)
StreamGrid: Streaming Point Cloud Analytics via Compulsory Splitting and Deterministic Termination
por: Feng, Yu, et al.
Publicado: (2025)
por: Feng, Yu, et al.
Publicado: (2025)
Lumina: Real-Time Mobile Neural Rendering by Exploiting Computational Redundancy
por: Feng, Yu, et al.
Publicado: (2025)
por: Feng, Yu, et al.
Publicado: (2025)
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
por: Park, Dahoon, et al.
Publicado: (2026)
por: Park, Dahoon, et al.
Publicado: (2026)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
por: Karfakis, George, et al.
Publicado: (2026)
por: Karfakis, George, et al.
Publicado: (2026)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
por: Koo, Jahyun, et al.
Publicado: (2024)
por: Koo, Jahyun, et al.
Publicado: (2024)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
por: İslamoğlu, Gamze, et al.
Publicado: (2025)
por: İslamoğlu, Gamze, et al.
Publicado: (2025)
HERO: Hardware-Efficient RL-based Optimization Framework for NeRF Quantization
por: Zhang, Yipu, et al.
Publicado: (2025)
por: Zhang, Yipu, et al.
Publicado: (2025)
Design of a 6-bit Threshold Inverter Quantization (TIQ) Flash Analog to Digital Converter (ADC)
por: Sarkar, Noyon Kumar, et al.
Publicado: (2025)
por: Sarkar, Noyon Kumar, et al.
Publicado: (2025)
Efficient yet Accurate End-to-End SC Accelerator Design
por: Li, Meng, et al.
Publicado: (2024)
por: Li, Meng, et al.
Publicado: (2024)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
por: Ramachandran, Akshat, et al.
Publicado: (2024)
por: Ramachandran, Akshat, et al.
Publicado: (2024)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
por: Zhang, Junming, et al.
Publicado: (2026)
por: Zhang, Junming, et al.
Publicado: (2026)
bitSMM: A bit-Serial Matrix Multiplication Accelerator
por: Antunes, Pedro, et al.
Publicado: (2026)
por: Antunes, Pedro, et al.
Publicado: (2026)
Splatonic: Architecture Support for 3D Gaussian Splatting SLAM via Sparse Processing
por: Huang, Xiaotong, et al.
Publicado: (2025)
por: Huang, Xiaotong, et al.
Publicado: (2025)
Potamoi: Accelerating Neural Rendering via a Unified Streaming Architecture
por: Feng, Yu, et al.
Publicado: (2024)
por: Feng, Yu, et al.
Publicado: (2024)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
por: Kida, Misaki, et al.
Publicado: (2025)
por: Kida, Misaki, et al.
Publicado: (2025)
AutoPDR: Circuit-Aware Solver Configuration Prediction for Hardware Model Checking
por: Hu, Guangyu, et al.
Publicado: (2026)
por: Hu, Guangyu, et al.
Publicado: (2026)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
por: Shan, Haoxuan, et al.
Publicado: (2025)
por: Shan, Haoxuan, et al.
Publicado: (2025)
QiMeng-CPU-v2: Automated Superscalar Processor Design by Learning Data Dependencies
por: Cheng, Shuyao, et al.
Publicado: (2025)
por: Cheng, Shuyao, et al.
Publicado: (2025)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
por: Wei, Jianyu, et al.
Publicado: (2025)
por: Wei, Jianyu, et al.
Publicado: (2025)
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
por: Zhang, Wenlun, et al.
Publicado: (2025)
por: Zhang, Wenlun, et al.
Publicado: (2025)
Nebula: Enable City-Scale 3D Gaussian Splatting in Virtual Reality via Collaborative Rendering and Accelerated Stereo Rasterization
por: Zhu, He, et al.
Publicado: (2025)
por: Zhu, He, et al.
Publicado: (2025)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
por: Duan, Bowen, et al.
Publicado: (2026)
por: Duan, Bowen, et al.
Publicado: (2026)
Is Finer Better? The Limits of Microscaling Formats in Large Language Models
por: Fasoli, Andrea, et al.
Publicado: (2026)
por: Fasoli, Andrea, et al.
Publicado: (2026)
AGON: Automated Design Framework for Customizing Processors from ISA Documents
por: Li, Chongxiao, et al.
Publicado: (2024)
por: Li, Chongxiao, et al.
Publicado: (2024)
Mixed Structural Choice Operator: Enhancing Technology Mapping with Heterogeneous Representations
por: Hu, Zhang, et al.
Publicado: (2025)
por: Hu, Zhang, et al.
Publicado: (2025)
Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency
por: Hübner, Paul, et al.
Publicado: (2025)
por: Hübner, Paul, et al.
Publicado: (2025)
Fletch: File-System Metadata Caching in Programmable Switches
por: Liu, Qingxiu, et al.
Publicado: (2025)
por: Liu, Qingxiu, et al.
Publicado: (2025)
Trimma: Trimming Metadata Storage and Latency for Hybrid Memory Systems
por: Li, Yiwei, et al.
Publicado: (2024)
por: Li, Yiwei, et al.
Publicado: (2024)
Ejemplares similares
-
M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type
por: Hu, Weiming, et al.
Publicado: (2025) -
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
por: Cheng, Jianyi, et al.
Publicado: (2023) -
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
por: Wipfli, Max, et al.
Publicado: (2026) -
Ascend HiFloat8 Format for Deep Learning
por: Luo, Yuanyong, et al.
Publicado: (2024) -
Characterization and Mitigation of Training Instabilities in Microscaling Formats
por: Su, Huangyuan, et al.
Publicado: (2025)