MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Zou, Jiaxiang, Chen, Yonghao, Wu, Ruilong, Chen, Xinyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4
di: Cim, Musa, et al.
Pubblicazione: (2026)
di: Cim, Musa, et al.
Pubblicazione: (2026)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
di: Mukunoki, Daichi
Pubblicazione: (2025)
di: Mukunoki, Daichi
Pubblicazione: (2025)
Faster Inference of LLMs using FP8 on the Intel Gaudi
di: Lee, Joonhyung, et al.
Pubblicazione: (2025)
di: Lee, Joonhyung, et al.
Pubblicazione: (2025)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
di: Liu, Shih-yang, et al.
Pubblicazione: (2023)
di: Liu, Shih-yang, et al.
Pubblicazione: (2023)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
di: Su, Zhongling, et al.
Pubblicazione: (2025)
di: Su, Zhongling, et al.
Pubblicazione: (2025)
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
di: Gowda, Bindu G, et al.
Pubblicazione: (2025)
di: Gowda, Bindu G, et al.
Pubblicazione: (2025)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
di: Xia, Haojun, et al.
Pubblicazione: (2024)
di: Xia, Haojun, et al.
Pubblicazione: (2024)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
di: Zhao, Liang, et al.
Pubblicazione: (2026)
di: Zhao, Liang, et al.
Pubblicazione: (2026)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
di: Nikolić, Miloš, et al.
Pubblicazione: (2022)
di: Nikolić, Miloš, et al.
Pubblicazione: (2022)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
Affordable HPC: Leveraging Small Clusters for Big Data and Graph Computing
di: Wu, Ruilong, et al.
Pubblicazione: (2024)
di: Wu, Ruilong, et al.
Pubblicazione: (2024)
Enhanced LPDDR4X PHY in 12 nm FinFET
di: Feldmann, Johannes, et al.
Pubblicazione: (2025)
di: Feldmann, Johannes, et al.
Pubblicazione: (2025)
Mixed Structural Choice Operator: Enhancing Technology Mapping with Heterogeneous Representations
di: Hu, Zhang, et al.
Pubblicazione: (2025)
di: Hu, Zhang, et al.
Pubblicazione: (2025)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
COMET: Towards Partical W4A4KV4 LLMs Serving
di: Liu, Lian, et al.
Pubblicazione: (2024)
di: Liu, Lian, et al.
Pubblicazione: (2024)
Corrigendum to: A Systematic Study of DDR4 DRAM Faults in the Field
di: Beigi, Majed Valad, et al.
Pubblicazione: (2024)
di: Beigi, Majed Valad, et al.
Pubblicazione: (2024)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
di: Dong, Pingcheng, et al.
Pubblicazione: (2026)
di: Dong, Pingcheng, et al.
Pubblicazione: (2026)
TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
di: Bertuletti, Marco, et al.
Pubblicazione: (2026)
di: Bertuletti, Marco, et al.
Pubblicazione: (2026)
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale
di: Zheng, Ziyang, et al.
Pubblicazione: (2025)
di: Zheng, Ziyang, et al.
Pubblicazione: (2025)
$R^4$: A Racetrack Register File with Runtime Software Reconfiguration
di: Hakert, Christian, et al.
Pubblicazione: (2025)
di: Hakert, Christian, et al.
Pubblicazione: (2025)
Iterative LLM-Based Assertion Generation Using Syntax-Semantic Representations for Functional Coverage-Guided Verification
di: Wang, Yonghao, et al.
Pubblicazione: (2026)
di: Wang, Yonghao, et al.
Pubblicazione: (2026)
Towards Reliable Systems: A Scalable Approach to AXI4 Transaction Monitoring
di: Liang, Chaoqun, et al.
Pubblicazione: (2025)
di: Liang, Chaoqun, et al.
Pubblicazione: (2025)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
di: Kida, Misaki, et al.
Pubblicazione: (2025)
di: Kida, Misaki, et al.
Pubblicazione: (2025)
Block-SSD: A New Block-Based Blocking SSD Architecture
di: Wong, Ryan, et al.
Pubblicazione: (2024)
di: Wong, Ryan, et al.
Pubblicazione: (2024)
A Benchmarking Platform for DDR4 Memory Performance in Data-Center-Class FPGAs
di: Galimberti, Andrea, et al.
Pubblicazione: (2025)
di: Galimberti, Andrea, et al.
Pubblicazione: (2025)
A High-Throughput Hardware Accelerator for Lempel-Ziv 4 Compression Algorithm
di: Chen, Tao, et al.
Pubblicazione: (2024)
di: Chen, Tao, et al.
Pubblicazione: (2024)
CoverAssert: Iterative LLM Assertion Generation Driven by Functional Coverage via Syntax-Semantic Representations
di: Wang, Yonghao, et al.
Pubblicazione: (2026)
di: Wang, Yonghao, et al.
Pubblicazione: (2026)
Evaluation of GPU Video Encoder for Low-Latency Real-Time 4K UHD Encoding
di: Arunruangsirilert, Kasidis, et al.
Pubblicazione: (2025)
di: Arunruangsirilert, Kasidis, et al.
Pubblicazione: (2025)
EA4RCA:Efficient AIE accelerator design framework for Regular Communication-Avoiding Algorithm
di: Zhang, W. B., et al.
Pubblicazione: (2024)
di: Zhang, W. B., et al.
Pubblicazione: (2024)
A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices
di: Park, Haneul, et al.
Pubblicazione: (2025)
di: Park, Haneul, et al.
Pubblicazione: (2025)
DAE4HLS: Exposing Memory-Level Parallelism for High-Level Synthesis using Explicit Decoupling
di: Metz, David, et al.
Pubblicazione: (2026)
di: Metz, David, et al.
Pubblicazione: (2026)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
di: Dai, Xilai, et al.
Pubblicazione: (2024)
di: Dai, Xilai, et al.
Pubblicazione: (2024)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
4T2R X-ReRAM CiM Array for Variation-tolerant, Low-power, Massively Parallel MAC Operation
di: Kihara, Fuyuki, et al.
Pubblicazione: (2025)
di: Kihara, Fuyuki, et al.
Pubblicazione: (2025)
Think with Self-Decoupling and Self-Verification: Automated RTL Design with Backtrack-ToT
di: Chao, Zhiteng, et al.
Pubblicazione: (2025)
di: Chao, Zhiteng, et al.
Pubblicazione: (2025)
AssertMiner: Module-Level Spec Generation and Assertion Mining using Static Analysis Guided LLMs
di: Lyu, Hongqin, et al.
Pubblicazione: (2025)
di: Lyu, Hongqin, et al.
Pubblicazione: (2025)
DeepAssert: An LLM-Aided Verification Framework with Fine-Grained Assertion Generation for Modules with Extracted Module Specifications
di: Wang, Yonghao, et al.
Pubblicazione: (2025)
di: Wang, Yonghao, et al.
Pubblicazione: (2025)
AssertFix: Empowering Automated Assertion Fix via Large Language Models
di: Lyu, Hongqin, et al.
Pubblicazione: (2025)
di: Lyu, Hongqin, et al.
Pubblicazione: (2025)
RidgeWalker: Perfectly Pipelined Graph Random Walks on FPGAs
di: Tan, Hongshi, et al.
Pubblicazione: (2026)
di: Tan, Hongshi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4
di: Cim, Musa, et al.
Pubblicazione: (2026) -
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
di: Park, Gunho, et al.
Pubblicazione: (2025) -
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
di: Mukunoki, Daichi
Pubblicazione: (2025) -
Faster Inference of LLMs using FP8 on the Intel Gaudi
di: Lee, Joonhyung, et al.
Pubblicazione: (2025) -
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
di: Liu, Shih-yang, et al.
Pubblicazione: (2023)