Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4
Fuente:
arXiv
Salvato in:
| Autori principali: | Cim, Musa, Topcu, Burak, Kandemir, Mahmut Taylan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Pretraining large language models with MXFP4 on Native FP4 Hardware
di: Cim, Musa, et al.
Pubblicazione: (2026)
di: Cim, Musa, et al.
Pubblicazione: (2026)
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
di: Zou, Jiaxiang, et al.
Pubblicazione: (2026)
di: Zou, Jiaxiang, et al.
Pubblicazione: (2026)
Parallel Context Compaction for Long-Horizon LLM Agent Serving
di: Cim, Musa, et al.
Pubblicazione: (2026)
di: Cim, Musa, et al.
Pubblicazione: (2026)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
di: Chhugani, Jatin, et al.
Pubblicazione: (2026)
di: Chhugani, Jatin, et al.
Pubblicazione: (2026)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
di: Chen, Yao, et al.
Pubblicazione: (2022)
di: Chen, Yao, et al.
Pubblicazione: (2022)
Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA
di: Zhu, Xuqi, et al.
Pubblicazione: (2024)
di: Zhu, Xuqi, et al.
Pubblicazione: (2024)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
di: Kim, Jiyoon, et al.
Pubblicazione: (2025)
di: Kim, Jiyoon, et al.
Pubblicazione: (2025)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
di: Liu, Shih-yang, et al.
Pubblicazione: (2023)
di: Liu, Shih-yang, et al.
Pubblicazione: (2023)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
di: Xia, Haojun, et al.
Pubblicazione: (2024)
di: Xia, Haojun, et al.
Pubblicazione: (2024)
Efficient LLM inference solution on Intel GPU
di: Wu, Hui, et al.
Pubblicazione: (2023)
di: Wu, Hui, et al.
Pubblicazione: (2023)
A 28 nm AI microcontroller with tightly coupled zero-standby power weight memory featuring standard logic compatible 4 Mb 4-bits/cell embedded flash technology
di: Kim, Daewung, et al.
Pubblicazione: (2025)
di: Kim, Daewung, et al.
Pubblicazione: (2025)
DPUV4E: High-Throughput DPU Architecture Design for CNN on Versal ACAP
di: Li, Guoyu, et al.
Pubblicazione: (2025)
di: Li, Guoyu, et al.
Pubblicazione: (2025)
LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation
di: Zhong, Ruizhe, et al.
Pubblicazione: (2023)
di: Zhong, Ruizhe, et al.
Pubblicazione: (2023)
LLM4SecHW: Leveraging Domain Specific Large Language Model for Hardware Debugging
di: Fu, Weimin, et al.
Pubblicazione: (2024)
di: Fu, Weimin, et al.
Pubblicazione: (2024)
Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours
di: The Verkor Team, et al.
Pubblicazione: (2026)
di: The Verkor Team, et al.
Pubblicazione: (2026)
Full-stack evaluation of Machine Learning inference workloads for RISC-V systems
di: Bhattacharjee, Debjyoti, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Debjyoti, et al.
Pubblicazione: (2024)
HiFloat4 Format for Language Model Inference
di: Luo, Yuanyong, et al.
Pubblicazione: (2026)
di: Luo, Yuanyong, et al.
Pubblicazione: (2026)
Salient Store: Enabling Smart Storage for Continuous Learning Edge Servers
di: Mishra, Cyan Subhra, et al.
Pubblicazione: (2024)
di: Mishra, Cyan Subhra, et al.
Pubblicazione: (2024)
AutoGNN: End-to-End Hardware-Driven Graph Preprocessing for Enhanced GNN Performance
di: Kang, Seungkwan, et al.
Pubblicazione: (2026)
di: Kang, Seungkwan, et al.
Pubblicazione: (2026)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
Multibit neural inference in a N-ary crossbar architecture
di: Moureaux, Anatole, et al.
Pubblicazione: (2026)
di: Moureaux, Anatole, et al.
Pubblicazione: (2026)
rule4ml: An Open-Source Tool for Resource Utilization and Latency Estimation for ML Models on FPGA
di: Rahimifar, Mohammad Mehdi, et al.
Pubblicazione: (2024)
di: Rahimifar, Mohammad Mehdi, et al.
Pubblicazione: (2024)
Hardware-software co-exploration with racetrack memory based in-memory computing for CNN inference in embedded systems
di: Choong, Benjamin Chen Ming, et al.
Pubblicazione: (2025)
di: Choong, Benjamin Chen Ming, et al.
Pubblicazione: (2025)
Adaptive Robotic Arm Control with a Spiking Recurrent Neural Network on a Digital Accelerator
di: Linares-Barranco, Alejandro, et al.
Pubblicazione: (2024)
di: Linares-Barranco, Alejandro, et al.
Pubblicazione: (2024)
Towards Cognitive AI Systems: a Survey and Prospective on Neuro-Symbolic AI
di: Wan, Zishen, et al.
Pubblicazione: (2024)
di: Wan, Zishen, et al.
Pubblicazione: (2024)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
di: Lin, Jiawei, et al.
Pubblicazione: (2025)
di: Lin, Jiawei, et al.
Pubblicazione: (2025)
Estimating Voltage Drop: Models, Features and Data Representation Towards a Neural Surrogate
di: Jin, Yifei, et al.
Pubblicazione: (2025)
di: Jin, Yifei, et al.
Pubblicazione: (2025)
Rome was Not Built in a Single Step: Hierarchical Prompting for LLM-based Chip Design
di: Nakkab, Andre, et al.
Pubblicazione: (2024)
di: Nakkab, Andre, et al.
Pubblicazione: (2024)
Insights from Verification: Training a Verilog Generation LLM with Reinforcement Learning with Testbench Feedback
di: Wang, Ning, et al.
Pubblicazione: (2025)
di: Wang, Ning, et al.
Pubblicazione: (2025)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
di: Zeng, Shulin, et al.
Pubblicazione: (2024)
di: Zeng, Shulin, et al.
Pubblicazione: (2024)
ROMA: a Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
di: Wang, Wenqiang, et al.
Pubblicazione: (2025)
di: Wang, Wenqiang, et al.
Pubblicazione: (2025)
Hey AI, Generate Me a Hardware Code! Agentic AI-based Hardware Design & Verification
di: Gadde, Deepak Narayan, et al.
Pubblicazione: (2025)
di: Gadde, Deepak Narayan, et al.
Pubblicazione: (2025)
Design Conductor: An agent autonomously builds a 1.5 GHz Linux-capable RISC-V CPU
di: The Verkor Team, et al.
Pubblicazione: (2026)
di: The Verkor Team, et al.
Pubblicazione: (2026)
Natural Language to Verilog: Design of a Recurrent Spiking Neural Network using Large Language Models and ChatGPT
di: Vitolo, Paola, et al.
Pubblicazione: (2024)
di: Vitolo, Paola, et al.
Pubblicazione: (2024)
DeepV: A Model-Agnostic Retrieval-Augmented Framework for Verilog Code Generation with a High-Quality Knowledge Base
di: Ibnat, Zahin, et al.
Pubblicazione: (2025)
di: Ibnat, Zahin, et al.
Pubblicazione: (2025)
wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation
di: Hawks, Benjamin, et al.
Pubblicazione: (2025)
di: Hawks, Benjamin, et al.
Pubblicazione: (2025)
Spiker+: a framework for the generation of efficient Spiking Neural Networks FPGA accelerators for inference at the edge
di: Carpegna, Alessio, et al.
Pubblicazione: (2024)
di: Carpegna, Alessio, et al.
Pubblicazione: (2024)
POET: Power-Oriented Evolutionary Tuning for LLM-Based RTL PPA Optimization
di: Ping, Heng, et al.
Pubblicazione: (2026)
di: Ping, Heng, et al.
Pubblicazione: (2026)
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
di: Lou, Binglei, et al.
Pubblicazione: (2026)
di: Lou, Binglei, et al.
Pubblicazione: (2026)
HYPERHEURIST: A Simulated Annealing-Based Control Framework for LLM-Driven Code Generation in Optimized Hardware Design
di: Ahir, Shiva, et al.
Pubblicazione: (2026)
di: Ahir, Shiva, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Pretraining large language models with MXFP4 on Native FP4 Hardware
di: Cim, Musa, et al.
Pubblicazione: (2026) -
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
di: Zou, Jiaxiang, et al.
Pubblicazione: (2026) -
Parallel Context Compaction for Long-Horizon LLM Agent Serving
di: Cim, Musa, et al.
Pubblicazione: (2026) -
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
di: Chhugani, Jatin, et al.
Pubblicazione: (2026) -
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
di: Chen, Yao, et al.
Pubblicazione: (2022)