Refining Datapath for Microscaling ViTs
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Can, Cheng, Jianyi, Zhao, Aaron |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
by: Cheng, Jianyi, et al.
Published: (2023)
by: Cheng, Jianyi, et al.
Published: (2023)
Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
Combining Power and Arithmetic Optimization via Datapath Rewriting
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025)
by: Dumoulin, Joren, et al.
Published: (2025)
Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics
by: Morsali, Mehrdad, et al.
Published: (2025)
by: Morsali, Mehrdad, et al.
Published: (2025)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
by: Wipfli, Max, et al.
Published: (2026)
by: Wipfli, Max, et al.
Published: (2026)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
by: Karfakis, George, et al.
Published: (2026)
by: Karfakis, George, et al.
Published: (2026)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
by: Su, Huangyuan, et al.
Published: (2025)
by: Su, Huangyuan, et al.
Published: (2025)
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
by: İslamoğlu, Gamze, et al.
Published: (2025)
by: İslamoğlu, Gamze, et al.
Published: (2025)
M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
by: Hu, Weiming, et al.
Published: (2026)
by: Hu, Weiming, et al.
Published: (2026)
M$^2$-ViT: Accelerating Hybrid Vision Transformers with Two-Level Mixed Quantization
by: Liang, Yanbiao, et al.
Published: (2024)
by: Liang, Yanbiao, et al.
Published: (2024)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
by: Lee, Jungi, et al.
Published: (2025)
by: Lee, Jungi, et al.
Published: (2025)
Datapath Combinational Equivalence Checking With Hybrid Sweeping Engines and Parallelization
by: Chen, Zhihan, et al.
Published: (2024)
by: Chen, Zhihan, et al.
Published: (2024)
Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
by: Wu, Haoran, et al.
Published: (2025)
by: Wu, Haoran, et al.
Published: (2025)
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs
by: Wu, Haoran, et al.
Published: (2026)
by: Wu, Haoran, et al.
Published: (2026)
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
by: Park, Dahoon, et al.
Published: (2026)
by: Park, Dahoon, et al.
Published: (2026)
ViPSN 2.0: A Reconfigurable Battery-free IoT Platform for Vibration Energy Harvesting
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
by: Koo, Jahyun, et al.
Published: (2024)
by: Koo, Jahyun, et al.
Published: (2024)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
by: Parikh, Dhruv, et al.
Published: (2024)
by: Parikh, Dhruv, et al.
Published: (2024)
RayFlex: An Open-Source RTL Implementation of the Hardware Ray Tracer Datapath
by: Shen, Fangjia, et al.
Published: (2024)
by: Shen, Fangjia, et al.
Published: (2024)
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation
by: Zhang, Zixi, et al.
Published: (2023)
by: Zhang, Zixi, et al.
Published: (2023)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
by: Yu, Zhewen, et al.
Published: (2024)
by: Yu, Zhewen, et al.
Published: (2024)
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
by: Miao, Wenxuan, et al.
Published: (2025)
by: Miao, Wenxuan, et al.
Published: (2025)
NeuroVM: Dynamic Neuromorphic Hardware Virtualization
by: Isik, Murat, et al.
Published: (2024)
by: Isik, Murat, et al.
Published: (2024)
3D-Carbon: An Analytical Carbon Modeling Tool for 3D and 2.5D Integrated Circuits
by: Zhao, Yujie, et al.
Published: (2023)
by: Zhao, Yujie, et al.
Published: (2023)
Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
by: Jarmusch, Aaron, et al.
Published: (2025)
by: Jarmusch, Aaron, et al.
Published: (2025)
Is Finer Better? The Limits of Microscaling Formats in Large Language Models
by: Fasoli, Andrea, et al.
Published: (2026)
by: Fasoli, Andrea, et al.
Published: (2026)
Accelerating Sensor Fusion in Neuromorphic Computing: A Case Study on Loihi-2
by: Isik, Murat, et al.
Published: (2024)
by: Isik, Murat, et al.
Published: (2024)
Link Quality Aware Pathfinding for Chiplet Interconnects
by: Yen, Aaron, et al.
Published: (2026)
by: Yen, Aaron, et al.
Published: (2026)
A Review of Memory Wall for Neuromorphic Computing
by: Le, Dexter, et al.
Published: (2025)
by: Le, Dexter, et al.
Published: (2025)
A Reconfigurable Framework for AI-FPGA Agent Integration and Acceleration
by: Yunusoglu, Aybars, et al.
Published: (2026)
by: Yunusoglu, Aybars, et al.
Published: (2026)
Adaptive Hybrid FFT: A Novel Pipeline and Memory-Based Architecture for Radix-$2^k$ FFT in Large Size Processing
by: Zhao, Fangyu, et al.
Published: (2025)
by: Zhao, Fangyu, et al.
Published: (2025)
An Event-Driven Spiking Compute-In-Memory Macro based on SOT-MRAM
by: Yu, Deyang, et al.
Published: (2025)
by: Yu, Deyang, et al.
Published: (2025)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024)
by: Shao, Haikuo, et al.
Published: (2024)
Efficient Nonlinear Function Approximation in Analog Resistive Crossbars for Recurrent Neural Networks
by: Yang, Junyi, et al.
Published: (2024)
by: Yang, Junyi, et al.
Published: (2024)
NPU Design for Diffusion Language Model Inference
by: Lou, Binglei, et al.
Published: (2026)
by: Lou, Binglei, et al.
Published: (2026)
ViTAD: Timing Violation-Aware Debugging of RTL Code using Large Language Models
by: Lv, Wenhao, et al.
Published: (2025)
by: Lv, Wenhao, et al.
Published: (2025)
Similar Items
-
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
by: Cheng, Jianyi, et al.
Published: (2023) -
Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
by: Cuyckens, Stef, et al.
Published: (2025) -
Combining Power and Arithmetic Optimization via Datapath Rewriting
by: Coward, Samuel, et al.
Published: (2024) -
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
by: Moitra, Abhishek, et al.
Published: (2024) -
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025)