A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Jianyi, Zhang, Cheng, Yu, Zhewen, Bouganis, Christos-Savvas, Constantinides, George A., Zhao, Yiren |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
von: Zheng, Keran, et al.
Veröffentlicht: (2025)
von: Zheng, Keran, et al.
Veröffentlicht: (2025)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
von: Biggs, Benjamin, et al.
Veröffentlicht: (2023)
von: Biggs, Benjamin, et al.
Veröffentlicht: (2023)
Refining Datapath for Microscaling ViTs
von: Xiao, Can, et al.
Veröffentlicht: (2025)
von: Xiao, Can, et al.
Veröffentlicht: (2025)
SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip Eviction
von: Toupas, Petros, et al.
Veröffentlicht: (2024)
von: Toupas, Petros, et al.
Veröffentlicht: (2024)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
von: Hu, Weiming, et al.
Veröffentlicht: (2026)
von: Hu, Weiming, et al.
Veröffentlicht: (2026)
Revet: A Language and Compiler for Dataflow Threads
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
von: Wipfli, Max, et al.
Veröffentlicht: (2026)
von: Wipfli, Max, et al.
Veröffentlicht: (2026)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
von: Su, Huangyuan, et al.
Veröffentlicht: (2025)
von: Su, Huangyuan, et al.
Veröffentlicht: (2025)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
von: Andronic, Marta, et al.
Veröffentlicht: (2023)
von: Andronic, Marta, et al.
Veröffentlicht: (2023)
fpgaHART: A toolflow for throughput-oriented acceleration of 3D CNNs for HAR onto FPGAs
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
von: Lee, Jungi, et al.
Veröffentlicht: (2025)
von: Lee, Jungi, et al.
Veröffentlicht: (2025)
Soft GPGPU versus IP cores: Quantifying and Reducing the Performance Gap
von: Langhammer, Martin, et al.
Veröffentlicht: (2024)
von: Langhammer, Martin, et al.
Veröffentlicht: (2024)
Banked Memories for Soft SIMT Processors
von: Langhammer, Martin, et al.
Veröffentlicht: (2025)
von: Langhammer, Martin, et al.
Veröffentlicht: (2025)
A Statically and Dynamically Scalable Soft GPGPU
von: Langhammer, Martin, et al.
Veröffentlicht: (2024)
von: Langhammer, Martin, et al.
Veröffentlicht: (2024)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation
von: Zhang, Zixi, et al.
Veröffentlicht: (2023)
von: Zhang, Zixi, et al.
Veröffentlicht: (2023)
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
von: Park, Dahoon, et al.
Veröffentlicht: (2026)
von: Park, Dahoon, et al.
Veröffentlicht: (2026)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
von: Li, Fang
Veröffentlicht: (2025)
von: Li, Fang
Veröffentlicht: (2025)
HARFLOW3D: A Latency-Oriented 3D-CNN Accelerator Toolflow for HAR on FPGA Devices
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
ROVER: RTL Optimization via Verified E-Graph Rewriting
von: Coward, Samuel, et al.
Veröffentlicht: (2024)
von: Coward, Samuel, et al.
Veröffentlicht: (2024)
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
von: Cuyckens, Stef, et al.
Veröffentlicht: (2025)
von: Cuyckens, Stef, et al.
Veröffentlicht: (2025)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
von: Karfakis, George, et al.
Veröffentlicht: (2026)
von: Karfakis, George, et al.
Veröffentlicht: (2026)
Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
von: Wu, Haoran, et al.
Veröffentlicht: (2025)
von: Wu, Haoran, et al.
Veröffentlicht: (2025)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
von: Andronic, Marta, et al.
Veröffentlicht: (2024)
von: Andronic, Marta, et al.
Veröffentlicht: (2024)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
von: Gu, Hang, et al.
Veröffentlicht: (2026)
von: Gu, Hang, et al.
Veröffentlicht: (2026)
Combining Power and Arithmetic Optimization via Datapath Rewriting
von: Coward, Samuel, et al.
Veröffentlicht: (2024)
von: Coward, Samuel, et al.
Veröffentlicht: (2024)
Fast Cross-Operator Optimization of Attention Dataflow
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
von: Yu, Zhewen, et al.
Veröffentlicht: (2024) -
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
von: Zheng, Keran, et al.
Veröffentlicht: (2025) -
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
von: Biggs, Benjamin, et al.
Veröffentlicht: (2023) -
Refining Datapath for Microscaling ViTs
von: Xiao, Can, et al.
Veröffentlicht: (2025) -
SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip Eviction
von: Toupas, Petros, et al.
Veröffentlicht: (2024)