Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Jiayi, Lu, Jinming, Li, Hai, Wang, Xiangwei, Hao, Cong, Young, Ian, Zhang, Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
by: Zhang, Jinsong, et al.
Published: (2025)
by: Zhang, Jinsong, et al.
Published: (2025)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
by: Zhang, Weichuang, et al.
Published: (2024)
by: Zhang, Weichuang, et al.
Published: (2024)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
by: Shin, Duckgyu, et al.
Published: (2026)
by: Shin, Duckgyu, et al.
Published: (2026)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024)
by: Shao, Haikuo, et al.
Published: (2024)
Holistic Optimization Framework for FPGA Accelerators
by: Pouget, Stéphane, et al.
Published: (2025)
by: Pouget, Stéphane, et al.
Published: (2025)
A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODE
by: Okubo, Ikumi, et al.
Published: (2024)
by: Okubo, Ikumi, et al.
Published: (2024)
Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
LLM-Aided Compilation for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2024)
by: Hong, Charles, et al.
Published: (2024)
TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
by: Shen, Chaoyao, et al.
Published: (2026)
by: Shen, Chaoyao, et al.
Published: (2026)
RealProbe: An Automated and Lightweight Performance Profiler for In-FPGA Execution of High-Level Synthesis Designs
by: Kim, Jiho, et al.
Published: (2025)
by: Kim, Jiho, et al.
Published: (2025)
An Irredundant and Compressed Data Layout to Optimize Bandwidth Utilization of FPGA Accelerators
by: Ferry, Corentin, et al.
Published: (2024)
by: Ferry, Corentin, et al.
Published: (2024)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
by: Cheng, Feng, et al.
Published: (2025)
by: Cheng, Feng, et al.
Published: (2025)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
by: Gupta, Neelesh, et al.
Published: (2026)
by: Gupta, Neelesh, et al.
Published: (2026)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025)
by: Soi, Rupanshu, et al.
Published: (2025)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
by: Sohn, Gina, et al.
Published: (2025)
by: Sohn, Gina, et al.
Published: (2025)
LaZagna: An Open-Source Framework for Flexible 3D FPGA Architectural Exploration
by: Youssef, Ismael, et al.
Published: (2025)
by: Youssef, Ismael, et al.
Published: (2025)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
by: Huang, Sixiao, et al.
Published: (2025)
by: Huang, Sixiao, et al.
Published: (2025)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
by: Andronic, Marta, et al.
Published: (2023)
by: Andronic, Marta, et al.
Published: (2023)
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
FPGA Technology Mapping Using Sketch-Guided Program Synthesis
by: Smith, Gus Henry, et al.
Published: (2024)
by: Smith, Gus Henry, et al.
Published: (2024)
SSR: Spatial Sequential Hybrid Architecture for Latency Throughput Tradeoff in Transformer Acceleration
by: Zhuang, Jinming, et al.
Published: (2024)
by: Zhuang, Jinming, et al.
Published: (2024)
An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
by: Wang, Miaoxin, et al.
Published: (2024)
by: Wang, Miaoxin, et al.
Published: (2024)
Efficient FPGA Implementation of Time-Domain Popcount for Low-Complexity Machine Learning
by: Duan, Shengyu, et al.
Published: (2025)
by: Duan, Shengyu, et al.
Published: (2025)
FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference
by: Hsieh, Fen-Yu, et al.
Published: (2025)
by: Hsieh, Fen-Yu, et al.
Published: (2025)
FPGA-based Hyrbid Memory Emulation System
by: Wen, Fei, et al.
Published: (2020)
by: Wen, Fei, et al.
Published: (2020)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
by: Zhang, Zehuan, et al.
Published: (2024)
by: Zhang, Zehuan, et al.
Published: (2024)
A2H-MAS: An Algorithm-to-HLS Multi-Agent System for Automated and Reliable FPGA Implementation
by: Lei, Jie, et al.
Published: (2025)
by: Lei, Jie, et al.
Published: (2025)
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
by: Li, Jindong, et al.
Published: (2025)
by: Li, Jindong, et al.
Published: (2025)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
by: Duan, Bowen, et al.
Published: (2026)
by: Duan, Bowen, et al.
Published: (2026)
Demystifying FPGA Hard NoC Performance
by: Liu, Sihao, et al.
Published: (2025)
by: Liu, Sihao, et al.
Published: (2025)
Real Time FPGA Based Transformers & VLMs for Vision Tasks: SOTA Designs and Optimizations
by: Sali, Safa Mohammed, et al.
Published: (2025)
by: Sali, Safa Mohammed, et al.
Published: (2025)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
by: Ding, Hong, et al.
Published: (2025)
by: Ding, Hong, et al.
Published: (2025)
An Efficient Data Reuse with Tile-Based Adaptive Stationary for Transformer Accelerators
by: Li, Tseng-Jen, et al.
Published: (2025)
by: Li, Tseng-Jen, et al.
Published: (2025)
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
by: Dong, Jiale, et al.
Published: (2025)
by: Dong, Jiale, et al.
Published: (2025)
Exploring FPGA designs for MX and beyond
by: Samson, Ebby, et al.
Published: (2024)
by: Samson, Ebby, et al.
Published: (2024)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
by: Chen, Hongzheng, et al.
Published: (2023)
by: Chen, Hongzheng, et al.
Published: (2023)
Scaling Laws for Floating Point Quantization Training
by: Sun, Xingwu, et al.
Published: (2025)
by: Sun, Xingwu, et al.
Published: (2025)
Similar Items
-
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
by: Lu, Jinming, et al.
Published: (2025) -
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
by: Lu, Jinming, et al.
Published: (2025) -
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
by: Zhang, Jinsong, et al.
Published: (2025) -
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
by: Zhang, Weichuang, et al.
Published: (2024) -
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
by: Shin, Duckgyu, et al.
Published: (2026)