Implementation and Evaluation of Stable Diffusion on a General-Purpose CGLA Accelerator
Fuente:
arXiv
Salvato in:
| Autori principali: | Ando, Takuto, Eto, Yu, Nakashima, Yasuhiko |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Kernel Mapping and Comprehensive System Evaluation of LLM Acceleration on a CGLA
di: Ando, Takuto, et al.
Pubblicazione: (2025)
di: Ando, Takuto, et al.
Pubblicazione: (2025)
Energy-Efficient Hardware Acceleration of Whisper ASR on a CGLA
di: Ando, Takuto, et al.
Pubblicazione: (2025)
di: Ando, Takuto, et al.
Pubblicazione: (2025)
Facial Expression Recognition System Using DNN Accelerator with Multi-threading on FPGA
di: Ando, Takuto, et al.
Pubblicazione: (2025)
di: Ando, Takuto, et al.
Pubblicazione: (2025)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
di: Le, Tran Xuan Hieu, et al.
Pubblicazione: (2024)
di: Le, Tran Xuan Hieu, et al.
Pubblicazione: (2024)
Trinity: A General Purpose FHE Accelerator
di: Deng, Xianglong, et al.
Pubblicazione: (2024)
di: Deng, Xianglong, et al.
Pubblicazione: (2024)
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
di: Langarita, Rubén, et al.
Pubblicazione: (2025)
di: Langarita, Rubén, et al.
Pubblicazione: (2025)
SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations
di: Wang, Zhican, et al.
Pubblicazione: (2025)
di: Wang, Zhican, et al.
Pubblicazione: (2025)
Implementation and Analysis of Thermometer Encoding in DWN FPGA Accelerators
di: Mecik, Michael, et al.
Pubblicazione: (2025)
di: Mecik, Michael, et al.
Pubblicazione: (2025)
Exploring the Limitations of Kolmogorov-Arnold Networks in Classification: Insights to Software Training and Hardware Implementation
di: Tran, Van Duy, et al.
Pubblicazione: (2024)
di: Tran, Van Duy, et al.
Pubblicazione: (2024)
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
di: Ghosh, Arkapravo, et al.
Pubblicazione: (2025)
di: Ghosh, Arkapravo, et al.
Pubblicazione: (2025)
Efficient Implementation of an Adaptive Transformer Accelerator for Massive MIMO Outdoor Localization
di: Yaman, Ilayda, et al.
Pubblicazione: (2026)
di: Yaman, Ilayda, et al.
Pubblicazione: (2026)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
di: Wang, Jiayi, et al.
Pubblicazione: (2026)
di: Wang, Jiayi, et al.
Pubblicazione: (2026)
Design, Implementation and Evaluation of the SVNAPOT Extension on a RISC-V Processor
di: Papadopoulos, Nikolaos-Charalampos, et al.
Pubblicazione: (2024)
di: Papadopoulos, Nikolaos-Charalampos, et al.
Pubblicazione: (2024)
FQsun: A Configurable Wave Function-Based Quantum Emulator for Power-Efficient Quantum Simulations
di: Vu, Tuan Hai, et al.
Pubblicazione: (2024)
di: Vu, Tuan Hai, et al.
Pubblicazione: (2024)
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
di: Kim, Seeyeon, et al.
Pubblicazione: (2026)
di: Kim, Seeyeon, et al.
Pubblicazione: (2026)
Voyager: An End-to-End Framework for Design-Space Exploration and Generation of DNN Accelerators
di: Prabhu, Kartik, et al.
Pubblicazione: (2025)
di: Prabhu, Kartik, et al.
Pubblicazione: (2025)
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
di: Li, Cong, et al.
Pubblicazione: (2026)
di: Li, Cong, et al.
Pubblicazione: (2026)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
di: Li, Meng, et al.
Pubblicazione: (2026)
di: Li, Meng, et al.
Pubblicazione: (2026)
Accelerating Diffusion Models for Generative AI Applications with Silicon Photonics
di: Suresh, Tharini, et al.
Pubblicazione: (2026)
di: Suresh, Tharini, et al.
Pubblicazione: (2026)
GEM3D CIM General Purpose Matrix Computation Using 3D Integrated SRAM eDRAM Hybrid Compute In Memory on Memory Architecture
di: Chakraborty, Subhradip, et al.
Pubblicazione: (2026)
di: Chakraborty, Subhradip, et al.
Pubblicazione: (2026)
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
di: Zhang, Wenlun, et al.
Pubblicazione: (2025)
di: Zhang, Wenlun, et al.
Pubblicazione: (2025)
Towards Generalized On-Chip Communication for Programmable Accelerators in Heterogeneous Architectures
di: Zuckerman, Joseph, et al.
Pubblicazione: (2024)
di: Zuckerman, Joseph, et al.
Pubblicazione: (2024)
GTA: a new General Tensor Accelerator with Better Area Efficiency and Data Reuse
di: Ai, Chenyang, et al.
Pubblicazione: (2024)
di: Ai, Chenyang, et al.
Pubblicazione: (2024)
DiffuSE: Cross-Layer Design Space Exploration of DNN Accelerator via Diffusion-Driven Optimization
di: Ren, Yi, et al.
Pubblicazione: (2025)
di: Ren, Yi, et al.
Pubblicazione: (2025)
ASiM: Modeling and Analyzing Inference Accuracy of SRAM-Based Analog CiM Circuits
di: Zhang, Wenlun, et al.
Pubblicazione: (2024)
di: Zhang, Wenlun, et al.
Pubblicazione: (2024)
A 28.6 mJ/iter Stable Diffusion Processor for Text-to-Image Generation with Patch Similarity-based Sparsity Augmentation and Text-based Mixed-Precision
di: Choi, Jiwon, et al.
Pubblicazione: (2024)
di: Choi, Jiwon, et al.
Pubblicazione: (2024)
Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
di: Yoon, Jieon, et al.
Pubblicazione: (2026)
di: Yoon, Jieon, et al.
Pubblicazione: (2026)
GSIM: Accelerating RTL Simulation for Large-Scale Designs
di: Chen, Lu, et al.
Pubblicazione: (2025)
di: Chen, Lu, et al.
Pubblicazione: (2025)
Hardware Efficient Accelerator for Spiking Transformer With Reconfigurable Parallel Time Step Computing
di: Chen, Bo-Yu, et al.
Pubblicazione: (2025)
di: Chen, Bo-Yu, et al.
Pubblicazione: (2025)
Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Capacity
di: Xue, Zi Yu, et al.
Pubblicazione: (2023)
di: Xue, Zi Yu, et al.
Pubblicazione: (2023)
Graphitron: A Domain Specific Language for FPGA-based Graph Processing Accelerator Generation
di: Zhang, Xinmiao, et al.
Pubblicazione: (2024)
di: Zhang, Xinmiao, et al.
Pubblicazione: (2024)
A Flexible Template for Edge Generative AI with High-Accuracy Accelerated Softmax & GELU
di: Belano, Andrea, et al.
Pubblicazione: (2024)
di: Belano, Andrea, et al.
Pubblicazione: (2024)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
di: Latif, Imran, et al.
Pubblicazione: (2024)
di: Latif, Imran, et al.
Pubblicazione: (2024)
Crypto-RV: High-Efficiency FPGA-Based RISC-V Cryptographic Co-Processor for IoT Security
di: Pham, Anh Kiet, et al.
Pubblicazione: (2026)
di: Pham, Anh Kiet, et al.
Pubblicazione: (2026)
A Review of SRAM-based Compute-in-Memory Circuits
di: Yoshioka, Kentaro, et al.
Pubblicazione: (2024)
di: Yoshioka, Kentaro, et al.
Pubblicazione: (2024)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
di: Geens, Robin, et al.
Pubblicazione: (2026)
di: Geens, Robin, et al.
Pubblicazione: (2026)
SA-DS: A Dataset for Large Language Model-Driven AI Accelerator Design Generation
di: Vungarala, Deepak, et al.
Pubblicazione: (2024)
di: Vungarala, Deepak, et al.
Pubblicazione: (2024)
SimulatorCoder: DNN Accelerator Simulator Code Generation and Optimization via Large Language Models
di: Xia, Yuhuan, et al.
Pubblicazione: (2026)
di: Xia, Yuhuan, et al.
Pubblicazione: (2026)
General-Purpose Multicore Architectures
di: Ghose, Saugata
Pubblicazione: (2024)
di: Ghose, Saugata
Pubblicazione: (2024)
Documenti analoghi
-
Efficient Kernel Mapping and Comprehensive System Evaluation of LLM Acceleration on a CGLA
di: Ando, Takuto, et al.
Pubblicazione: (2025) -
Energy-Efficient Hardware Acceleration of Whisper ASR on a CGLA
di: Ando, Takuto, et al.
Pubblicazione: (2025) -
Facial Expression Recognition System Using DNN Accelerator with Multi-threading on FPGA
di: Ando, Takuto, et al.
Pubblicazione: (2025) -
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
di: Le, Tran Xuan Hieu, et al.
Pubblicazione: (2024) -
Trinity: A General Purpose FHE Accelerator
di: Deng, Xianglong, et al.
Pubblicazione: (2024)