A 28.6 mJ/iter Stable Diffusion Processor for Text-to-Image Generation with Patch Similarity-based Sparsity Augmentation and Text-based Mixed-Precision
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Jiwon, Jo, Wooyoung, Hong, Seongyon, Kwon, Beomseok, Park, Wonhoon, Yoo, Hoi-Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
by: Dai, Xilai, et al.
Published: (2024)
by: Dai, Xilai, et al.
Published: (2024)
A 0.5V, 6.2$μ$W, 0.059mm$^{2}$ Sinusoidal Current Generator IC with 0.088% THD for Bio-Impedance Sensing
by: Kim, Kwantae, et al.
Published: (2024)
by: Kim, Kwantae, et al.
Published: (2024)
A Bit Level Weight Reordering Strategy Based on Column Similarity to Explore Weight Sparsity in RRAM-based NN Accelerator
by: Yang, Weiping, et al.
Published: (2025)
by: Yang, Weiping, et al.
Published: (2025)
HARP: Hadamard-Domain Write-and-Verify for Noise-Robust RRAM Programming
by: Choi, Ilhuan, et al.
Published: (2026)
by: Choi, Ilhuan, et al.
Published: (2026)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
by: Wang, Chuanning, et al.
Published: (2024)
by: Wang, Chuanning, et al.
Published: (2024)
SPEED: A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
by: Wang, Chuanning, et al.
Published: (2024)
by: Wang, Chuanning, et al.
Published: (2024)
SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
by: He, Zicheng, et al.
Published: (2026)
by: He, Zicheng, et al.
Published: (2026)
SimFuzz: Similarity-guided Block-level Mutation for RISC-V Processor Fuzzing
by: Lyu, Hao, et al.
Published: (2026)
by: Lyu, Hao, et al.
Published: (2026)
Sensitivity-Aware Mixed-Precision Quantization for ReRAM-based Computing-in-Memory
by: Chen, Guan-Cheng, et al.
Published: (2025)
by: Chen, Guan-Cheng, et al.
Published: (2025)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
by: Han, Wontak, et al.
Published: (2024)
by: Han, Wontak, et al.
Published: (2024)
CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
by: Min, Kyeongpil, et al.
Published: (2026)
by: Min, Kyeongpil, et al.
Published: (2026)
Large Processor Chip Model
by: Chang, Kaiyan, et al.
Published: (2025)
by: Chang, Kaiyan, et al.
Published: (2025)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
by: Li, Bohan, et al.
Published: (2026)
by: Li, Bohan, et al.
Published: (2026)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
by: Moon, Seungjae, et al.
Published: (2024)
by: Moon, Seungjae, et al.
Published: (2024)
SAMIPS: A Synthesised Asynchronous Processor
by: Zhang, Qianyi, et al.
Published: (2024)
by: Zhang, Qianyi, et al.
Published: (2024)
Banked Memories for Soft SIMT Processors
by: Langhammer, Martin, et al.
Published: (2025)
by: Langhammer, Martin, et al.
Published: (2025)
A Precision-Scalable RISC-V DNN Processor with On-Device Learning Capability at the Extreme Edge
by: Huang, Longwei, et al.
Published: (2023)
by: Huang, Longwei, et al.
Published: (2023)
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
by: Kim, Seeyeon, et al.
Published: (2026)
by: Kim, Seeyeon, et al.
Published: (2026)
Ellora: Exploring Low-Power OFDM-based Radar Processors using Approximate Computing
by: Bhattacharjya, Rajat, et al.
Published: (2023)
by: Bhattacharjya, Rajat, et al.
Published: (2023)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
by: Yoon, Jieon, et al.
Published: (2026)
by: Yoon, Jieon, et al.
Published: (2026)
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
by: Kim, Jae-Young, et al.
Published: (2025)
by: Kim, Jae-Young, et al.
Published: (2025)
A 950 MHz SIMT Soft Processor
by: Langhammer, Martin, et al.
Published: (2025)
by: Langhammer, Martin, et al.
Published: (2025)
Hypervisor Extension for a RISC-V Processor
by: Gauchola, Jaume, et al.
Published: (2024)
by: Gauchola, Jaume, et al.
Published: (2024)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
by: Wang, Chuanzhen, et al.
Published: (2026)
by: Wang, Chuanzhen, et al.
Published: (2026)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
by: Seo, Minseok, et al.
Published: (2024)
by: Seo, Minseok, et al.
Published: (2024)
FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF Rendering
by: Noh, Seock-Hwan, et al.
Published: (2025)
by: Noh, Seock-Hwan, et al.
Published: (2025)
Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks
by: Wei, Chiyue, et al.
Published: (2025)
by: Wei, Chiyue, et al.
Published: (2025)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
by: Garg, Raveesh, et al.
Published: (2025)
by: Garg, Raveesh, et al.
Published: (2025)
Implementation of Compute Intensive Algorithms on Software Configurable Processor
by: Ganesha, et al.
Published: (2025)
by: Ganesha, et al.
Published: (2025)
Duet: Creating Harmony between Processors and Embedded FPGAs
by: Li, Ang, et al.
Published: (2023)
by: Li, Ang, et al.
Published: (2023)
Web-Based Simulator of Superscalar RISC-V Processors
by: Jaros, Jiri, et al.
Published: (2024)
by: Jaros, Jiri, et al.
Published: (2024)
Neuromorphic Processor Employing FPGA Technology with Universal Interconnections
by: Harlikar, Pracheta, et al.
Published: (2025)
by: Harlikar, Pracheta, et al.
Published: (2025)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
by: Yu, Feng, et al.
Published: (2026)
by: Yu, Feng, et al.
Published: (2026)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
by: Wu, Jiajun, et al.
Published: (2024)
by: Wu, Jiajun, et al.
Published: (2024)
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
by: Wu, Michael, et al.
Published: (2025)
by: Wu, Michael, et al.
Published: (2025)
Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice Sparsity
by: Kam, Dongyun, et al.
Published: (2024)
by: Kam, Dongyun, et al.
Published: (2024)
Image processing Application Development on Software Configurable Processor Array
by: Prabhu, Ganesh, et al.
Published: (2025)
by: Prabhu, Ganesh, et al.
Published: (2025)
Functional ISS-Driven Verification of Superscalar RISC-V Processors
by: Galimberti, Andrea, et al.
Published: (2024)
by: Galimberti, Andrea, et al.
Published: (2024)
Similar Items
-
SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
by: Choi, Yuseon, et al.
Published: (2025) -
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
by: Dai, Xilai, et al.
Published: (2024) -
A 0.5V, 6.2$μ$W, 0.059mm$^{2}$ Sinusoidal Current Generator IC with 0.088% THD for Bio-Impedance Sensing
by: Kim, Kwantae, et al.
Published: (2024) -
A Bit Level Weight Reordering Strategy Based on Column Similarity to Explore Weight Sparsity in RRAM-based NN Accelerator
by: Yang, Weiping, et al.
Published: (2025) -
HARP: Hadamard-Domain Write-and-Verify for Noise-Robust RRAM Programming
by: Choi, Ilhuan, et al.
Published: (2026)