A 28nm 0.22μJ/token memory-compute-intensity-aware CNN-Transformer accelerator with hybrid-attention-based layer-fusion and cascaded pruning for semantic-segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Pingcheng, Tan, Yonghao, Liu, Xuejiao, Luo, Peng, Liu, Yu, Liang, Luhong, Zhou, Yitong, Pang, Di, Yung, Man-To, Zhang, Dong, Huang, Xijie, Liu, Shih-Yang, Wu, Yongkun, Tian, Fengshi, Tsui, Chi-Ying, Tu, Fengbin, Cheng, Kwang-Ting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
by: Tan, Yonghao, et al.
Published: (2025)
by: Tan, Yonghao, et al.
Published: (2025)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
by: Dong, Pingcheng, et al.
Published: (2026)
by: Dong, Pingcheng, et al.
Published: (2026)
Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
by: Dong, Pingcheng, et al.
Published: (2024)
by: Dong, Pingcheng, et al.
Published: (2024)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
by: Ma, Songchen, et al.
Published: (2026)
by: Ma, Songchen, et al.
Published: (2026)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
by: Liu, Shih-yang, et al.
Published: (2023)
by: Liu, Shih-yang, et al.
Published: (2023)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
A 10.60 $μ$W 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia Detection
by: Qin, Yifan, et al.
Published: (2024)
by: Qin, Yifan, et al.
Published: (2024)
T-REX: A 68-567 μs/token, 0.41-3.95 μJ/token Transformer Accelerator with Reduced External Memory Access and Enhanced Hardware Utilization in 16nm FinFET
by: Moon, Seunghyun, et al.
Published: (2025)
by: Moon, Seunghyun, et al.
Published: (2025)
Memory Efficient Transformer Adapter for Dense Predictions
by: Zhang, Dong, et al.
Published: (2025)
by: Zhang, Dong, et al.
Published: (2025)
Towards Customized Knowledge Distillation for Chip-Level Dense Image Predictions
by: Zhang, Dong, et al.
Published: (2024)
by: Zhang, Dong, et al.
Published: (2024)
Compressing CNN models for resource-constrained systems by channel and layer pruning
by: Sadaqa, Ahmed, et al.
Published: (2025)
by: Sadaqa, Ahmed, et al.
Published: (2025)
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
by: Jiang, Xixi, et al.
Published: (2025)
by: Jiang, Xixi, et al.
Published: (2025)
SR-LIO++: Efficient LiDAR-Inertial Odometry and Quantized Mapping with Sweep Reconstruction
by: Yuan, Zikang, et al.
Published: (2025)
by: Yuan, Zikang, et al.
Published: (2025)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
by: Huang, Xijie, et al.
Published: (2024)
by: Huang, Xijie, et al.
Published: (2024)
Efficient and Robust Quantization-aware Training via Adaptive Coreset Selection
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight Combination
by: Zhao, Liang, et al.
Published: (2025)
by: Zhao, Liang, et al.
Published: (2025)
Multimodal semantic retrieval for product search
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
Joint stereo 3D object detection and implicit surface reconstruction
by: Li, Shichao, et al.
Published: (2021)
by: Li, Shichao, et al.
Published: (2021)
Practical token pruning for foundation models in few-shot conversational virtual assistant systems
by: Qi, Haode, et al.
Published: (2024)
by: Qi, Haode, et al.
Published: (2024)
DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
by: Shao, Kunming, et al.
Published: (2025)
by: Shao, Kunming, et al.
Published: (2025)
Chapter 6 Environmental movements in Asia
by: Wu, Fengshi
Published: (2024)
by: Wu, Fengshi
Published: (2024)
Game semantics for the constructive $μ$-calculus
by: Pacheco, Leonardo
Published: (2023)
by: Pacheco, Leonardo
Published: (2023)
Sub-token ViT Embedding via Stochastic Resonance Transformers
by: Lao, Dong, et al.
Published: (2023)
by: Lao, Dong, et al.
Published: (2023)
Cyclic Contrastive Knowledge Transfer for Open-Vocabulary Object Detection
by: Zhang, Chuhan, et al.
Published: (2025)
by: Zhang, Chuhan, et al.
Published: (2025)
Learning effective pruning at initialization from iterative pruning
by: Liu, Shengkai, et al.
Published: (2024)
by: Liu, Shengkai, et al.
Published: (2024)
A laser plasma soliton fusion scheme
by: Chen, Pisin, et al.
Published: (2026)
by: Chen, Pisin, et al.
Published: (2026)
PRS2Net: an efficient intelligent carrot detection model via filter pruning and attention mechanisms
by: Huayu Fu, et al.
Published: (2025)
by: Huayu Fu, et al.
Published: (2025)
Datasets and codes
by: Liu, Yonghao
Published: (2025)
by: Liu, Yonghao
Published: (2025)
Zero-shot sketch-based remote sensing image retrieval based on multi-level and attention-guided tokenization
by: Yang, Bo, et al.
Published: (2024)
by: Yang, Bo, et al.
Published: (2024)
Multi-Modal Brain Tumor Segmentation via 3D Multi-Scale Self-attention and Cross-attention
by: Huang, Yonghao, et al.
Published: (2025)
by: Huang, Yonghao, et al.
Published: (2025)
Research on feature fusion and multimodal patent text based on graph attention network
by: Song, Zhenzhen, et al.
Published: (2025)
by: Song, Zhenzhen, et al.
Published: (2025)
Lightweight crop disease identification network based on frequency domain and channel mixing attention and cross‐scale semantic fusion
by: Tiancan Jian, et al.
Published: (2025)
by: Tiancan Jian, et al.
Published: (2025)
On the token distance modeling ability of higher RoPE attention dimension
by: Hong, Xiangyu, et al.
Published: (2024)
by: Hong, Xiangyu, et al.
Published: (2024)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
by: Zhao, Liang, et al.
Published: (2026)
by: Zhao, Liang, et al.
Published: (2026)
An adversarial feature learning based semantic communication method for Human 3D Reconstruction
by: Liu, Shaojiang, et al.
Published: (2024)
by: Liu, Shaojiang, et al.
Published: (2024)
GQA-μP: The maximal parameterization update for grouped query attention
by: Chickering, Kyle R., et al.
Published: (2026)
by: Chickering, Kyle R., et al.
Published: (2026)
Multi-modal and Multi-view Fundus Image Fusion for Retinopathy Diagnosis via Multi-scale Cross-attention and Shifted Window Self-attention
by: Huang, Yonghao, et al.
Published: (2025)
by: Huang, Yonghao, et al.
Published: (2025)
Game semantics for lattice-based modal μ-calculus
by: Ding, Yiwen, et al.
Published: (2023)
by: Ding, Yiwen, et al.
Published: (2023)
Road Traffic Sign Recognition method using Siamese network Combining Efficient-CNN based Encoder
by: Xi, Zhenghao, et al.
Published: (2025)
by: Xi, Zhenghao, et al.
Published: (2025)
SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
by: Shao, Kunming, et al.
Published: (2024)
by: Shao, Kunming, et al.
Published: (2024)
Similar Items
-
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
by: Tan, Yonghao, et al.
Published: (2025) -
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
by: Dong, Pingcheng, et al.
Published: (2026) -
Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
by: Dong, Pingcheng, et al.
Published: (2024) -
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
by: Ma, Songchen, et al.
Published: (2026) -
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
by: Liu, Shih-yang, et al.
Published: (2023)