Saved in:
| Main Authors: | Wu, Junchi, Wan, Xinfei, Li, Zhuoran, Jin, Yuyang, Sun, Guangyu, Liang, Yun, Zhou, Diyu, Zhuo, Youwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.24112 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
by: Li, Zhuoran, et al.
Published: (2026)
by: Li, Zhuoran, et al.
Published: (2026)
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
by: Zhu, Jingchen, et al.
Published: (2024)
by: Zhu, Jingchen, et al.
Published: (2024)
Aquas: Enhancing Domain Specialization through Holistic Hardware-Software Co-Optimization based on MLIR
by: Zou, Yuyang, et al.
Published: (2025)
by: Zou, Yuyang, et al.
Published: (2025)
Towards Generalized On-Chip Communication for Programmable Accelerators in Heterogeneous Architectures
by: Zuckerman, Joseph, et al.
Published: (2024)
by: Zuckerman, Joseph, et al.
Published: (2024)
AGS: Accelerating 3D Gaussian Splatting SLAM via CODEC-Assisted Frame Covisibility Detection
by: He, Houshu, et al.
Published: (2025)
by: He, Houshu, et al.
Published: (2025)
Search-in-Memory (SiM): Reliable, Versatile, and Efficient Data Matching in SSD's NAND Flash Memory Chip for Data Indexing Acceleration
by: Chen, Yun-Chih, et al.
Published: (2024)
by: Chen, Yun-Chih, et al.
Published: (2024)
Energy-Efficient QoS-Aware Scheduling for S-NUCA Many-Cores
by: Wasala, Sudam M., et al.
Published: (2025)
by: Wasala, Sudam M., et al.
Published: (2025)
TAMI-MPC:Trusted Acceleration of Minimal-Interaction MPC for Efficient Nonlinear Inference
by: Li, Zhuoran, et al.
Published: (2026)
by: Li, Zhuoran, et al.
Published: (2026)
WISP: Image Segmentation-Based Whitespace Diagnosis for Optimal Rectilinear Floorplanning
by: Zhao, Xiaotian, et al.
Published: (2025)
by: Zhao, Xiaotian, et al.
Published: (2025)
Efficient yet Accurate End-to-End SC Accelerator Design
by: Li, Meng, et al.
Published: (2024)
by: Li, Meng, et al.
Published: (2024)
Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
by: He, Siyuan, et al.
Published: (2025)
by: He, Siyuan, et al.
Published: (2025)
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
by: Abdelmaksoud, Ahmed J., et al.
Published: (2026)
by: Abdelmaksoud, Ahmed J., et al.
Published: (2026)
Monad: Towards Cost-effective Specialization for Chiplet-based Spatial Accelerators
by: Hao, Xiaochen, et al.
Published: (2023)
by: Hao, Xiaochen, et al.
Published: (2023)
Hot-LEGO: Architect Microfluidic Cooling Equipped 3DICs with Pre-RTL Thermal Simulation
by: Wang, Runxi, et al.
Published: (2024)
by: Wang, Runxi, et al.
Published: (2024)
Triangel: A High-Performance, Accurate, Timely On-Chip Temporal Prefetcher
by: Ainsworth, Sam, et al.
Published: (2024)
by: Ainsworth, Sam, et al.
Published: (2024)
Resilient and Secure Programmable System-on-Chip Accelerator Offload
by: Gouveia, Inês Pinto, et al.
Published: (2024)
by: Gouveia, Inês Pinto, et al.
Published: (2024)
ControlPULP: A RISC-V On-Chip Parallel Power Controller for Many-Core HPC Processors with FPGA-Based Hardware-In-The-Loop Power and Thermal Emulation
by: Ottaviano, Alessandro, et al.
Published: (2023)
by: Ottaviano, Alessandro, et al.
Published: (2023)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
by: Wang, Zhao, et al.
Published: (2021)
by: Wang, Zhao, et al.
Published: (2021)
DAS-MP: Enabling High-Quality Macro Placement with Enhanced Dataflow Awareness
by: Zhao, Xiaotian, et al.
Published: (2025)
by: Zhao, Xiaotian, et al.
Published: (2025)
HFRWKV: A High-Performance Fully On-Chip Hardware Accelerator for RWKV
by: Shijie, Liu, et al.
Published: (2026)
by: Shijie, Liu, et al.
Published: (2026)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
FedChip: Federated LLM for Artificial Intelligence Accelerator Chip Design
by: Nazzal, Mahmoud, et al.
Published: (2025)
by: Nazzal, Mahmoud, et al.
Published: (2025)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
Cool-3D: An End-to-End Thermal-Aware Framework for Early-Phase Design Space Exploration of Microfluidic-Cooled 3DICs
by: Wang, Runxi, et al.
Published: (2025)
by: Wang, Runxi, et al.
Published: (2025)
Shift-Left Techniques in Electronic Design Automation: A Survey
by: Wu, Xinyue, et al.
Published: (2025)
by: Wu, Xinyue, et al.
Published: (2025)
Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators
by: Huang, Zongle, et al.
Published: (2026)
by: Huang, Zongle, et al.
Published: (2026)
Work-In-Progress: Accelerating Numpy With OpenBLAS For Open-Source RISC-V Chips
by: Koenig, Cyril, et al.
Published: (2025)
by: Koenig, Cyril, et al.
Published: (2025)
Large Processor Chip Model
by: Chang, Kaiyan, et al.
Published: (2025)
by: Chang, Kaiyan, et al.
Published: (2025)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
by: Tong, Jianming, et al.
Published: (2024)
by: Tong, Jianming, et al.
Published: (2024)
Efficient Implementation of an Adaptive Transformer Accelerator for Massive MIMO Outdoor Localization
by: Yaman, Ilayda, et al.
Published: (2026)
by: Yaman, Ilayda, et al.
Published: (2026)
FireBridge: Cycle-Accurate Hardware + Firmware Co-Verification for Modern Accelerators
by: Abarajithan, G, et al.
Published: (2026)
by: Abarajithan, G, et al.
Published: (2026)
Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators
by: Zhang, Chenguang, et al.
Published: (2024)
by: Zhang, Chenguang, et al.
Published: (2024)
Generalized Ping-Pong: Off-Chip Memory Bandwidth Centric Pipelining Strategy for Processing-In-Memory Accelerators
by: Wang, Ruibao, et al.
Published: (2024)
by: Wang, Ruibao, et al.
Published: (2024)
An Energy-Efficient Artefact Detection Accelerator on FPGAs for Hyper-Spectral Satellite Imagery
by: Castelino, Cornell, et al.
Published: (2024)
by: Castelino, Cornell, et al.
Published: (2024)
NOVA: Coordinated Test Selection and Bayes-Optimized Constrained Randomization for Accelerated Coverage Closure
by: Peng, Weijie, et al.
Published: (2025)
by: Peng, Weijie, et al.
Published: (2025)
Efficient and Accurate Graph Classification with Hyperdimensional Computing on FPGA
by: Arockiaraj, Jebacyril, et al.
Published: (2025)
by: Arockiaraj, Jebacyril, et al.
Published: (2025)
A Time- and Energy-Efficient CNN with Dense Connections on Memristor-Based Chips
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
by: Wang, Aotao, et al.
Published: (2025)
by: Wang, Aotao, et al.
Published: (2025)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
by: Liu, Xingyu, et al.
Published: (2025)
by: Liu, Xingyu, et al.
Published: (2025)
ZynqParrot: A Scale-Down Approach to Cycle-Accurate, FPGA-Accelerated Co-Emulation
by: Ruelas-Petrisko, Daniel, et al.
Published: (2025)
by: Ruelas-Petrisko, Daniel, et al.
Published: (2025)
Similar Items
-
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
by: Li, Zhuoran, et al.
Published: (2026) -
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
by: Zhu, Jingchen, et al.
Published: (2024) -
Aquas: Enhancing Domain Specialization through Holistic Hardware-Software Co-Optimization based on MLIR
by: Zou, Yuyang, et al.
Published: (2025) -
Towards Generalized On-Chip Communication for Programmable Accelerators in Heterogeneous Architectures
by: Zuckerman, Joseph, et al.
Published: (2024) -
AGS: Accelerating 3D Gaussian Splatting SLAM via CODEC-Assisted Frame Covisibility Detection
by: He, Houshu, et al.
Published: (2025)