Saved in:
| Main Authors: | Chen, Jiajie, Qu, Peng, Zhang, Youhui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.13900 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Branch Predictor for Graph Applications
by: Upasna, et al.
Published: (2026)
by: Upasna, et al.
Published: (2026)
Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency
by: Hübner, Paul, et al.
Published: (2025)
by: Hübner, Paul, et al.
Published: (2025)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
by: Luo, Weile, et al.
Published: (2024)
by: Luo, Weile, et al.
Published: (2024)
Evaluation of Posits for Spectral Analysis Using a Software-Defined Dataflow Architecture
by: Deshmukh, Sameer, et al.
Published: (2024)
by: Deshmukh, Sameer, et al.
Published: (2024)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
by: Zhou, Zhe, et al.
Published: (2024)
by: Zhou, Zhe, et al.
Published: (2024)
Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
Architecture, Simulation and Software Stack to Support Post-CMOS Accelerators: The ARCHYTAS Project
by: Agosta, Giovanni, et al.
Published: (2025)
by: Agosta, Giovanni, et al.
Published: (2025)
Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
by: Jarmusch, Aaron, et al.
Published: (2025)
by: Jarmusch, Aaron, et al.
Published: (2025)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
by: Alsop, Johnathan, et al.
Published: (2023)
by: Alsop, Johnathan, et al.
Published: (2023)
APACHE: A Processing-Near-Memory Architecture for Multi-Scheme Fully Homomorphic Encryption
by: Ding, Lin, et al.
Published: (2024)
by: Ding, Lin, et al.
Published: (2024)
Exposing Shadow Branches
by: Pepi, Chrysanthos, et al.
Published: (2024)
by: Pepi, Chrysanthos, et al.
Published: (2024)
Study on the Particle Sorting Performance for Reactor Monte Carlo Neutron Transport on Apple Unified Memory GPUs
by: Liu, Changyuan
Published: (2024)
by: Liu, Changyuan
Published: (2024)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
by: Luo, Weile, et al.
Published: (2025)
by: Luo, Weile, et al.
Published: (2025)
Optimizing Scalable Multi-Cluster Architectures for Next-Generation Wireless Sensing and Communication
by: Riedel, Samuel, et al.
Published: (2025)
by: Riedel, Samuel, et al.
Published: (2025)
A Fully-Configurable Open-Source Software-Defined Digital Quantized Spiking Neural Core Architecture
by: Matinizadeh, Shadi, et al.
Published: (2024)
by: Matinizadeh, Shadi, et al.
Published: (2024)
Workload Characterization for Branch Predictability
by: Vikas, FNU, et al.
Published: (2025)
by: Vikas, FNU, et al.
Published: (2025)
Aquas: Enhancing Domain Specialization through Holistic Hardware-Software Co-Optimization based on MLIR
by: Zou, Yuyang, et al.
Published: (2025)
by: Zou, Yuyang, et al.
Published: (2025)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
by: Xu, Haocheng, et al.
Published: (2024)
by: Xu, Haocheng, et al.
Published: (2024)
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025)
by: Papalamprou, Ilias, et al.
Published: (2025)
GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing
by: Pei, Minnan, et al.
Published: (2025)
by: Pei, Minnan, et al.
Published: (2025)
Evolution, Challenges, and Optimization in Computer Architecture: The Role of Reconfigurable Systems
by: Ederhion, Jefferson, et al.
Published: (2024)
by: Ederhion, Jefferson, et al.
Published: (2024)
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
by: Zhang, Tao, et al.
Published: (2026)
by: Zhang, Tao, et al.
Published: (2026)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
by: Raj, Ritik, et al.
Published: (2025)
by: Raj, Ritik, et al.
Published: (2025)
Branch Target Buffer Reverse Engineering on Arm
by: Wan, Junpeng
Published: (2024)
by: Wan, Junpeng
Published: (2024)
Automated HEMT Model Construction from Datasheets via Multi-Modal Intelligence and Prior-Knowledge-Free Optimization
by: Peng, Yuang, et al.
Published: (2025)
by: Peng, Yuang, et al.
Published: (2025)
The Survey of Chiplet-based Integrated Architecture: An EDA perspective
by: Chen, Shixin, et al.
Published: (2024)
by: Chen, Shixin, et al.
Published: (2024)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Thermal Analysis for NVIDIA GTX480 Fermi GPU Architecture
by: Nagendra, Savinay
Published: (2024)
by: Nagendra, Savinay
Published: (2024)
The Non-Predictability of Mispredicted Branches using Timing Information
by: Constantinou, Ioannis, et al.
Published: (2026)
by: Constantinou, Ioannis, et al.
Published: (2026)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
by: Sarda, Giuseppe M., et al.
Published: (2024)
by: Sarda, Giuseppe M., et al.
Published: (2024)
Optimized Memory System Architecture for VESA VDC-M Decoder with Multi-Slice Support
by: Yang, Hannah, et al.
Published: (2025)
by: Yang, Hannah, et al.
Published: (2025)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
by: Luo, Shuqing, et al.
Published: (2026)
by: Luo, Shuqing, et al.
Published: (2026)
Reconfigurable Stream Network Architecture
by: Wang, Chengyue, et al.
Published: (2024)
by: Wang, Chengyue, et al.
Published: (2024)
Workload-Aware Early-Stage Power Delivery Network Optimization via Architectural Power Traces
by: Hayes, Oran, et al.
Published: (2026)
by: Hayes, Oran, et al.
Published: (2026)
Branch Prediction in Hardcaml for a RISC-V 32im CPU
by: Saveau, Alex
Published: (2023)
by: Saveau, Alex
Published: (2023)
Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
by: Jang, Yongjoo, et al.
Published: (2025)
by: Jang, Yongjoo, et al.
Published: (2025)
A Protocol-Independent Transport Architecture
by: Mohammadtaheri, Kimiya, et al.
Published: (2026)
by: Mohammadtaheri, Kimiya, et al.
Published: (2026)
System-Technology Co-Optimization of Bitline Routing and Bonding Pathways in Monolithic 3D DRAM Architectures
by: Lee, Kiseok, et al.
Published: (2026)
by: Lee, Kiseok, et al.
Published: (2026)
HOPE: Holistic STT-RAM Architecture Exploration Framework for Future Cross-Platform Analysis
by: SeyedFaraji, Saeed, et al.
Published: (2024)
by: SeyedFaraji, Saeed, et al.
Published: (2024)
Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
by: Wei, Chiyue, et al.
Published: (2025)
by: Wei, Chiyue, et al.
Published: (2025)
Similar Items
-
Optimizing Branch Predictor for Graph Applications
by: Upasna, et al.
Published: (2026) -
Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency
by: Hübner, Paul, et al.
Published: (2025) -
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
by: Luo, Weile, et al.
Published: (2024) -
Evaluation of Posits for Spectral Analysis Using a Software-Defined Dataflow Architecture
by: Deshmukh, Sameer, et al.
Published: (2024) -
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
by: Zhou, Zhe, et al.
Published: (2024)