Saved in:
| Main Authors: | Guan, Yue, Yu, Hongtao, Chen, Peng, Shi, Daohang, Manivannan, Karthik, Riasanovsky, Nicholas J, Ren, Manman, Wang, Lei, Nay, Shane, Kanuparthy, Partha, Pan, Zaifeng, Hu, Zhengding, Ding, Yufei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.10905 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Real Time Evolvable Hardware for Optimal Reconfiguration of Cusp-Like Pulse Shapers
by: Lanchares, Juan, et al.
Published: (2024)
by: Lanchares, Juan, et al.
Published: (2024)
Bombyx: OpenCilk Compilation for FPGA Hardware Acceleration
by: Shahawy, Mohamed, et al.
Published: (2025)
by: Shahawy, Mohamed, et al.
Published: (2025)
Generation of Compiler Backends from Formal Models of Hardware
by: Smith, Gus Henry
Published: (2024)
by: Smith, Gus Henry
Published: (2024)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
by: Zhou, Zhe, et al.
Published: (2024)
by: Zhou, Zhe, et al.
Published: (2024)
Look-Up Table based Neural Network Hardware
by: Sen, Ovishake, et al.
Published: (2024)
by: Sen, Ovishake, et al.
Published: (2024)
HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation
by: Zhou, Yaoyun, et al.
Published: (2025)
by: Zhou, Yaoyun, et al.
Published: (2025)
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
by: Pu, Huanzhi, et al.
Published: (2025)
by: Pu, Huanzhi, et al.
Published: (2025)
Hardware-Aware Neural Network Compilation with Learned Optimization: A RISC-V Accelerator Approach
by: Ganti, Ravindra, et al.
Published: (2025)
by: Ganti, Ravindra, et al.
Published: (2025)
SIP: Autotuning GPU Native Schedules via Stochastic Instruction Perturbation
by: He, Guoliang, et al.
Published: (2024)
by: He, Guoliang, et al.
Published: (2024)
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving
by: Yu, Zhongkai, et al.
Published: (2026)
by: Yu, Zhongkai, et al.
Published: (2026)
ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design
by: Yu, Zhongkai, et al.
Published: (2026)
by: Yu, Zhongkai, et al.
Published: (2026)
ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation
by: Yu, Zhongkai, et al.
Published: (2026)
by: Yu, Zhongkai, et al.
Published: (2026)
Vorion: A RISC-V GPU with Hardware-Accelerated 3D Gaussian Rendering and Training
by: Wang, Yipeng, et al.
Published: (2025)
by: Wang, Yipeng, et al.
Published: (2025)
Design Environment of Quantization-Aware Edge AI Hardware for Few-Shot Learning
by: Kanda, R., et al.
Published: (2026)
by: Kanda, R., et al.
Published: (2026)
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
by: Yu, Zhongkai, et al.
Published: (2025)
by: Yu, Zhongkai, et al.
Published: (2025)
OneAdapt: Adaptive Compilation for Resource-Constrained Photonic One-Way Quantum Computing
by: Zhang, Hezi, et al.
Published: (2025)
by: Zhang, Hezi, et al.
Published: (2025)
Bit-Width-Aware Design Environment for Few-Shot Learning on Edge AI Hardware
by: Kanda, R., et al.
Published: (2026)
by: Kanda, R., et al.
Published: (2026)
Parallax: A Compiler for Neutral Atom Quantum Computers under Hardware Constraints
by: Ludmir, Jason, et al.
Published: (2024)
by: Ludmir, Jason, et al.
Published: (2024)
CHOSEN: Compilation to Hardware Optimization Stack for Efficient Vision Transformer Inference
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
by: Qiu, Tong Dong, et al.
Published: (2023)
by: Qiu, Tong Dong, et al.
Published: (2023)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
by: Gouk, Donghyun, et al.
Published: (2025)
by: Gouk, Donghyun, et al.
Published: (2025)
RoboGPU: Accelerating GPU Collision Detection for Robotics
by: Liu, Lufei, et al.
Published: (2026)
by: Liu, Lufei, et al.
Published: (2026)
Sustainable Hardware Specialization
by: Dangi, Pranav, et al.
Published: (2024)
by: Dangi, Pranav, et al.
Published: (2024)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
by: You, Dean, et al.
Published: (2025)
by: You, Dean, et al.
Published: (2025)
DRCY: Agentic Hardware Design Reviews
by: Dumont, Kyle, et al.
Published: (2026)
by: Dumont, Kyle, et al.
Published: (2026)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
by: Zhang, Weichuang, et al.
Published: (2026)
by: Zhang, Weichuang, et al.
Published: (2026)
Compilation and Execution of an Embeddable YOLO-NAS on the VTA
by: Faure-Gignoux, Anthony, et al.
Published: (2026)
by: Faure-Gignoux, Anthony, et al.
Published: (2026)
Revet: A Language and Compiler for Dataflow Threads
by: Rucker, Alexander, et al.
Published: (2023)
by: Rucker, Alexander, et al.
Published: (2023)
Analyzing Modern NVIDIA GPU cores
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
by: Shao, Kunming, et al.
Published: (2024)
by: Shao, Kunming, et al.
Published: (2024)
An FPGA Compiler for On-the-Fly Adaptive CNN Deployment and Reconfiguration
by: Mazouz, Alaa, et al.
Published: (2025)
by: Mazouz, Alaa, et al.
Published: (2025)
Design of a GPU with Heterogeneous Cores for Graphics
by: Tomás, Aurora, et al.
Published: (2026)
by: Tomás, Aurora, et al.
Published: (2026)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
by: Luo, Weile, et al.
Published: (2024)
by: Luo, Weile, et al.
Published: (2024)
COOK Access Control on an embedded Volta GPU
by: Lesage, Benjamin, et al.
Published: (2024)
by: Lesage, Benjamin, et al.
Published: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
by: Sun, Xiaotian, et al.
Published: (2024)
by: Sun, Xiaotian, et al.
Published: (2024)
Analyzing and Improving Hardware Modeling of Accel-Sim
by: Huerta, Rodrigo, et al.
Published: (2024)
by: Huerta, Rodrigo, et al.
Published: (2024)
Hardware and software build flow with SoCMake
by: Pejašinović, Risto, et al.
Published: (2025)
by: Pejašinović, Risto, et al.
Published: (2025)
QED: Scalable Verification of Hardware Memory Consistency
by: Ravi, Gokulan, et al.
Published: (2024)
by: Ravi, Gokulan, et al.
Published: (2024)
NeuroVM: Dynamic Neuromorphic Hardware Virtualization
by: Isik, Murat, et al.
Published: (2024)
by: Isik, Murat, et al.
Published: (2024)
Similar Items
-
Real Time Evolvable Hardware for Optimal Reconfiguration of Cusp-Like Pulse Shapers
by: Lanchares, Juan, et al.
Published: (2024) -
Bombyx: OpenCilk Compilation for FPGA Hardware Acceleration
by: Shahawy, Mohamed, et al.
Published: (2025) -
Generation of Compiler Backends from Formal Models of Hardware
by: Smith, Gus Henry
Published: (2024) -
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
by: Zhou, Zhe, et al.
Published: (2024) -
Look-Up Table based Neural Network Hardware
by: Sen, Ovishake, et al.
Published: (2024)