HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Yaoyun, Wang, Qian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HERO: Hardware-Efficient RL-based Optimization Framework for NeRF Quantization
por: Zhang, Yipu, et al.
Publicado: (2025)
por: Zhang, Yipu, et al.
Publicado: (2025)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
por: Zhang, Kaixuan, et al.
Publicado: (2026)
por: Zhang, Kaixuan, et al.
Publicado: (2026)
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
por: Guan, Yue, et al.
Publicado: (2026)
por: Guan, Yue, et al.
Publicado: (2026)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
por: Zhang, Weichuang, et al.
Publicado: (2026)
por: Zhang, Weichuang, et al.
Publicado: (2026)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
por: Cheng, Jianyi, et al.
Publicado: (2023)
por: Cheng, Jianyi, et al.
Publicado: (2023)
Hardware-Aware Neural Network Compilation with Learned Optimization: A RISC-V Accelerator Approach
por: Ganti, Ravindra, et al.
Publicado: (2025)
por: Ganti, Ravindra, et al.
Publicado: (2025)
GPU-Accelerated Simulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
por: Gonul, Yilmaz Ege, et al.
Publicado: (2025)
por: Gonul, Yilmaz Ege, et al.
Publicado: (2025)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
por: Gouk, Donghyun, et al.
Publicado: (2025)
por: Gouk, Donghyun, et al.
Publicado: (2025)
OpenACM: An Open-Source SRAM-Based Approximate CiM Compiler
por: Zhou, Yiqi, et al.
Publicado: (2026)
por: Zhou, Yiqi, et al.
Publicado: (2026)
Evaluation of GPU Video Encoder for Low-Latency Real-Time 4K UHD Encoding
por: Arunruangsirilert, Kasidis, et al.
Publicado: (2025)
por: Arunruangsirilert, Kasidis, et al.
Publicado: (2025)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
por: Luo, Weile, et al.
Publicado: (2024)
por: Luo, Weile, et al.
Publicado: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
por: Sun, Xiaotian, et al.
Publicado: (2024)
por: Sun, Xiaotian, et al.
Publicado: (2024)
A Time- and Energy-Efficient CNN with Dense Connections on Memristor-Based Chips
por: Zhou, Wenyong, et al.
Publicado: (2025)
por: Zhou, Wenyong, et al.
Publicado: (2025)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
por: Zhao, Shixin, et al.
Publicado: (2025)
por: Zhao, Shixin, et al.
Publicado: (2025)
Compilation and Execution of an Embeddable YOLO-NAS on the VTA
por: Faure-Gignoux, Anthony, et al.
Publicado: (2026)
por: Faure-Gignoux, Anthony, et al.
Publicado: (2026)
Revet: A Language and Compiler for Dataflow Threads
por: Rucker, Alexander, et al.
Publicado: (2023)
por: Rucker, Alexander, et al.
Publicado: (2023)
Analyzing Modern NVIDIA GPU cores
por: Huerta, Rodrigo, et al.
Publicado: (2025)
por: Huerta, Rodrigo, et al.
Publicado: (2025)
Bi-SamplerZ: A Hardware-Efficient Gaussian Sampler Architecture for Quantum-Resistant Falcon Signatures
por: Zhao, Binke, et al.
Publicado: (2025)
por: Zhao, Binke, et al.
Publicado: (2025)
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
por: Fu, Zizhuo, et al.
Publicado: (2026)
por: Fu, Zizhuo, et al.
Publicado: (2026)
Bombyx: OpenCilk Compilation for FPGA Hardware Acceleration
por: Shahawy, Mohamed, et al.
Publicado: (2025)
por: Shahawy, Mohamed, et al.
Publicado: (2025)
An FPGA Compiler for On-the-Fly Adaptive CNN Deployment and Reconfiguration
por: Mazouz, Alaa, et al.
Publicado: (2025)
por: Mazouz, Alaa, et al.
Publicado: (2025)
PipeRTL: Timing-Aware Pipeline Optimization at IR-Level for RTL Generation
por: Yin, Shuo, et al.
Publicado: (2026)
por: Yin, Shuo, et al.
Publicado: (2026)
SPPAM: Signature Pattern Prediction and Access-Map Prefetcher
por: Merrell, Maccoy, et al.
Publicado: (2026)
por: Merrell, Maccoy, et al.
Publicado: (2026)
RoboGPU: Accelerating GPU Collision Detection for Robotics
por: Liu, Lufei, et al.
Publicado: (2026)
por: Liu, Lufei, et al.
Publicado: (2026)
Design of a GPU with Heterogeneous Cores for Graphics
por: Tomás, Aurora, et al.
Publicado: (2026)
por: Tomás, Aurora, et al.
Publicado: (2026)
COOK Access Control on an embedded Volta GPU
por: Lesage, Benjamin, et al.
Publicado: (2024)
por: Lesage, Benjamin, et al.
Publicado: (2024)
SEGA-DCIM: Design Space Exploration-Guided Automatic Digital CIM Compiler with Multiple Precision Support
por: Diao, Haikang, et al.
Publicado: (2025)
por: Diao, Haikang, et al.
Publicado: (2025)
Efficient LLM inference solution on Intel GPU
por: Wu, Hui, et al.
Publicado: (2023)
por: Wu, Hui, et al.
Publicado: (2023)
Leveraging Application-Specific Knowledge for Energy-Efficient Deep Learning Accelerators on Resource-Constrained FPGAs
por: Qian, Chao
Publicado: (2025)
por: Qian, Chao
Publicado: (2025)
Multiport Support for Vortex OpenGPU Memory Hierarchy
por: Shin, Injae, et al.
Publicado: (2025)
por: Shin, Injae, et al.
Publicado: (2025)
CuLifter: Lifting GPU Binaries to Typed IR
por: Zhao, Jisheng, et al.
Publicado: (2026)
por: Zhao, Jisheng, et al.
Publicado: (2026)
Capstone: Power-Capped Pipelining for Coarse-Grained Reconfigurable Array Compilers
por: Yarzada, Sabrina, et al.
Publicado: (2026)
por: Yarzada, Sabrina, et al.
Publicado: (2026)
Building a Reusable and Extensible Automatic Compiler Infrastructure for Reconfigurable Devices
por: Zang, Zhenya, et al.
Publicado: (2023)
por: Zang, Zhenya, et al.
Publicado: (2023)
SimulatorCoder: DNN Accelerator Simulator Code Generation and Optimization via Large Language Models
por: Xia, Yuhuan, et al.
Publicado: (2026)
por: Xia, Yuhuan, et al.
Publicado: (2026)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
por: Zhang, Kaixuan, et al.
Publicado: (2026)
por: Zhang, Kaixuan, et al.
Publicado: (2026)
GAP-LA: GPU-Accelerated Performance-Driven Layer Assignment
por: Zhao, Chunyuan, et al.
Publicado: (2025)
por: Zhao, Chunyuan, et al.
Publicado: (2025)
Thermal Analysis for NVIDIA GTX480 Fermi GPU Architecture
por: Nagendra, Savinay
Publicado: (2024)
por: Nagendra, Savinay
Publicado: (2024)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
por: Zhao, Wei, et al.
Publicado: (2024)
por: Zhao, Wei, et al.
Publicado: (2024)
From PyTorch to Calyx: An Open-Source Compiler Toolchain for ML Accelerators
por: Xie, Jiahan, et al.
Publicado: (2025)
por: Xie, Jiahan, et al.
Publicado: (2025)
Adapting Atmospheric Chemistry Components for Efficient GPU Accelerators
por: Ruiz, Christian Guzman, et al.
Publicado: (2024)
por: Ruiz, Christian Guzman, et al.
Publicado: (2024)
Ejemplares similares
-
HERO: Hardware-Efficient RL-based Optimization Framework for NeRF Quantization
por: Zhang, Yipu, et al.
Publicado: (2025) -
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
por: Zhang, Kaixuan, et al.
Publicado: (2026) -
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
por: Guan, Yue, et al.
Publicado: (2026) -
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
por: Zhang, Weichuang, et al.
Publicado: (2026) -
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
por: Cheng, Jianyi, et al.
Publicado: (2023)