Learning Generalizable Program and Architecture Representations for Performance Modeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Lingda, Flynn, Thomas, Hoisie, Adolfy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Less is More: Hop-Wise Graph Attention for Scalable and Generalizable Learning on Circuits
por: Deng, Chenhui, et al.
Publicado: (2024)
por: Deng, Chenhui, et al.
Publicado: (2024)
Federated Knowledge Distillation for Multi-Model Architectures Lithography Hotspot Detection
por: Li, Yuqi, et al.
Publicado: (2025)
por: Li, Yuqi, et al.
Publicado: (2025)
DeepSeq2: Enhanced Sequential Circuit Learning with Disentangled Representations
por: Khan, Sadaf, et al.
Publicado: (2024)
por: Khan, Sadaf, et al.
Publicado: (2024)
Accelerating Computer Architecture Simulation through Machine Learning
por: Ali, Wajid, et al.
Publicado: (2024)
por: Ali, Wajid, et al.
Publicado: (2024)
Learning Library Cell Representations in Vector Space
por: Liang, Rongjian, et al.
Publicado: (2025)
por: Liang, Rongjian, et al.
Publicado: (2025)
Explainable Fuzzy Neural Network with Multi-Fidelity Reinforcement Learning for Micro-Architecture Design Space Exploration
por: Fan, Hanwei, et al.
Publicado: (2024)
por: Fan, Hanwei, et al.
Publicado: (2024)
NAS-Cap: Deep-Learning Driven 3-D Capacitance Extraction with Neural Architecture Search and Data Augmentation
por: Li, Haoyuan, et al.
Publicado: (2024)
por: Li, Haoyuan, et al.
Publicado: (2024)
Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
por: Xu, Bin, et al.
Publicado: (2025)
por: Xu, Bin, et al.
Publicado: (2025)
TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
por: Shen, Chaoyao, et al.
Publicado: (2026)
por: Shen, Chaoyao, et al.
Publicado: (2026)
AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
por: Sharma, Amit
Publicado: (2025)
por: Sharma, Amit
Publicado: (2025)
Hierarchical Mixture of Experts: Generalizable Learning for High-Level Synthesis
por: Li, Weikai, et al.
Publicado: (2024)
por: Li, Weikai, et al.
Publicado: (2024)
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale
por: Zheng, Ziyang, et al.
Publicado: (2025)
por: Zheng, Ziyang, et al.
Publicado: (2025)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
por: Wolters, Christopher, et al.
Publicado: (2024)
por: Wolters, Christopher, et al.
Publicado: (2024)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
por: Duan, Bowen, et al.
Publicado: (2026)
por: Duan, Bowen, et al.
Publicado: (2026)
Cross-Modality Program Representation Learning for Electronic Design Automation with High-Level Synthesis
por: Qin, Zongyue, et al.
Publicado: (2024)
por: Qin, Zongyue, et al.
Publicado: (2024)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
por: Ibrahim, Muhammad Sohail, et al.
Publicado: (2024)
por: Ibrahim, Muhammad Sohail, et al.
Publicado: (2024)
M100: An Orchestrated Dataflow Architecture Powering General AI Computing
por: Xie, Yan, et al.
Publicado: (2026)
por: Xie, Yan, et al.
Publicado: (2026)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
por: Yang, Xiaoxuan, et al.
Publicado: (2025)
por: Yang, Xiaoxuan, et al.
Publicado: (2025)
CIMFlow: An Integrated Framework for Systematic Design and Evaluation of Digital CIM Architectures
por: Qi, Yingjie, et al.
Publicado: (2025)
por: Qi, Yingjie, et al.
Publicado: (2025)
FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction
por: Kulp, Gabriel, et al.
Publicado: (2024)
por: Kulp, Gabriel, et al.
Publicado: (2024)
NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction
por: Wadle, Shayne, et al.
Publicado: (2025)
por: Wadle, Shayne, et al.
Publicado: (2025)
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
por: Qiao, Ye, et al.
Publicado: (2025)
por: Qiao, Ye, et al.
Publicado: (2025)
Atleus: Accelerating Transformers on the Edge Enabled by 3D Heterogeneous Manycore Architectures
por: Dhingra, Pratyush, et al.
Publicado: (2025)
por: Dhingra, Pratyush, et al.
Publicado: (2025)
From LLM to Silicon: RL-Driven ASIC Architecture Exploration for On-Device AI Inference
por: Ganti, Ravindra, et al.
Publicado: (2026)
por: Ganti, Ravindra, et al.
Publicado: (2026)
AxMoE: Characterizing the Impact of Approximate Multipliers on Mixture-of-Experts DNN Architectures
por: Shende, Omkar B, et al.
Publicado: (2026)
por: Shende, Omkar B, et al.
Publicado: (2026)
LLM-VeriPPA: Power, Performance, and Area Optimization aware Verilog Code Generation with Large Language Models
por: Thorat, Kiran, et al.
Publicado: (2025)
por: Thorat, Kiran, et al.
Publicado: (2025)
PACiM: A Sparsity-Centric Hybrid Compute-in-Memory Architecture via Probabilistic Approximation
por: Zhang, Wenlun, et al.
Publicado: (2024)
por: Zhang, Wenlun, et al.
Publicado: (2024)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
por: Zhang, Jintao, et al.
Publicado: (2026)
por: Zhang, Jintao, et al.
Publicado: (2026)
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
por: Liu, Qingyuan, et al.
Publicado: (2025)
por: Liu, Qingyuan, et al.
Publicado: (2025)
Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications
por: Weng, Olivia, et al.
Publicado: (2024)
por: Weng, Olivia, et al.
Publicado: (2024)
Toward Attention-based TinyML: A Heterogeneous Accelerated Architecture and Automated Deployment Flow
por: Wiese, Philip, et al.
Publicado: (2024)
por: Wiese, Philip, et al.
Publicado: (2024)
The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
por: Morisaki, Keita
Publicado: (2026)
por: Morisaki, Keita
Publicado: (2026)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
por: Wu, Qizhe, et al.
Publicado: (2024)
por: Wu, Qizhe, et al.
Publicado: (2024)
Improving Simulation Regression Efficiency using a Machine Learning-based Method in Design Verification
por: Gadde, Deepak Narayan, et al.
Publicado: (2024)
por: Gadde, Deepak Narayan, et al.
Publicado: (2024)
TorR: Towards Brain-Inspired Task-Oriented Reasoning via Cache-Oriented Algorithm-Architecture Co-design
por: Oh, Hyunwoo, et al.
Publicado: (2026)
por: Oh, Hyunwoo, et al.
Publicado: (2026)
OptINC: Optical In-Network-Computing for Scalable Distributed Learning
por: Fei, Sijie, et al.
Publicado: (2026)
por: Fei, Sijie, et al.
Publicado: (2026)
H3DFact: Heterogeneous 3D Integrated CIM for Factorization with Holographic Perceptual Representations
por: Wan, Zishen, et al.
Publicado: (2024)
por: Wan, Zishen, et al.
Publicado: (2024)
An FPGA-Based SoC Architecture with a RISC-V Controller for Energy-Efficient Temporal-Coding Spiking Neural Networks
por: Sekonji, Mohammad Javad, et al.
Publicado: (2026)
por: Sekonji, Mohammad Javad, et al.
Publicado: (2026)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
por: Mueller, Lion, et al.
Publicado: (2025)
por: Mueller, Lion, et al.
Publicado: (2025)
Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
por: Pelke, Rebecca, et al.
Publicado: (2025)
por: Pelke, Rebecca, et al.
Publicado: (2025)
Ejemplares similares
-
Less is More: Hop-Wise Graph Attention for Scalable and Generalizable Learning on Circuits
por: Deng, Chenhui, et al.
Publicado: (2024) -
Federated Knowledge Distillation for Multi-Model Architectures Lithography Hotspot Detection
por: Li, Yuqi, et al.
Publicado: (2025) -
DeepSeq2: Enhanced Sequential Circuit Learning with Disentangled Representations
por: Khan, Sadaf, et al.
Publicado: (2024) -
Accelerating Computer Architecture Simulation through Machine Learning
por: Ali, Wajid, et al.
Publicado: (2024) -
Learning Library Cell Representations in Vector Space
por: Liang, Rongjian, et al.
Publicado: (2025)