A High-Level Compiler Integration Approach for Deep Learning Accelerators Supporting Abstraction and Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Ahmadifarsani, Samira, Mueller-Gritschneder, Daniel, Schlichtmann, Ulf |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MLonMCU: TinyML Benchmarking with Fast Retargeting
por: van Kempen, Philipp, et al.
Publicado: (2023)
por: van Kempen, Philipp, et al.
Publicado: (2023)
Hybrid Convolution and Vision Transformer NAS Search Space for TinyML Image Classification
por: Djajapermana, Mikhael, et al.
Publicado: (2025)
por: Djajapermana, Mikhael, et al.
Publicado: (2025)
A Continual and Incremental Learning Approach for TinyML On-device Training Using Dataset Distillation and Model Size Adaption
por: Rüb, Marcus, et al.
Publicado: (2024)
por: Rüb, Marcus, et al.
Publicado: (2024)
Advancing On-Device Neural Network Training with TinyPropv2: Dynamic, Sparse, and Efficient Backpropagation
por: Rüb, Marcus, et al.
Publicado: (2024)
por: Rüb, Marcus, et al.
Publicado: (2024)
DRIP: DRop unImportant data Points -- Enhancing Machine Learning Efficiency with Grad-CAM-Based Real-Time Data Prioritization for On-Device Training
por: Rüb, Marcus, et al.
Publicado: (2025)
por: Rüb, Marcus, et al.
Publicado: (2025)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
por: Wolters, Christopher, et al.
Publicado: (2024)
por: Wolters, Christopher, et al.
Publicado: (2024)
PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
por: Ding, Ruogu, et al.
Publicado: (2025)
por: Ding, Ruogu, et al.
Publicado: (2025)
OptINC: Optical In-Network-Computing for Scalable Distributed Learning
por: Fei, Sijie, et al.
Publicado: (2026)
por: Fei, Sijie, et al.
Publicado: (2026)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
por: Li, Jianhui, et al.
Publicado: (2023)
por: Li, Jianhui, et al.
Publicado: (2023)
EncodingNet: A Novel Encoding-based MAC Design for Efficient Neural Network Acceleration
por: Liu, Bo, et al.
Publicado: (2024)
por: Liu, Bo, et al.
Publicado: (2024)
BasisN: Reprogramming-Free RRAM-Based In-Memory-Computing by Basis Combination for Deep Neural Networks
por: Eldebiky, Amro, et al.
Publicado: (2024)
por: Eldebiky, Amro, et al.
Publicado: (2024)
ACRoBat: Optimizing Auto-batching of Dynamic Deep Learning at Compile Time
por: Fegade, Pratik, et al.
Publicado: (2023)
por: Fegade, Pratik, et al.
Publicado: (2023)
EGIC: Enhanced Low-Bit-Rate Generative Image Compression Guided by Semantic Segmentation
por: Körber, Nikolai, et al.
Publicado: (2023)
por: Körber, Nikolai, et al.
Publicado: (2023)
Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
por: Hou, Bohan, et al.
Publicado: (2026)
por: Hou, Bohan, et al.
Publicado: (2026)
Learning Markov State Abstractions for Deep Reinforcement Learning
por: Allen, Cameron, et al.
Publicado: (2021)
por: Allen, Cameron, et al.
Publicado: (2021)
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
por: Jin, Hongyi, et al.
Publicado: (2026)
por: Jin, Hongyi, et al.
Publicado: (2026)
XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
por: Palenicek, Daniel, et al.
Publicado: (2025)
por: Palenicek, Daniel, et al.
Publicado: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
por: Chen, Chuangtao, et al.
Publicado: (2026)
por: Chen, Chuangtao, et al.
Publicado: (2026)
Deep Learning-Accelerated Surrogate Optimization for High-Dimensional Well Control in Stress-Sensitive Reservoirs
por: Valiyev, Mahammad, et al.
Publicado: (2026)
por: Valiyev, Mahammad, et al.
Publicado: (2026)
A Unified Perspective for Learning Graph Representations Across Multi-Level Abstractions
por: Amar, Mohamed Mahmoud, et al.
Publicado: (2026)
por: Amar, Mohamed Mahmoud, et al.
Publicado: (2026)
A Stochastic Approach to Bi-Level Optimization for Hyperparameter Optimization and Meta Learning
por: Kim, Minyoung, et al.
Publicado: (2024)
por: Kim, Minyoung, et al.
Publicado: (2024)
CompilerDream: Learning a Compiler World Model for General Code Optimization
por: Deng, Chaoyi, et al.
Publicado: (2024)
por: Deng, Chaoyi, et al.
Publicado: (2024)
SENSEi: Input-Sensitive Compilation for Accelerating GNNs
por: Lenadora, Damitha, et al.
Publicado: (2023)
por: Lenadora, Damitha, et al.
Publicado: (2023)
Self-Abstraction Learning for Effective and Stable Training of Deep Neural Networks
por: Cho, Wonyong, et al.
Publicado: (2026)
por: Cho, Wonyong, et al.
Publicado: (2026)
Enhancing Model Fairness and Accuracy with Similarity Networks: A Methodological Approach
por: Maghool, Samira, et al.
Publicado: (2024)
por: Maghool, Samira, et al.
Publicado: (2024)
Reducing Bias in Deep Learning Optimization: The RSGDM Approach
por: Qin, Honglin, et al.
Publicado: (2024)
por: Qin, Honglin, et al.
Publicado: (2024)
Dynamic Co-Optimization Compiler: Leveraging Multi-Agent Reinforcement Learning for Enhanced DNN Accelerator Performance
por: Fayyazi, Arya, et al.
Publicado: (2024)
por: Fayyazi, Arya, et al.
Publicado: (2024)
Multiple Abstraction Level Retrieve Augment Generation
por: Zheng, Zheng, et al.
Publicado: (2025)
por: Zheng, Zheng, et al.
Publicado: (2025)
An Invariant Compiler for Neural ODEs in AI-Accelerated Scientific Simulation
por: Yu, Fangzhou, et al.
Publicado: (2026)
por: Yu, Fangzhou, et al.
Publicado: (2026)
Deep Reinforcement Learning: A Convex Optimization Approach
por: Gattami, Ather
Publicado: (2024)
por: Gattami, Ather
Publicado: (2024)
A Triple-Inertial Accelerated Alternating Optimization Method for Deep Learning Training
por: Yan, Chengcheng, et al.
Publicado: (2025)
por: Yan, Chengcheng, et al.
Publicado: (2025)
Contrastive Abstraction for Reinforcement Learning
por: Patil, Vihang, et al.
Publicado: (2024)
por: Patil, Vihang, et al.
Publicado: (2024)
Temporal Abstraction in Reinforcement Learning with Offline Data
por: Ayyagari, Ranga Shaarad, et al.
Publicado: (2024)
por: Ayyagari, Ranga Shaarad, et al.
Publicado: (2024)
LLM-Aided Compilation for Tensor Accelerators
por: Hong, Charles, et al.
Publicado: (2024)
por: Hong, Charles, et al.
Publicado: (2024)
Development of Deep Learning Optimizers: Approaches, Concepts, and Update Rules
por: Altınel, Doğay
Publicado: (2025)
por: Altınel, Doğay
Publicado: (2025)
Neuro-Symbolic Imitation Learning: Discovering Symbolic Abstractions for Skill Learning
por: Keller, Leon, et al.
Publicado: (2025)
por: Keller, Leon, et al.
Publicado: (2025)
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
por: Chen, Simin, et al.
Publicado: (2025)
por: Chen, Simin, et al.
Publicado: (2025)
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
por: Yau, Chung-Yiu, et al.
Publicado: (2026)
por: Yau, Chung-Yiu, et al.
Publicado: (2026)
Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement Learning
por: Azran, Guy, et al.
Publicado: (2023)
por: Azran, Guy, et al.
Publicado: (2023)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
por: You, Bozhi, et al.
Publicado: (2025)
por: You, Bozhi, et al.
Publicado: (2025)
Ejemplares similares
-
MLonMCU: TinyML Benchmarking with Fast Retargeting
por: van Kempen, Philipp, et al.
Publicado: (2023) -
Hybrid Convolution and Vision Transformer NAS Search Space for TinyML Image Classification
por: Djajapermana, Mikhael, et al.
Publicado: (2025) -
A Continual and Incremental Learning Approach for TinyML On-device Training Using Dataset Distillation and Model Size Adaption
por: Rüb, Marcus, et al.
Publicado: (2024) -
Advancing On-Device Neural Network Training with TinyPropv2: Dynamic, Sparse, and Efficient Backpropagation
por: Rüb, Marcus, et al.
Publicado: (2024) -
DRIP: DRop unImportant data Points -- Enhancing Machine Learning Efficiency with Grad-CAM-Based Real-Time Data Prioritization for On-Device Training
por: Rüb, Marcus, et al.
Publicado: (2025)