Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Zhaolan, Schleiser, Kaspar, Myung, Gyungmin, Baccelli, Emmanuel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
U-TOE: Universal TinyML On-board Evaluation Toolkit for Low-Power IoT
di: Huang, Zhaolan, et al.
Pubblicazione: (2023)
di: Huang, Zhaolan, et al.
Pubblicazione: (2023)
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
di: Daghero, Francesco, et al.
Pubblicazione: (2025)
di: Daghero, Francesco, et al.
Pubblicazione: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
di: McDonald, Jesse, et al.
Pubblicazione: (2024)
di: McDonald, Jesse, et al.
Pubblicazione: (2024)
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
di: Dutta, Akash, et al.
Pubblicazione: (2024)
di: Dutta, Akash, et al.
Pubblicazione: (2024)
Distributed Matrix-Based Sampling for Graph Neural Network Training
di: Tripathy, Alok, et al.
Pubblicazione: (2023)
di: Tripathy, Alok, et al.
Pubblicazione: (2023)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
di: Zhang, Yaozheng, et al.
Pubblicazione: (2025)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
di: Sheng, Peiyao, et al.
Pubblicazione: (2024)
di: Sheng, Peiyao, et al.
Pubblicazione: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
di: Gupta, Ahan, et al.
Pubblicazione: (2026)
di: Gupta, Ahan, et al.
Pubblicazione: (2026)
Automated Programmatic Performance Analysis of Parallel Programs
di: Cankur, Onur, et al.
Pubblicazione: (2024)
di: Cankur, Onur, et al.
Pubblicazione: (2024)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
di: Pilliat, Emmanuel
Pubblicazione: (2026)
di: Pilliat, Emmanuel
Pubblicazione: (2026)
msf-CNN: Patch-based Multi-Stage Fusion with Convolutional Neural Networks for TinyML
di: Huang, Zhaolan, et al.
Pubblicazione: (2025)
di: Huang, Zhaolan, et al.
Pubblicazione: (2025)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
di: Anik, Md Saidul Hoque, et al.
Pubblicazione: (2024)
di: Anik, Md Saidul Hoque, et al.
Pubblicazione: (2024)
Optimal Parallel Scheduling under Concave Speedup Functions
di: Li, Chengzhang, et al.
Pubblicazione: (2025)
di: Li, Chengzhang, et al.
Pubblicazione: (2025)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
THAPI: Tracing Heterogeneous APIs
di: Bekele, Solomon, et al.
Pubblicazione: (2025)
di: Bekele, Solomon, et al.
Pubblicazione: (2025)
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
di: Chien, Steven W. D., et al.
Pubblicazione: (2024)
di: Chien, Steven W. D., et al.
Pubblicazione: (2024)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
di: Lacey, Dane C., et al.
Pubblicazione: (2024)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
di: Oztop, Beste, et al.
Pubblicazione: (2026)
di: Oztop, Beste, et al.
Pubblicazione: (2026)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
di: Dutt, Anurag, et al.
Pubblicazione: (2025)
di: Dutt, Anurag, et al.
Pubblicazione: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
di: Wang, Yuxin, et al.
Pubblicazione: (2023)
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
di: Banegas-Luna, Antonio Jesús, et al.
Pubblicazione: (2025)
di: Banegas-Luna, Antonio Jesús, et al.
Pubblicazione: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
di: Proaño, Andrès Rubio, et al.
Pubblicazione: (2024)
di: Proaño, Andrès Rubio, et al.
Pubblicazione: (2024)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
di: Papavasileiou, Ioannis, et al.
Pubblicazione: (2026)
di: Papavasileiou, Ioannis, et al.
Pubblicazione: (2026)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
di: Veeravalli, Bharadwaj
Pubblicazione: (2026)
di: Veeravalli, Bharadwaj
Pubblicazione: (2026)
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
di: Poptani, Akash, et al.
Pubblicazione: (2025)
di: Poptani, Akash, et al.
Pubblicazione: (2025)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
di: Nicusan, Andrei-Leonard, et al.
Pubblicazione: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
di: Ather, Hammad, et al.
Pubblicazione: (2024)
di: Ather, Hammad, et al.
Pubblicazione: (2024)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
di: Grbic, Dragana
Pubblicazione: (2026)
di: Grbic, Dragana
Pubblicazione: (2026)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
di: Titopoulos, Vasileios, et al.
Pubblicazione: (2025)
di: Titopoulos, Vasileios, et al.
Pubblicazione: (2025)
DIAL: Decentralized I/O AutoTuning via Learned Client-side Local Metrics for Parallel File System
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
di: Rashid, Md Hasanur, et al.
Pubblicazione: (2026)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
di: Thüring, Tim, et al.
Pubblicazione: (2026)
di: Thüring, Tim, et al.
Pubblicazione: (2026)
GPU Cluster Scheduling for Network-Sensitive Deep Learning
di: Sharma, Aakash, et al.
Pubblicazione: (2024)
di: Sharma, Aakash, et al.
Pubblicazione: (2024)
Energy-Aware Computing in the Year 2026
di: Tchakoute, Roblex Nana, et al.
Pubblicazione: (2026)
di: Tchakoute, Roblex Nana, et al.
Pubblicazione: (2026)
Multi-DNN Inference of Sparse Models on Edge SoCs
di: Luo, Jiawei, et al.
Pubblicazione: (2026)
di: Luo, Jiawei, et al.
Pubblicazione: (2026)
Hiku: Pull-Based Scheduling for Serverless Computing
di: Akbari, Saman, et al.
Pubblicazione: (2025)
di: Akbari, Saman, et al.
Pubblicazione: (2025)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
di: Lin, Zhongyi, et al.
Pubblicazione: (2024)
di: Lin, Zhongyi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
U-TOE: Universal TinyML On-board Evaluation Toolkit for Low-Power IoT
di: Huang, Zhaolan, et al.
Pubblicazione: (2023) -
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
di: Daghero, Francesco, et al.
Pubblicazione: (2025) -
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024) -
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
di: McDonald, Jesse, et al.
Pubblicazione: (2024) -
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
di: Dutta, Akash, et al.
Pubblicazione: (2024)