AEG: A Baremetal Framework for AI Acceleration via Direct Hardware Access in Heterogeneous Accelerators
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Hua, Mandal, Sayan, Kirincich, Brandon, Varadarajan, Govind |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gathering Semi-Synchronously Scheduled Two-State Robots
by: Otaka, Kohei, et al.
Published: (2024)
by: Otaka, Kohei, et al.
Published: (2024)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
Method for determining the acceleration of a parallel specialised computer system based on Amdahl's law
by: Filipchenko, Aleksandr S.
Published: (2024)
by: Filipchenko, Aleksandr S.
Published: (2024)
The Impact of Partial Computations on the Red-Blue Pebble Game
by: Papp, Pál András, et al.
Published: (2025)
by: Papp, Pál András, et al.
Published: (2025)
Red-Blue Pebbling with Multiple Processors: Time, Communication and Memory Trade-offs
by: Böhnlein, Toni, et al.
Published: (2024)
by: Böhnlein, Toni, et al.
Published: (2024)
Characterising resource management performance in Kubernetes
by: Medel, Víctor, et al.
Published: (2024)
by: Medel, Víctor, et al.
Published: (2024)
GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN
by: Yin, Jia, et al.
Published: (2025)
by: Yin, Jia, et al.
Published: (2025)
KAITIAN: A Unified Communication Framework for Enabling Efficient Collaboration Across Heterogeneous Accelerators in Embodied AI Systems
by: Lin, Jieke, et al.
Published: (2025)
by: Lin, Jieke, et al.
Published: (2025)
Multiprocessor Scheduling with Memory Constraints: Fundamental Properties and Finding Optimal Solutions
by: Papp, Pál András, et al.
Published: (2025)
by: Papp, Pál András, et al.
Published: (2025)
Tiga: Accelerating Geo-Distributed Transactions with Synchronized Clocks [Technical Report]
by: Geng, Jinkun, et al.
Published: (2025)
by: Geng, Jinkun, et al.
Published: (2025)
Mind the Gap: Revealing Inconsistencies Across Heterogeneous AI Accelerators
by: Wen, Elliott, et al.
Published: (2025)
by: Wen, Elliott, et al.
Published: (2025)
Accelerating Microswimmer Simulations via a Heterogeneous Pipelined Parallel-in-Time Framework
by: Huang, Ruixiang, et al.
Published: (2026)
by: Huang, Ruixiang, et al.
Published: (2026)
Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
by: Reisecker, Maximilian, et al.
Published: (2025)
by: Reisecker, Maximilian, et al.
Published: (2025)
Deterministic Distributed DFS and Other Problems via Cycle Separators in Planar Graphs
by: Jauregui, Benjamin, et al.
Published: (2025)
by: Jauregui, Benjamin, et al.
Published: (2025)
Enhancing Cluster Resilience: LLM-agent Based Autonomous Intelligent Cluster Diagnosis System and Evaluation Framework
by: Shi, Honghao, et al.
Published: (2024)
by: Shi, Honghao, et al.
Published: (2024)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
by: Wang, Hansheng, et al.
Published: (2024)
by: Wang, Hansheng, et al.
Published: (2024)
A Morton-Type Space-Filling Curve for Pyramid Subdivision and Hybrid Adaptive Mesh Refinement
by: Knapp, David, et al.
Published: (2026)
by: Knapp, David, et al.
Published: (2026)
GPU-Accelerated Algorithms for Process Mapping
by: Samoldekin, Petr, et al.
Published: (2025)
by: Samoldekin, Petr, et al.
Published: (2025)
Algorithms for Parallel Shared-Memory Sparse Matrix-Vector Multiplication on Unstructured Matrices
by: Bergmans, Kobe, et al.
Published: (2025)
by: Bergmans, Kobe, et al.
Published: (2025)
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
by: Park, Jeongmin Brian, et al.
Published: (2023)
by: Park, Jeongmin Brian, et al.
Published: (2023)
iOS as Acceleration
by: Chen, Alexander K.
Published: (2025)
by: Chen, Alexander K.
Published: (2025)
InTec: integrated things-edge computing: a framework for distributing machine learning pipelines in edge AI systems
by: Larian, Habib, et al.
Published: (2025)
by: Larian, Habib, et al.
Published: (2025)
Analysing cycloids using linear algebra
by: Valk, Rüdiger
Published: (2024)
by: Valk, Rüdiger
Published: (2024)
Machine Learning-Driven Predictive Resource Management in Complex Science Workflows
by: Chowdhury, Tasnuva, et al.
Published: (2025)
by: Chowdhury, Tasnuva, et al.
Published: (2025)
Accelerated Digital Twin Learning for Edge AI: A Comparison of FPGA and Mobile GPU
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
Keep Your Friends Close: Leveraging Affinity Groups to Accelerate AI Inference Workflows
by: Garrett, Thiago, et al.
Published: (2023)
by: Garrett, Thiago, et al.
Published: (2023)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
by: Wang, Tianze, et al.
Published: (2025)
by: Wang, Tianze, et al.
Published: (2025)
Towards Scalable GPU-Accelerated SNN Training via Temporal Fusion
by: Li, Yanchen, et al.
Published: (2024)
by: Li, Yanchen, et al.
Published: (2024)
A Unified Programming Model for Heterogeneous Computing with CPU and Accelerator Technologies
by: Xiong, Yuqing
Published: (2022)
by: Xiong, Yuqing
Published: (2022)
SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
by: Chen, Fahao, et al.
Published: (2025)
by: Chen, Fahao, et al.
Published: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
by: Shakeri, Heman, et al.
Published: (2026)
by: Shakeri, Heman, et al.
Published: (2026)
Accelerating Latency-Critical Applications with AI-Powered Semi-Automatic Fine-Grained Parallelization on SMT Processors
by: Los, Denis, et al.
Published: (2025)
by: Los, Denis, et al.
Published: (2025)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
by: Chen, Le, et al.
Published: (2025)
by: Chen, Le, et al.
Published: (2025)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
by: Zhao, Long, et al.
Published: (2026)
by: Zhao, Long, et al.
Published: (2026)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
Accelerating LLM Inference with Precomputed Query Storage
by: Park, Jay H., et al.
Published: (2025)
by: Park, Jay H., et al.
Published: (2025)
Federated Learning within Global Energy Budget over Heterogeneous Edge Accelerators
by: Banerjee, Roopkatha, et al.
Published: (2025)
by: Banerjee, Roopkatha, et al.
Published: (2025)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
Act While Thinking: Accelerating LLM Agents via Pattern-Aware Speculative Tool Execution
by: Sui, Yifan, et al.
Published: (2026)
by: Sui, Yifan, et al.
Published: (2026)
Similar Items
-
Gathering Semi-Synchronously Scheduled Two-State Robots
by: Otaka, Kohei, et al.
Published: (2024) -
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025) -
Method for determining the acceleration of a parallel specialised computer system based on Amdahl's law
by: Filipchenko, Aleksandr S.
Published: (2024) -
The Impact of Partial Computations on the Red-Blue Pebble Game
by: Papp, Pál András, et al.
Published: (2025) -
Red-Blue Pebbling with Multiple Processors: Time, Communication and Memory Trade-offs
by: Böhnlein, Toni, et al.
Published: (2024)