A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Yufan, De La Pierre, Marco, Barnard, Amanda S., Barca, Giuseppe Maria Junior |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
by: Xia, Yufan, et al.
Published: (2024)
by: Xia, Yufan, et al.
Published: (2024)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
by: Imani, HamidReza, et al.
Published: (2025)
by: Imani, HamidReza, et al.
Published: (2025)
FedTilt: Towards Multi-Level Fairness-Preserving and Robust Federated Learning
by: Zhang, Binghui, et al.
Published: (2025)
by: Zhang, Binghui, et al.
Published: (2025)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
by: Liu, Ruitao, et al.
Published: (2026)
by: Liu, Ruitao, et al.
Published: (2026)
Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications
by: Alshaer, Samer, et al.
Published: (2025)
by: Alshaer, Samer, et al.
Published: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
by: Lu, Yishun, et al.
Published: (2026)
by: Lu, Yishun, et al.
Published: (2026)
Towards a Flexible and High-Fidelity Approach to Distributed DNN Training Emulation
by: Liu, Banruo, et al.
Published: (2024)
by: Liu, Banruo, et al.
Published: (2024)
Have Your Cake and Eat It Too: Toward Efficient and Accurate Split Federated Learning
by: Yan, Dengke, et al.
Published: (2023)
by: Yan, Dengke, et al.
Published: (2023)
SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
Heterogeneity: An Open Challenge for Federated On-board Machine Learning
by: Hartmann, Maria, et al.
Published: (2024)
by: Hartmann, Maria, et al.
Published: (2024)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
by: Kokolis, Apostolos, et al.
Published: (2024)
by: Kokolis, Apostolos, et al.
Published: (2024)
Rashomon Sets and Model Multiplicity in Federated Learning
by: Heilmann, Xenia, et al.
Published: (2026)
by: Heilmann, Xenia, et al.
Published: (2026)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
Private, Augmentation-Robust and Task-Agnostic Data Valuation Approach for Data Marketplace
by: Jahani-Nezhad, Tayyebeh, et al.
Published: (2024)
by: Jahani-Nezhad, Tayyebeh, et al.
Published: (2024)
Learning Semantics, Not Addresses: Runtime Neural Prefetching for Far Memory
by: Huang, Yutong, et al.
Published: (2025)
by: Huang, Yutong, et al.
Published: (2025)
Machine Learning for Consistency Violation Faults Analysis
by: Giri, Kamal, et al.
Published: (2025)
by: Giri, Kamal, et al.
Published: (2025)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
by: Lin, Zhongyi, et al.
Published: (2024)
by: Lin, Zhongyi, et al.
Published: (2024)
Towards Client Driven Federated Learning
by: Li, Songze, et al.
Published: (2024)
by: Li, Songze, et al.
Published: (2024)
Provenance Tracking in Large-Scale Machine Learning Systems
by: Padovani, Gabriele, et al.
Published: (2025)
by: Padovani, Gabriele, et al.
Published: (2025)
ShardTensor: Domain Parallelism for Scientific Machine Learning
by: Adams, Corey, et al.
Published: (2026)
by: Adams, Corey, et al.
Published: (2026)
Training Machine Learning models at the Edge: A Survey
by: Khouas, Aymen Rayane, et al.
Published: (2024)
by: Khouas, Aymen Rayane, et al.
Published: (2024)
PiPar: Pipeline Parallelism for Collaborative Machine Learning
by: Zhang, Zihan, et al.
Published: (2022)
by: Zhang, Zihan, et al.
Published: (2022)
Algorithms for Collaborative Machine Learning under Statistical Heterogeneity
by: Hahn, Seok-Ju
Published: (2024)
by: Hahn, Seok-Ju
Published: (2024)
Towards Integrated Fine-tuning and Inference when Generative AI meets Edge Intelligence
by: Chen, Ning, et al.
Published: (2024)
by: Chen, Ning, et al.
Published: (2024)
Spectral Sentinel: Scalable Byzantine-Robust Decentralized Federated Learning via Sketched Random Matrix Theory on Blockchain
by: Mishra, Animesh
Published: (2025)
by: Mishra, Animesh
Published: (2025)
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
Towards Efficient Replay in Federated Incremental Learning
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
Unlearning during Learning: An Efficient Federated Machine Unlearning Method
by: Gu, Hanlin, et al.
Published: (2024)
by: Gu, Hanlin, et al.
Published: (2024)
Machine Learning-Based Research on the Adaptability of Adolescents to Online Education
by: Wang, Mingwei, et al.
Published: (2024)
by: Wang, Mingwei, et al.
Published: (2024)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
Dion2: A Simple Method to Shrink Matrix in Muon
by: Ahn, Kwangjun, et al.
Published: (2025)
by: Ahn, Kwangjun, et al.
Published: (2025)
Adaptive Federated Learning via New Entropy Approach
by: Zheng, Shensheng, et al.
Published: (2023)
by: Zheng, Shensheng, et al.
Published: (2023)
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
by: Zhou, Yuanchang, et al.
Published: (2026)
by: Zhou, Yuanchang, et al.
Published: (2026)
yProv4ML: Effortless Provenance Tracking for Machine Learning Systems
by: Padovani, Gabriele, et al.
Published: (2025)
by: Padovani, Gabriele, et al.
Published: (2025)
Adaptive Resolution Inference (ARI): Energy-Efficient Machine Learning for Internet of Things
by: Wang, Ziheng, et al.
Published: (2024)
by: Wang, Ziheng, et al.
Published: (2024)
Efficient Distributed Learning over Decentralized Networks with Convoluted Support Vector Machine
by: Chen, Canyi, et al.
Published: (2025)
by: Chen, Canyi, et al.
Published: (2025)
Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems
by: Furutanpey, Alireza, et al.
Published: (2025)
by: Furutanpey, Alireza, et al.
Published: (2025)
Similar Items
-
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
by: Xia, Yufan, et al.
Published: (2024) -
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
by: Imani, HamidReza, et al.
Published: (2025) -
FedTilt: Towards Multi-Level Fairness-Preserving and Robust Federated Learning
by: Zhang, Binghui, et al.
Published: (2025) -
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
by: Liu, Ruitao, et al.
Published: (2026) -
Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications
by: Alshaer, Samer, et al.
Published: (2025)