Batched DGEMMs for scientific codes running on long vector architectures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Banchelli, Fabio, Garcia-Gasulla, Marta, Mantovani, Filippo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploiting long vectors with a CFD code: a co-design show case
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025)
TALP-Pages: An easy-to-integrate continuous performance monitoring framework
von: Seitz, Valentin, et al.
Veröffentlicht: (2025)
von: Seitz, Valentin, et al.
Veröffentlicht: (2025)
Leveraging HPC Profiling & Tracing Tools to Understand the Performance of Particle-in-Cell Monte Carlo Simulations
von: Williams, Jeremy J., et al.
Veröffentlicht: (2023)
von: Williams, Jeremy J., et al.
Veröffentlicht: (2023)
Energy efficiency optimization of task-parallel codes on asymmetric architectures
von: Costero, Luis, et al.
Veröffentlicht: (2024)
von: Costero, Luis, et al.
Veröffentlicht: (2024)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
von: Rahimi, Ghazal, et al.
Veröffentlicht: (2026)
CRIU -- Checkpoint Restore in Userspace for computational simulations and scientific applications
von: Andrijauskas, Fabio, et al.
Veröffentlicht: (2024)
von: Andrijauskas, Fabio, et al.
Veröffentlicht: (2024)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
Towards observability of scientific applications
von: Balis, Bartosz, et al.
Veröffentlicht: (2024)
von: Balis, Bartosz, et al.
Veröffentlicht: (2024)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
Towards cloud-native scientific workflow management
von: Orzechowski, Michal, et al.
Veröffentlicht: (2024)
von: Orzechowski, Michal, et al.
Veröffentlicht: (2024)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
Batch Denoising for AIGC Service Provisioning in Wireless Edge Networks
von: Xu, Jinghang, et al.
Veröffentlicht: (2025)
von: Xu, Jinghang, et al.
Veröffentlicht: (2025)
Are Your Epochs Too Epic? Batch Free Can Be Harmful
von: Kim, Daewoo, et al.
Veröffentlicht: (2024)
von: Kim, Daewoo, et al.
Veröffentlicht: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
A Reinforcement Learning Based Backfilling Strategy for HPC Batch Jobs
von: Kolker-Hicks, Elliot, et al.
Veröffentlicht: (2024)
von: Kolker-Hicks, Elliot, et al.
Veröffentlicht: (2024)
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
Herring: Parallel Batch-Order-Fairness on DAG-based Blockchain Consensus
von: Putnik, Marko, et al.
Veröffentlicht: (2026)
von: Putnik, Marko, et al.
Veröffentlicht: (2026)
Extension of ACETONE C code generator for multi-core architectures
von: Aït-Aïssa, Yanis, et al.
Veröffentlicht: (2026)
von: Aït-Aïssa, Yanis, et al.
Veröffentlicht: (2026)
Fog enabled distributed training architecture for federated learning
von: Kumar, Aditya, et al.
Veröffentlicht: (2024)
von: Kumar, Aditya, et al.
Veröffentlicht: (2024)
An approach to provide serverless scientific pipelines within the context of SKA
von: Ríos-Monje, Carlos, et al.
Veröffentlicht: (2023)
von: Ríos-Monje, Carlos, et al.
Veröffentlicht: (2023)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
Lifting to tensors when compiling scientific computing workloads for AI Engines
von: Brown, Nick, et al.
Veröffentlicht: (2026)
von: Brown, Nick, et al.
Veröffentlicht: (2026)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
von: Sakip, Akhmed, et al.
Veröffentlicht: (2026)
von: Sakip, Akhmed, et al.
Veröffentlicht: (2026)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
von: Duan, Tao, et al.
Veröffentlicht: (2025)
von: Duan, Tao, et al.
Veröffentlicht: (2025)
Batch Query Processing and Optimization for Agentic Workflows
von: Shen, Junyi, et al.
Veröffentlicht: (2025)
von: Shen, Junyi, et al.
Veröffentlicht: (2025)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
GPU-Accelerated Vecchia Approximations of Gaussian Processes for Geospatial Data using Batched Matrix Computations
von: Pan, Qilong, et al.
Veröffentlicht: (2024)
von: Pan, Qilong, et al.
Veröffentlicht: (2024)
Efficiently Parallelizable Strassen-Based Multiplication of a Matrix by its Transpose
von: Arrigoni, Viviana, et al.
Veröffentlicht: (2021)
von: Arrigoni, Viviana, et al.
Veröffentlicht: (2021)
Batch-Schedule-Execute: On Optimizing Concurrent Deterministic Scheduling for Blockchains (Extended Version)
von: Hay, Yaron, et al.
Veröffentlicht: (2024)
von: Hay, Yaron, et al.
Veröffentlicht: (2024)
Rafture: Erasure-coded Raft with Post-Dissemination Pruning
von: Kerur, Rithwik, et al.
Veröffentlicht: (2026)
von: Kerur, Rithwik, et al.
Veröffentlicht: (2026)
Cross-architecture universal feature coding via distribution alignment
von: Gao, Changsheng, et al.
Veröffentlicht: (2025)
von: Gao, Changsheng, et al.
Veröffentlicht: (2025)
Neuro-Inspired Task Offloading in Edge-IoT Networks Using Spiking Neural Networks
von: Rossi, Fabio Diniz
Veröffentlicht: (2025)
von: Rossi, Fabio Diniz
Veröffentlicht: (2025)
Stochastic Modeling for Energy-Efficient Edge Infrastructure
von: Rossi, Fabio Diniz
Veröffentlicht: (2025)
von: Rossi, Fabio Diniz
Veröffentlicht: (2025)
DMRlib: Easy-coding and Efficient Resource Management for Job Malleability
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Exploiting long vectors with a CFD code: a co-design show case
von: Blancafort, Marc, et al.
Veröffentlicht: (2024) -
Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
von: Banchelli, Fabio, et al.
Veröffentlicht: (2025) -
TALP-Pages: An easy-to-integrate continuous performance monitoring framework
von: Seitz, Valentin, et al.
Veröffentlicht: (2025) -
Leveraging HPC Profiling & Tracing Tools to Understand the Performance of Particle-in-Cell Monte Carlo Simulations
von: Williams, Jeremy J., et al.
Veröffentlicht: (2023) -
Energy efficiency optimization of task-parallel codes on asymmetric architectures
von: Costero, Luis, et al.
Veröffentlicht: (2024)