Balanced segmentation of CNNs for multi-TPU inference
Fuente:
arXiv
Saved in:
| Main Authors: | Villarrubia, Jorge, Costero, Luis, Igual, Francisco D., Olcoz, Katzalin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving inference time in multi-TPU systems with profiled model segmentation
by: Villarrubia, Jorge, et al.
Published: (2025)
by: Villarrubia, Jorge, et al.
Published: (2025)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
by: Villarrubia, Jorge, et al.
Published: (2026)
by: Villarrubia, Jorge, et al.
Published: (2026)
Leveraging Multi-Instance GPUs through moldable task scheduling
by: Villarrubia, Jorge, et al.
Published: (2025)
by: Villarrubia, Jorge, et al.
Published: (2025)
Leveraging knowledge-as-a-service (KaaS) for QoS-aware resource management in multi-user video transcoding
by: Costero, Luis, et al.
Published: (2024)
by: Costero, Luis, et al.
Published: (2024)
Energy efficiency optimization of task-parallel codes on asymmetric architectures
by: Costero, Luis, et al.
Published: (2024)
by: Costero, Luis, et al.
Published: (2024)
PIM-STM: Software Transactional Memory for Processing-In-Memory Systems
by: Lopes, André, et al.
Published: (2024)
by: Lopes, André, et al.
Published: (2024)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
by: Ferreira, João Dinis, et al.
Published: (2021)
by: Ferreira, João Dinis, et al.
Published: (2021)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
Stream parallel skeleton optimization
by: Aldinucci, Marco, et al.
Published: (2024)
by: Aldinucci, Marco, et al.
Published: (2024)
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020)
by: Colonnelli, Iacopo, et al.
Published: (2020)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022)
by: Li, Chendi, et al.
Published: (2022)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024)
by: Xu, Yao, et al.
Published: (2024)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
by: de Castro, Manuel, et al.
Published: (2024)
by: de Castro, Manuel, et al.
Published: (2024)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
by: Gallone, Anna, et al.
Published: (2025)
by: Gallone, Anna, et al.
Published: (2025)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
by: Mantha, Pradeep, et al.
Published: (2026)
by: Mantha, Pradeep, et al.
Published: (2026)
Work-Efficient Parallel Non-Maximum Suppression Kernels
by: Oro, David, et al.
Published: (2025)
by: Oro, David, et al.
Published: (2025)
Aurora: Architecting Argonne's First Exascale Supercomputer for Accelerated Scientific Discovery
by: Allcock, William E., et al.
Published: (2025)
by: Allcock, William E., et al.
Published: (2025)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
N2N: A Parallel Framework for Large-Scale MILP under Distributed Memory
by: Wang, Longfei, et al.
Published: (2025)
by: Wang, Longfei, et al.
Published: (2025)
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
by: Singh, Ajay, et al.
Published: (2026)
by: Singh, Ajay, et al.
Published: (2026)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
Construction of a Byzantine Linearizable SWMR Atomic Register from SWSR Atomic Registers
by: Kshemkalyani, Ajay D., et al.
Published: (2024)
by: Kshemkalyani, Ajay D., et al.
Published: (2024)
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
A C++17 Thread Pool for High-Performance Scientific Computing
by: Shoshany, Barak
Published: (2021)
by: Shoshany, Barak
Published: (2021)
Scalable Concurrent Queues for GPU
by: Shetty, Pratheek Prakash, et al.
Published: (2026)
by: Shetty, Pratheek Prakash, et al.
Published: (2026)
SpaDA: A Spatial Dataflow Architecture Programming Language
by: Gianinazzi, Lukas, et al.
Published: (2025)
by: Gianinazzi, Lukas, et al.
Published: (2025)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
by: Li, Shigang, et al.
Published: (2020)
by: Li, Shigang, et al.
Published: (2020)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
On Advanced Monte Carlo Methods for Linear Algebra on Advanced Accelerator Architectures
by: Lebedev, Anton, et al.
Published: (2024)
by: Lebedev, Anton, et al.
Published: (2024)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Hardware-Level QoS Enforcement Features: Technologies, Use Cases, and Research Challenges
by: Larsson, Oliver, et al.
Published: (2025)
by: Larsson, Oliver, et al.
Published: (2025)
Rust vs. C for Python Libraries: Evaluating Rust-Compatible Bindings Toolchains
by: Amaral, Isabella Basso do, et al.
Published: (2025)
by: Amaral, Isabella Basso do, et al.
Published: (2025)
GPU-centric Communication Schemes for HPC and ML Applications
by: Namashivayam, Naveen
Published: (2025)
by: Namashivayam, Naveen
Published: (2025)
ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference
by: Zhu, Xiongwei, et al.
Published: (2026)
by: Zhu, Xiongwei, et al.
Published: (2026)
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication
by: Wang, Zongwu, et al.
Published: (2024)
by: Wang, Zongwu, et al.
Published: (2024)
A Performance Analysis of You Only Look Once Models for Deployment on Constrained Computational Edge Devices in Drone Applications
by: Rey, Lucas, et al.
Published: (2025)
by: Rey, Lucas, et al.
Published: (2025)
Shortest paths search method based on the projective description of unweighted mixed graphs
by: Melent'ev, V. A.
Published: (2023)
by: Melent'ev, V. A.
Published: (2023)
Similar Items
-
Improving inference time in multi-TPU systems with profiled model segmentation
by: Villarrubia, Jorge, et al.
Published: (2025) -
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
by: Villarrubia, Jorge, et al.
Published: (2026) -
Leveraging Multi-Instance GPUs through moldable task scheduling
by: Villarrubia, Jorge, et al.
Published: (2025) -
Leveraging knowledge-as-a-service (KaaS) for QoS-aware resource management in multi-user video transcoding
by: Costero, Luis, et al.
Published: (2024) -
Energy efficiency optimization of task-parallel codes on asymmetric architectures
by: Costero, Luis, et al.
Published: (2024)