TOAST: Fast and scalable auto-partitioning based on principled static analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Alabed, Sami, Grewe, Dominik, Rink, Norman Alexander, Samsikova, Masha, Sitdikov, Timur, Swietlik, Agnieszka, Vytiniotis, Dimitrios, Belov, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PartIR: Composing SPMD Partitioning Strategies for Machine Learning
by: Alabed, Sami, et al.
Published: (2024)
by: Alabed, Sami, et al.
Published: (2024)
HPC resources for CMS offline computing: An integration and scalability challenge for the Submission Infrastructure
by: Yzquierdo, Antonio Perez-Calero, et al.
Published: (2024)
by: Yzquierdo, Antonio Perez-Calero, et al.
Published: (2024)
Flotilla: A scalable, modular and resilient federated learning framework for heterogeneous resources
by: Banerjee, Roopkatha, et al.
Published: (2025)
by: Banerjee, Roopkatha, et al.
Published: (2025)
Trustworthy Scheduling for Big Data Applications
by: Tomaras, Dimitrios, et al.
Published: (2026)
by: Tomaras, Dimitrios, et al.
Published: (2026)
TIMBER: On supporting data pipelines in Mobile Cloud Environments
by: Tomaras, Dimitrios, et al.
Published: (2024)
by: Tomaras, Dimitrios, et al.
Published: (2024)
FirecREST v2: lessons learned from redesigning an API for scalable HPC resource access
by: Palme, Elia, et al.
Published: (2025)
by: Palme, Elia, et al.
Published: (2025)
Dynamic Size Counting in the Population Protocol Model
by: Kaaser, Dominik, et al.
Published: (2024)
by: Kaaser, Dominik, et al.
Published: (2024)
Prediction-driven resource provisioning for serverless container runtimes
by: Tomaras, Dimitrios, et al.
Published: (2024)
by: Tomaras, Dimitrios, et al.
Published: (2024)
Fast-HotStuff: A Fast and Resilient HotStuff Protocol
by: Jalalzai, Mohammad M., et al.
Published: (2020)
by: Jalalzai, Mohammad M., et al.
Published: (2020)
FastGraph: Optimized GPU-Enabled Algorithms for Fast Graph Building and Message Passing
by: Agarwal, Aarush, et al.
Published: (2025)
by: Agarwal, Aarush, et al.
Published: (2025)
A Fast Confirmation Rule (aka Fast Synchronous Finality) for the Ethereum Consensus Protocol
by: Asgaonkar, Aditya, et al.
Published: (2024)
by: Asgaonkar, Aditya, et al.
Published: (2024)
Scalable mRMR feature selection to handle high dimensional datasets: Vertical partitioning based Iterative MapReduce framework
by: Vivek, Yelleti, et al.
Published: (2022)
by: Vivek, Yelleti, et al.
Published: (2022)
Taming the Memory Footprint Crisis: System Design for Production Diffusion LLM Serving
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
Leveraging Core and Uncore Frequency Scaling for Power-Efficient Serverless Workflows
by: Tzenetopoulos, Achilleas, et al.
Published: (2024)
by: Tzenetopoulos, Achilleas, et al.
Published: (2024)
Fast Byzantine Total Order Broadcast
by: Monti, Matteo, et al.
Published: (2024)
by: Monti, Matteo, et al.
Published: (2024)
FastSet: Parallel Claim Settlement
by: Chen, Xiaohong, et al.
Published: (2025)
by: Chen, Xiaohong, et al.
Published: (2025)
Fast Transaction Scheduling in Blockchain Sharding
by: Adhikari, Ramesh, et al.
Published: (2024)
by: Adhikari, Ramesh, et al.
Published: (2024)
Asynchronous Latency and Fast Atomic Snapshot
by: Bezerra, João Paulo, et al.
Published: (2024)
by: Bezerra, João Paulo, et al.
Published: (2024)
Prime Collective Communications Library -- Technical Report
by: Keiblinger, Michael, et al.
Published: (2025)
by: Keiblinger, Michael, et al.
Published: (2025)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Modality Inflation: Energy Characterization and Optimization Opportunities for MLLM Inference
by: Moghadampanah, Mona, et al.
Published: (2025)
by: Moghadampanah, Mona, et al.
Published: (2025)
Cabinet: Dynamically Weighted Consensus Made Fast
by: Zhang, Gengrui, et al.
Published: (2025)
by: Zhang, Gengrui, et al.
Published: (2025)
Fast State Restoration in LLM Serving with HCache
by: Gao, Shiwei, et al.
Published: (2024)
by: Gao, Shiwei, et al.
Published: (2024)
Minimmit: Fast Finality with Even Faster Blocks
by: Chou, Brendan Kobayashi, et al.
Published: (2025)
by: Chou, Brendan Kobayashi, et al.
Published: (2025)
Mangrove: Fast and Parallelizable State Replication for Blockchains
by: Paramonov, Anton, et al.
Published: (2025)
by: Paramonov, Anton, et al.
Published: (2025)
Exploring Fast Fourier Transforms on the Tenstorrent Wormhole
by: Brown, Nick, et al.
Published: (2025)
by: Brown, Nick, et al.
Published: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
by: Jangda, Abhinav, et al.
Published: (2024)
by: Jangda, Abhinav, et al.
Published: (2024)
Carbon-Aware Mapping and Scheduling for Deadline-Constrained Workflows
by: Schweisgut, Dominik, et al.
Published: (2026)
by: Schweisgut, Dominik, et al.
Published: (2026)
Carbon-Aware Workflow Scheduling with Fixed Mapping and Deadline Constraint
by: Schweisgut, Dominik, et al.
Published: (2025)
by: Schweisgut, Dominik, et al.
Published: (2025)
Demeter: Resource-Efficient Distributed Stream Processing under Dynamic Loads with Multi-Configuration Optimization
by: Geldenhuys, Morgan, et al.
Published: (2024)
by: Geldenhuys, Morgan, et al.
Published: (2024)
ZEUS: An Efficient GPU Optimization Method Integrating PSO, BFGS, and Automatic Differentiation
by: Soos, Dominik, et al.
Published: (2026)
by: Soos, Dominik, et al.
Published: (2026)
SmartPQ: An Adaptive Concurrent Priority Queue for NUMA Architectures
by: Giannoula, Christina, et al.
Published: (2024)
by: Giannoula, Christina, et al.
Published: (2024)
Fast and Robust Information Spreading in the Noisy PULL Model
by: D'Archivio, Niccolò, et al.
Published: (2024)
by: D'Archivio, Niccolò, et al.
Published: (2024)
Towards a Scalable In Situ Fast Fourier Transform
by: Kulkarni, Sudhanshu, et al.
Published: (2024)
by: Kulkarni, Sudhanshu, et al.
Published: (2024)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Fast Iterative Graph Computing with Updated Neighbor States
by: Zhou, Yijie, et al.
Published: (2024)
by: Zhou, Yijie, et al.
Published: (2024)
Shoal++: High Throughput DAG BFT Can Be Fast!
by: Arun, Balaji, et al.
Published: (2024)
by: Arun, Balaji, et al.
Published: (2024)
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
by: Heidari, Sina, et al.
Published: (2026)
by: Heidari, Sina, et al.
Published: (2026)
Hamster: A Fast Synchronous Byzantine Fault Tolerance Protocol
by: Fu, Ximing, et al.
Published: (2024)
by: Fu, Ximing, et al.
Published: (2024)
Similar Items
-
PartIR: Composing SPMD Partitioning Strategies for Machine Learning
by: Alabed, Sami, et al.
Published: (2024) -
HPC resources for CMS offline computing: An integration and scalability challenge for the Submission Infrastructure
by: Yzquierdo, Antonio Perez-Calero, et al.
Published: (2024) -
Flotilla: A scalable, modular and resilient federated learning framework for heterogeneous resources
by: Banerjee, Roopkatha, et al.
Published: (2025) -
Trustworthy Scheduling for Big Data Applications
by: Tomaras, Dimitrios, et al.
Published: (2026) -
TIMBER: On supporting data pipelines in Mobile Cloud Environments
by: Tomaras, Dimitrios, et al.
Published: (2024)