Guardado en:
| Autores principales: | Lendve, Shardul, Bletsas, Konstantinos, Souto, Pedro F. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2410.17563 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multiprocessor Scheduling with Memory Constraints: Fundamental Properties and Finding Optimal Solutions
por: Papp, Pál András, et al.
Publicado: (2025)
por: Papp, Pál András, et al.
Publicado: (2025)
Efficient Parallel Scheduling for Sparse Triangular Solvers
por: Böhnlein, Toni, et al.
Publicado: (2025)
por: Böhnlein, Toni, et al.
Publicado: (2025)
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
por: Papp, Pál András, et al.
Publicado: (2024)
por: Papp, Pál András, et al.
Publicado: (2024)
Flex-MIG: Enabling Distributed Execution on MIG
por: Kim, Myeongsu, et al.
Publicado: (2025)
por: Kim, Myeongsu, et al.
Publicado: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
por: Shakeri, Heman, et al.
Publicado: (2026)
por: Shakeri, Heman, et al.
Publicado: (2026)
ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates
por: Lan, Tingfeng, et al.
Publicado: (2025)
por: Lan, Tingfeng, et al.
Publicado: (2025)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
por: Goldman, Amos, et al.
Publicado: (2026)
por: Goldman, Amos, et al.
Publicado: (2026)
Replication in Graph Partitioning and Scheduling Problems
por: Papp, Pál András, et al.
Publicado: (2026)
por: Papp, Pál András, et al.
Publicado: (2026)
GridPilot: Real-Time Grid-Responsive Control for AI Supercomputers
por: Constantinescu, Denisa-Andreea, et al.
Publicado: (2026)
por: Constantinescu, Denisa-Andreea, et al.
Publicado: (2026)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
CRDT-Based Game State Synchronization in Peer-to-Peer VR
por: Dantas, Abel, et al.
Publicado: (2025)
por: Dantas, Abel, et al.
Publicado: (2025)
Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
por: Luiz, Anderson de Lima, et al.
Publicado: (2025)
por: Luiz, Anderson de Lima, et al.
Publicado: (2025)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
por: Cheng, Long, et al.
Publicado: (2026)
por: Cheng, Long, et al.
Publicado: (2026)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
por: Cicirello, Vincent A.
Publicado: (2017)
por: Cicirello, Vincent A.
Publicado: (2017)
Parallel Self-Avoiding Walks for a Low-Autocorrelation Binary Sequences Problem
por: Bošković, Borko, et al.
Publicado: (2022)
por: Bošković, Borko, et al.
Publicado: (2022)
Evaluating Large Language Models for Workload Mapping and Scheduling in Heterogeneous HPC Systems
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
GPU-Initiated Networking for NCCL
por: Hamidouche, Khaled, et al.
Publicado: (2025)
por: Hamidouche, Khaled, et al.
Publicado: (2025)
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
por: Thomaz, Gabriel Fernandes, et al.
Publicado: (2026)
por: Thomaz, Gabriel Fernandes, et al.
Publicado: (2026)
Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
por: Gao, Yunquan, et al.
Publicado: (2025)
por: Gao, Yunquan, et al.
Publicado: (2025)
An Empirical Evaluation of Quantum-Inspired QUBO Methods for Heterogeneous HPC Workflow Mapping and Scheduling
por: Sharma, Aasish Kumar, et al.
Publicado: (2026)
por: Sharma, Aasish Kumar, et al.
Publicado: (2026)
push0: Scalable and Fault-Tolerant Orchestration for Zero-Knowledge Proof Generation
por: Ahmadvand, Mohsen, et al.
Publicado: (2026)
por: Ahmadvand, Mohsen, et al.
Publicado: (2026)
Serverless Cold Starts and Where to Find Them
por: Joosen, Artjom, et al.
Publicado: (2024)
por: Joosen, Artjom, et al.
Publicado: (2024)
Studying the Effect of Schedule Preemption on Dynamic Task Graph Scheduling
por: Khodabandehlou, Mohammadali, et al.
Publicado: (2026)
por: Khodabandehlou, Mohammadali, et al.
Publicado: (2026)
Leveraging Multi-Instance GPUs through moldable task scheduling
por: Villarrubia, Jorge, et al.
Publicado: (2025)
por: Villarrubia, Jorge, et al.
Publicado: (2025)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
por: Chillarón, Mónica, et al.
Publicado: (2024)
por: Chillarón, Mónica, et al.
Publicado: (2024)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
por: Li, Zonghang, et al.
Publicado: (2024)
por: Li, Zonghang, et al.
Publicado: (2024)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
por: Konopa, Michal, et al.
Publicado: (2025)
por: Konopa, Michal, et al.
Publicado: (2025)
Scheduler-Driven Job Atomization
por: Konopa, Michal, et al.
Publicado: (2025)
por: Konopa, Michal, et al.
Publicado: (2025)
Distributed Generalized Linear Models: A Privacy-Preserving Approach
por: Tinoco, Daniel, et al.
Publicado: (2025)
por: Tinoco, Daniel, et al.
Publicado: (2025)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
por: Topcu, Burak, et al.
Publicado: (2026)
por: Topcu, Burak, et al.
Publicado: (2026)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
por: Li, Xiangchen, et al.
Publicado: (2026)
por: Li, Xiangchen, et al.
Publicado: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
por: Li, Xiangchen, et al.
Publicado: (2026)
por: Li, Xiangchen, et al.
Publicado: (2026)
Decentralized Optimization in Time-Varying Networks with Arbitrary Delays
por: Ortega, Tomas, et al.
Publicado: (2024)
por: Ortega, Tomas, et al.
Publicado: (2024)
Laminar: A Probe-First Scheduling Paradigm with Deterministic Runtime Survival
por: Chu, Zhengyan
Publicado: (2026)
por: Chu, Zhengyan
Publicado: (2026)
Rank-Aware Resource Scheduling for Tightly-Coupled MPI Workloads on Kubernetes
por: Xie, Tianfang
Publicado: (2026)
por: Xie, Tianfang
Publicado: (2026)
Light Cone Consistency: Toward a Unified Theory of Consistency in Message-Passing Systems
por: Landers, Rob, et al.
Publicado: (2026)
por: Landers, Rob, et al.
Publicado: (2026)
Instance Configuration for Sustainable Job Shop Scheduling
por: Perez, Christian, et al.
Publicado: (2024)
por: Perez, Christian, et al.
Publicado: (2024)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
por: Li, Yinrong, et al.
Publicado: (2026)
por: Li, Yinrong, et al.
Publicado: (2026)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
por: Siddiqui, Mohammed Humaid, et al.
Publicado: (2025)
por: Siddiqui, Mohammed Humaid, et al.
Publicado: (2025)
Ejemplares similares
-
Multiprocessor Scheduling with Memory Constraints: Fundamental Properties and Finding Optimal Solutions
por: Papp, Pál András, et al.
Publicado: (2025) -
Efficient Parallel Scheduling for Sparse Triangular Solvers
por: Böhnlein, Toni, et al.
Publicado: (2025) -
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
por: Papp, Pál András, et al.
Publicado: (2024) -
Flex-MIG: Enabling Distributed Execution on MIG
por: Kim, Myeongsu, et al.
Publicado: (2025) -
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
por: Shakeri, Heman, et al.
Publicado: (2026)