Modeling the Effect of Data Redundancy on Speedup in MLFMA Near-Field Computation
Fuente:
arXiv
Guardado en:
| Autor principal: | Sadeghi, Morteza |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Optimizing Near Field Computation in the MLFMA Algorithm with Data Redundancy and Performance Modeling on a Single GPU
por: Sadeghi, Morteza, et al.
Publicado: (2024)
por: Sadeghi, Morteza, et al.
Publicado: (2024)
Optimal Parallel Scheduling under Concave Speedup Functions
por: Li, Chengzhang, et al.
Publicado: (2025)
por: Li, Chengzhang, et al.
Publicado: (2025)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
por: Zhuang, Chen, et al.
Publicado: (2025)
por: Zhuang, Chen, et al.
Publicado: (2025)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
por: Veeravalli, Bharadwaj
Publicado: (2026)
por: Veeravalli, Bharadwaj
Publicado: (2026)
Kino-PAX: Highly Parallel Kinodynamic Sampling-based Planner
por: Perrault, Nicolas, et al.
Publicado: (2024)
por: Perrault, Nicolas, et al.
Publicado: (2024)
Energy-Aware Computing in the Year 2026
por: Tchakoute, Roblex Nana, et al.
Publicado: (2026)
por: Tchakoute, Roblex Nana, et al.
Publicado: (2026)
Hiku: Pull-Based Scheduling for Serverless Computing
por: Akbari, Saman, et al.
Publicado: (2025)
por: Akbari, Saman, et al.
Publicado: (2025)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
por: Maurya, Avinash, et al.
Publicado: (2026)
por: Maurya, Avinash, et al.
Publicado: (2026)
Optimal Configuration of API Resources in Cloud Native Computing
por: Truyen, Eddy, et al.
Publicado: (2025)
por: Truyen, Eddy, et al.
Publicado: (2025)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
por: Kulkarni, Apurv Deepak, et al.
Publicado: (2025)
por: Kulkarni, Apurv Deepak, et al.
Publicado: (2025)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
por: Vatsavai, Sairam Sri, et al.
Publicado: (2025)
por: Vatsavai, Sairam Sri, et al.
Publicado: (2025)
Optimizations on Graph-Level for Domain Specific Computations in Julia and Application to QED
por: Reinhard, Anton, et al.
Publicado: (2025)
por: Reinhard, Anton, et al.
Publicado: (2025)
Universal Workers: A Vision for Eliminating Cold Starts in Serverless Computing
por: Akbari, Saman, et al.
Publicado: (2025)
por: Akbari, Saman, et al.
Publicado: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
por: McDonald, Jesse, et al.
Publicado: (2024)
por: McDonald, Jesse, et al.
Publicado: (2024)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
por: Ramesh, Risshab Srinivas
Publicado: (2024)
por: Ramesh, Risshab Srinivas
Publicado: (2024)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
por: Mao, Ying, et al.
Publicado: (2020)
por: Mao, Ying, et al.
Publicado: (2020)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
por: Tharwani, Jay, et al.
Publicado: (2024)
por: Tharwani, Jay, et al.
Publicado: (2024)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
por: Qi, S., et al.
Publicado: (2024)
por: Qi, S., et al.
Publicado: (2024)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
por: Andersson, Måns I., et al.
Publicado: (2025)
por: Andersson, Måns I., et al.
Publicado: (2025)
DREAMS: Decentralized Resource Allocation and Service Management across the Compute Continuum Using Service Affinity
por: Dinh-Tuan, Hai, et al.
Publicado: (2025)
por: Dinh-Tuan, Hai, et al.
Publicado: (2025)
The SAP Cloud Infrastructure Dataset: A Reality Check of Scheduling and Placement of VMs in Cloud Computing
por: Uhlig, Arno, et al.
Publicado: (2025)
por: Uhlig, Arno, et al.
Publicado: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
por: Zhang, Yaozheng, et al.
Publicado: (2025)
por: Zhang, Yaozheng, et al.
Publicado: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
por: Meneghin, Massimiliano, et al.
Publicado: (2025)
por: Meneghin, Massimiliano, et al.
Publicado: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
por: Islam, Tanzima Z., et al.
Publicado: (2024)
por: Islam, Tanzima Z., et al.
Publicado: (2024)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
por: Rose, Martin, et al.
Publicado: (2025)
por: Rose, Martin, et al.
Publicado: (2025)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
por: Bogart, Christopher, et al.
Publicado: (2025)
por: Bogart, Christopher, et al.
Publicado: (2025)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
por: Scheinert, Dominik, et al.
Publicado: (2023)
por: Scheinert, Dominik, et al.
Publicado: (2023)
Modeling and Characterizing Service Interference in Dynamic Infrastructures
por: Medel, VÍctor, et al.
Publicado: (2024)
por: Medel, VÍctor, et al.
Publicado: (2024)
Denoising Application Performance Models with Noise-Resilient Priors
por: de Morais, Gustavo, et al.
Publicado: (2025)
por: de Morais, Gustavo, et al.
Publicado: (2025)
Taking GPU Programming Models to Task for Performance Portability
por: Davis, Joshua H., et al.
Publicado: (2024)
por: Davis, Joshua H., et al.
Publicado: (2024)
Ridgeline: A 2D Roofline Model for Distributed Systems
por: Checconi, Fabio, et al.
Publicado: (2022)
por: Checconi, Fabio, et al.
Publicado: (2022)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
por: Chu, Xiaoyu, et al.
Publicado: (2025)
por: Chu, Xiaoyu, et al.
Publicado: (2025)
Taming Cold Starts: Proactive Serverless Scheduling with Model Predictive Control
por: Nguyen, Chanh, et al.
Publicado: (2025)
por: Nguyen, Chanh, et al.
Publicado: (2025)
Learning-Augmented Performance Model for Tensor Product Factorization in High-Order FEM
por: Ren, Xuanzhengbo, et al.
Publicado: (2026)
por: Ren, Xuanzhengbo, et al.
Publicado: (2026)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
por: Rashid, Md Hasanur, et al.
Publicado: (2026)
por: Rashid, Md Hasanur, et al.
Publicado: (2026)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
por: Afzal, Ayesha, et al.
Publicado: (2024)
por: Afzal, Ayesha, et al.
Publicado: (2024)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
por: Zhao, Xuanlei, et al.
Publicado: (2024)
por: Zhao, Xuanlei, et al.
Publicado: (2024)
Ejemplares similares
-
Optimizing Near Field Computation in the MLFMA Algorithm with Data Redundancy and Performance Modeling on a Single GPU
por: Sadeghi, Morteza, et al.
Publicado: (2024) -
Optimal Parallel Scheduling under Concave Speedup Functions
por: Li, Chengzhang, et al.
Publicado: (2025) -
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
por: Zhuang, Chen, et al.
Publicado: (2025) -
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
por: Papavasileiou, Ioannis, et al.
Publicado: (2026) -
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
por: Veeravalli, Bharadwaj
Publicado: (2026)