Scalability of 3D-DFT by block tensor-matrix multiplication on the JUWELS Cluster
Fuente:
arXiv
Guardado en:
| Autores principales: | Malapally, Nitin, Bolnykh, Viacheslav, Suarez, Estela, Carloni, Paolo, Lippert, Thomas, Mandelli, Davide |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A sparsity-aware distributed-memory algorithm for sparse-sparse matrix multiplication
por: Hong, Yuxi, et al.
Publicado: (2024)
por: Hong, Yuxi, et al.
Publicado: (2024)
Stardust: A Scalable and Extensible Simulator for the 3D Continuum
por: Pusztai, Thomas, et al.
Publicado: (2025)
por: Pusztai, Thomas, et al.
Publicado: (2025)
Byzantine Fault-Tolerant Min-Max Optimization
por: Liu, Shuo, et al.
Publicado: (2022)
por: Liu, Shuo, et al.
Publicado: (2022)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
por: Bai, Haoyu, et al.
Publicado: (2024)
por: Bai, Haoyu, et al.
Publicado: (2024)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
por: Kreuzer, Anke, et al.
Publicado: (2019)
por: Kreuzer, Anke, et al.
Publicado: (2019)
Byzantine Consensus in Directed Graphs with Message Authentication
por: Vaidya, Nitin H., et al.
Publicado: (2026)
por: Vaidya, Nitin H., et al.
Publicado: (2026)
Byzantine fault-tolerant distributed set intersection with redundancy
por: Liu, Shuo, et al.
Publicado: (2024)
por: Liu, Shuo, et al.
Publicado: (2024)
Extreme scaling of the metadynamics of paths algorithm on the pre-exascale JUWELS Booster supercomputer
por: Malapally, Nitin, et al.
Publicado: (2025)
por: Malapally, Nitin, et al.
Publicado: (2025)
Asynchronous Checkpoint for Eventually Consistent Databases
por: Ravishankar, Raaghav, et al.
Publicado: (2025)
por: Ravishankar, Raaghav, et al.
Publicado: (2025)
Approximate Byzantine Fault-Tolerance in Distributed Optimization
por: Liu, Shuo, et al.
Publicado: (2021)
por: Liu, Shuo, et al.
Publicado: (2021)
Scalable Readability Evaluation for Graph Layouts: 2D Geometric Distributed Algorithms
por: Yun, Sanggeon
Publicado: (2024)
por: Yun, Sanggeon
Publicado: (2024)
Fixing Non-blocking Data Structures for Better Compatibility with Memory Reclamation Schemes
por: Arovi, Md Amit Hasan, et al.
Publicado: (2025)
por: Arovi, Md Amit Hasan, et al.
Publicado: (2025)
Parallel Gaussian process with kernel approximation in CUDA
por: Carminati, Davide
Publicado: (2024)
por: Carminati, Davide
Publicado: (2024)
Lifting to tensors when compiling scientific computing workloads for AI Engines
por: Brown, Nick, et al.
Publicado: (2026)
por: Brown, Nick, et al.
Publicado: (2026)
Scalable and Performant Data Loading
por: Hira, Moto, et al.
Publicado: (2025)
por: Hira, Moto, et al.
Publicado: (2025)
PWDFT-SW: Extending the Limit of Plane-Wave DFT Calculations to 16K Atoms on the New Sunway Supercomputer
por: Jiang, Qingcai, et al.
Publicado: (2024)
por: Jiang, Qingcai, et al.
Publicado: (2024)
Pilotfish: Distributed Execution for Scalable Blockchains
por: Kniep, Quentin, et al.
Publicado: (2024)
por: Kniep, Quentin, et al.
Publicado: (2024)
Scalable Maxflow Processing for Dynamic Graphs
por: Kannappan, Shruthi, et al.
Publicado: (2025)
por: Kannappan, Shruthi, et al.
Publicado: (2025)
Robust and Scalable Renaming with Subquadratic Bits
por: Bai, Sirui, et al.
Publicado: (2025)
por: Bai, Sirui, et al.
Publicado: (2025)
Building State Machine Replication Using Practical Network Synchrony
por: Wan, Yiliang, et al.
Publicado: (2025)
por: Wan, Yiliang, et al.
Publicado: (2025)
ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding
por: Jin, ChiHeng, et al.
Publicado: (2026)
por: Jin, ChiHeng, et al.
Publicado: (2026)
ClusterLess: Deadline-Aware Serverless Workflow Orchestration on Federated Edge Clusters
por: Farahani, Reza, et al.
Publicado: (2026)
por: Farahani, Reza, et al.
Publicado: (2026)
State of practice: evaluating GPU performance of state vector and tensor network methods
por: Vallero, Marzio, et al.
Publicado: (2024)
por: Vallero, Marzio, et al.
Publicado: (2024)
Truncated multiplication and batch software SIMD AVX512 implementation for faster Montgomery multiplications and modular exponentiation
por: Didier, Laurent-Stéphane, et al.
Publicado: (2024)
por: Didier, Laurent-Stéphane, et al.
Publicado: (2024)
Towards a Testbed for Scalable FaaS Platforms
por: Schirmer, Trever, et al.
Publicado: (2025)
por: Schirmer, Trever, et al.
Publicado: (2025)
The National Research Platform: Stretched, Multi-Tenant, Scientific Kubernetes Cluster
por: Weitzel, Derek, et al.
Publicado: (2025)
por: Weitzel, Derek, et al.
Publicado: (2025)
A Scalable Clustered Architecture for Cyber-Physical Systems
por: Cabral, Bernardo
Publicado: (2024)
por: Cabral, Bernardo
Publicado: (2024)
Towards Efficient and Scalable Distributed Vector Search with RDMA
por: Zhi, Xiangyu, et al.
Publicado: (2025)
por: Zhi, Xiangyu, et al.
Publicado: (2025)
Scalable HPC Job Scheduling and Resource Management in SST
por: Abdurahman, Abubeker, et al.
Publicado: (2025)
por: Abdurahman, Abubeker, et al.
Publicado: (2025)
OSGym: Scalable OS Infra for Computer Use Agents
por: Qin, Zengyi, et al.
Publicado: (2025)
por: Qin, Zengyi, et al.
Publicado: (2025)
Arma: Byzantine Fault Tolerant Consensus with Horizontal Scalability
por: Manevich, Yacov, et al.
Publicado: (2024)
por: Manevich, Yacov, et al.
Publicado: (2024)
Towards a Scalable In Situ Fast Fourier Transform
por: Kulkarni, Sudhanshu, et al.
Publicado: (2024)
por: Kulkarni, Sudhanshu, et al.
Publicado: (2024)
A More Scalable Sparse Dynamic Data Exchange
por: Geyko, Andrew, et al.
Publicado: (2023)
por: Geyko, Andrew, et al.
Publicado: (2023)
Optimising Blockchain Scalability for Real-Time IoT Applications
por: Rhidoy, Hasan Mahmud, et al.
Publicado: (2026)
por: Rhidoy, Hasan Mahmud, et al.
Publicado: (2026)
Totoro$^+$: An Adaptive and Scalable Edge Federated Learning System
por: Ching, Cheng-Wei, et al.
Publicado: (2026)
por: Ching, Cheng-Wei, et al.
Publicado: (2026)
Off-Road Autonomy Validation Using Scalable Digital Twin Simulations Within High-Performance Computing Clusters
por: Samak, Tanmay Vilas, et al.
Publicado: (2024)
por: Samak, Tanmay Vilas, et al.
Publicado: (2024)
Predictable LLM Serving on GPU Clusters
por: Darzi, Erfan, et al.
Publicado: (2025)
por: Darzi, Erfan, et al.
Publicado: (2025)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
por: Xu, Minxian, et al.
Publicado: (2026)
por: Xu, Minxian, et al.
Publicado: (2026)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
por: Li, Zixuan, et al.
Publicado: (2026)
por: Li, Zixuan, et al.
Publicado: (2026)
FedOptimus: Optimizing Vertical Federated Learning for Scalability and Efficiency
por: Shrivastava, Nikita, et al.
Publicado: (2025)
por: Shrivastava, Nikita, et al.
Publicado: (2025)
Ejemplares similares
-
A sparsity-aware distributed-memory algorithm for sparse-sparse matrix multiplication
por: Hong, Yuxi, et al.
Publicado: (2024) -
Stardust: A Scalable and Extensible Simulator for the 3D Continuum
por: Pusztai, Thomas, et al.
Publicado: (2025) -
Byzantine Fault-Tolerant Min-Max Optimization
por: Liu, Shuo, et al.
Publicado: (2022) -
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
por: Bai, Haoyu, et al.
Publicado: (2024) -
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
por: Kreuzer, Anke, et al.
Publicado: (2019)