cuVegas: Accelerate Multidimensional Monte Carlo Integration through a Parallelized CUDA-based Implementation of the VEGAS Enhanced Algorithm
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tolotti, Emiliano, Jnini, Anas, Vella, Flavio, Passerone, Roberto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
cuConv: A CUDA Implementation of Convolution for CNN Inference
von: Jordà, Marc, et al.
Veröffentlicht: (2021)
von: Jordà, Marc, et al.
Veröffentlicht: (2021)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
von: Song, Boyang
Veröffentlicht: (2025)
von: Song, Boyang
Veröffentlicht: (2025)
Acceleration of Parallel Tempering for Markov Chain Monte Carlo methods
von: Ramos, Aingeru, et al.
Veröffentlicht: (2025)
von: Ramos, Aingeru, et al.
Veröffentlicht: (2025)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
von: Bellavita, Julian, et al.
Veröffentlicht: (2025)
von: Bellavita, Julian, et al.
Veröffentlicht: (2025)
Parallel Gaussian process with kernel approximation in CUDA
von: Carminati, Davide
Veröffentlicht: (2024)
von: Carminati, Davide
Veröffentlicht: (2024)
Taking Cryptography Out of the Data Path via Near-Memory Processing in DRAM
von: Barcarolo, Nicola, et al.
Veröffentlicht: (2026)
von: Barcarolo, Nicola, et al.
Veröffentlicht: (2026)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
von: Zwaka, Linus
Veröffentlicht: (2024)
von: Zwaka, Linus
Veröffentlicht: (2024)
Parallel Paradigms in Modern HPC: A Comparative Analysis of MPI, OpenMP, and CUDA
von: ALHafez, Nizar, et al.
Veröffentlicht: (2025)
von: ALHafez, Nizar, et al.
Veröffentlicht: (2025)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
von: Sojoodi, Amirhossein, et al.
Veröffentlicht: (2026)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
State of practice: evaluating GPU performance of state vector and tensor network methods
von: Vallero, Marzio, et al.
Veröffentlicht: (2024)
von: Vallero, Marzio, et al.
Veröffentlicht: (2024)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
von: Apanasevich, L., et al.
Veröffentlicht: (2024)
High Performance Unstructured SpMM Computation Using Tensor Cores
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
von: Ekelund, Jonah, et al.
Veröffentlicht: (2025)
cuNRTO: GPU-Accelerated Nonlinear Robust Trajectory Optimization
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
Warp-STAR: High-performance, Differentiable GPU-Accelerated Static Timing Analysis through Warp-oriented Parallel Orchestration
von: Huang, En-Ming, et al.
Veröffentlicht: (2026)
von: Huang, En-Ming, et al.
Veröffentlicht: (2026)
Efficient Parallel Implementation of the Pilot Assignment Problem in Massive MIMO Systems
von: Alqudah, Eman, et al.
Veröffentlicht: (2025)
von: Alqudah, Eman, et al.
Veröffentlicht: (2025)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
von: Wang, Zhigang, et al.
Veröffentlicht: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
Lessons Learned Migrating CUDA to SYCL: A HEP Case Study with ROOT RDataFrame
von: Chen, Jolly, et al.
Veröffentlicht: (2024)
von: Chen, Jolly, et al.
Veröffentlicht: (2024)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024)
von: Tang, Ding, et al.
Veröffentlicht: (2024)
cuSZ-$i$: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
von: Liu, Jinyang, et al.
Veröffentlicht: (2023)
von: Liu, Jinyang, et al.
Veröffentlicht: (2023)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
von: He, Jinghai, et al.
Veröffentlicht: (2024)
von: He, Jinghai, et al.
Veröffentlicht: (2024)
A Practical GPU-Accelerated Implementation of Orthogonal Matching Pursuit
von: Lubonja, Ariel, et al.
Veröffentlicht: (2024)
von: Lubonja, Ariel, et al.
Veröffentlicht: (2024)
Multi-GPU Acceleration of PALABOS Fluid Solver using C++ Standard Parallelism
von: Latt, Jonas, et al.
Veröffentlicht: (2025)
von: Latt, Jonas, et al.
Veröffentlicht: (2025)
Accelerating Microswimmer Simulations via a Heterogeneous Pipelined Parallel-in-Time Framework
von: Huang, Ruixiang, et al.
Veröffentlicht: (2026)
von: Huang, Ruixiang, et al.
Veröffentlicht: (2026)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
von: Xia, Mengchun, et al.
Veröffentlicht: (2026)
von: Xia, Mengchun, et al.
Veröffentlicht: (2026)
Communication-Avoiding SpGEMM via Trident Partitioning on Hierarchical GPU Interconnects
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control
von: Buo, Antoonio, et al.
Veröffentlicht: (2026)
von: Buo, Antoonio, et al.
Veröffentlicht: (2026)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
von: Tramm, John, et al.
Veröffentlicht: (2024)
von: Tramm, John, et al.
Veröffentlicht: (2024)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
UniPar: A Unified LLM-Based Framework for Parallel and Accelerated Code Translation in HPC
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
von: Jiang, Xuan, et al.
Veröffentlicht: (2024)
GPU Acceleration of Monte Carlo Tallies on Unstructured Meshes in OpenMC with PUMI-Tally
von: Hasan, Fuad, et al.
Veröffentlicht: (2025)
von: Hasan, Fuad, et al.
Veröffentlicht: (2025)
Comparative Analysis of Distributed Caching Algorithms: Performance Metrics and Implementation Considerations
von: Mayer, Helen, et al.
Veröffentlicht: (2025)
von: Mayer, Helen, et al.
Veröffentlicht: (2025)
pdGRASS: A Fast Parallel Density-Aware Algorithm for Graph Spectral Sparsification
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2025)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2025)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
Optimizing Resource Allocation and Energy Efficiency in Federated Fog Computing for IoT
von: Shah, Syed Sarmad, et al.
Veröffentlicht: (2025)
von: Shah, Syed Sarmad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
cuConv: A CUDA Implementation of Convolution for CNN Inference
von: Jordà, Marc, et al.
Veröffentlicht: (2021) -
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
von: Song, Boyang
Veröffentlicht: (2025) -
Acceleration of Parallel Tempering for Markov Chain Monte Carlo methods
von: Ramos, Aingeru, et al.
Veröffentlicht: (2025) -
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
von: Bellavita, Julian, et al.
Veröffentlicht: (2025) -
Parallel Gaussian process with kernel approximation in CUDA
von: Carminati, Davide
Veröffentlicht: (2024)