On The Performance of Prefix-Sum Parallel Kalman Filters and Smoothers on GPUs
Fuente:
arXiv
Salvato in:
| Autori principali: | Särkkä, Simo, García-Fernández, Ángel F. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Parallel state estimation for systems with integrated measurements
di: Yaghoobi, Fatemeh, et al.
Pubblicazione: (2024)
di: Yaghoobi, Fatemeh, et al.
Pubblicazione: (2024)
Temporal Parallelisation of the HJB Equation and Continuous-Time Linear Quadratic Control
di: Särkkä, Simo, et al.
Pubblicazione: (2022)
di: Särkkä, Simo, et al.
Pubblicazione: (2022)
Auxiliary MCMC and particle Gibbs samplers for parallelisable inference in latent dynamical systems
di: Corenflos, Adrien, et al.
Pubblicazione: (2023)
di: Corenflos, Adrien, et al.
Pubblicazione: (2023)
Parallelizing Maximal Clique Enumeration on GPUs
di: Almasri, Mohammad, et al.
Pubblicazione: (2022)
di: Almasri, Mohammad, et al.
Pubblicazione: (2022)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
di: Lee, Jin, et al.
Pubblicazione: (2026)
di: Lee, Jin, et al.
Pubblicazione: (2026)
Temporal parallelisation of continuous-time maximum-a-posteriori trajectory estimation
di: Razavi, Hassan, et al.
Pubblicazione: (2025)
di: Razavi, Hassan, et al.
Pubblicazione: (2025)
Communication Round and Computation Efficient Exclusive Prefix-Sums Algorithms (for MPI_Exscan)
di: Träff, Jesper Larsson
Pubblicazione: (2025)
di: Träff, Jesper Larsson
Pubblicazione: (2025)
Raptr: Prefix Consensus for Robust High-Performance BFT
di: Tonkikh, Andrei, et al.
Pubblicazione: (2025)
di: Tonkikh, Andrei, et al.
Pubblicazione: (2025)
KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
di: Pan, Zaifeng, et al.
Pubblicazione: (2025)
di: Pan, Zaifeng, et al.
Pubblicazione: (2025)
Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs
di: Xu, Yechen, et al.
Pubblicazione: (2026)
di: Xu, Yechen, et al.
Pubblicazione: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
di: Fan, Jiakun, et al.
Pubblicazione: (2025)
di: Fan, Jiakun, et al.
Pubblicazione: (2025)
Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems
di: Wang, Chen, et al.
Pubblicazione: (2024)
di: Wang, Chen, et al.
Pubblicazione: (2024)
Cuckoo-GPU: Accelerating Cuckoo Filters on Modern GPUs
di: Dortmann, Tim, et al.
Pubblicazione: (2026)
di: Dortmann, Tim, et al.
Pubblicazione: (2026)
GPUArmor: A Hardware-Software Co-design for Efficient and Scalable Memory Safety on GPUs
di: Ziad, Mohamed Tarek Ibn, et al.
Pubblicazione: (2025)
di: Ziad, Mohamed Tarek Ibn, et al.
Pubblicazione: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
di: Ernst, Dominik, et al.
Pubblicazione: (2022)
di: Ernst, Dominik, et al.
Pubblicazione: (2022)
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
di: Costanzo, Manuel, et al.
Pubblicazione: (2024)
di: Costanzo, Manuel, et al.
Pubblicazione: (2024)
DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks
di: Wang, Jinquan, et al.
Pubblicazione: (2025)
di: Wang, Jinquan, et al.
Pubblicazione: (2025)
Prefix Consensus For Censorship Resistant BFT
di: Xiang, Zhuolun, et al.
Pubblicazione: (2026)
di: Xiang, Zhuolun, et al.
Pubblicazione: (2026)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
di: Tramm, John, et al.
Pubblicazione: (2024)
di: Tramm, John, et al.
Pubblicazione: (2024)
Ksurf: Attention Kalman Filter and Principal Component Analysis for Prediction under Highly Variable Cloud Workloads
di: Dang'ana, Michael, et al.
Pubblicazione: (2024)
di: Dang'ana, Michael, et al.
Pubblicazione: (2024)
Faster Parallel Triangular Maximally Filtered Graphs and Hierarchical Clustering
di: Raphael, Steven, et al.
Pubblicazione: (2024)
di: Raphael, Steven, et al.
Pubblicazione: (2024)
BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
di: Li, Zhouzi, et al.
Pubblicazione: (2026)
di: Li, Zhouzi, et al.
Pubblicazione: (2026)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
di: Wu, Shixun, et al.
Pubblicazione: (2024)
di: Wu, Shixun, et al.
Pubblicazione: (2024)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
di: Sun, Mo, et al.
Pubblicazione: (2024)
di: Sun, Mo, et al.
Pubblicazione: (2024)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
di: Li, Pengbo, et al.
Pubblicazione: (2026)
di: Li, Pengbo, et al.
Pubblicazione: (2026)
An Adaptive Distributed Stencil Abstraction for GPUs
di: Bhosale, Aditya, et al.
Pubblicazione: (2025)
di: Bhosale, Aditya, et al.
Pubblicazione: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
di: Hsieh, Chou-Ying, et al.
Pubblicazione: (2024)
Optimizing sDTW for AMD GPUs
di: Latta-Lin, Daniel, et al.
Pubblicazione: (2024)
di: Latta-Lin, Daniel, et al.
Pubblicazione: (2024)
Combining Performance and Productivity: Accelerating the Network Sensing Graph Challenge with GPUs and Commodity Data Science Software
di: Samsi, Siddharth, et al.
Pubblicazione: (2025)
di: Samsi, Siddharth, et al.
Pubblicazione: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
di: Chen, Aodong, et al.
Pubblicazione: (2023)
di: Chen, Aodong, et al.
Pubblicazione: (2023)
Scaling State-Space Models on Multiple GPUs with Tensor Parallelism
di: Dutt, Anurag, et al.
Pubblicazione: (2026)
di: Dutt, Anurag, et al.
Pubblicazione: (2026)
Parallelizing Large-Scale Tensor Network Contraction on Multiple GPUs
di: Pan, Feng, et al.
Pubblicazione: (2026)
di: Pan, Feng, et al.
Pubblicazione: (2026)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
di: Jangda, Abhinav, et al.
Pubblicazione: (2024)
di: Jangda, Abhinav, et al.
Pubblicazione: (2024)
Optimal Workload Placement on Multi-Instance GPUs
di: Turkkan, Bekir, et al.
Pubblicazione: (2024)
di: Turkkan, Bekir, et al.
Pubblicazione: (2024)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
di: Wu, Tianyuan, et al.
Pubblicazione: (2024)
di: Wu, Tianyuan, et al.
Pubblicazione: (2024)
Parallel-in-Time Kalman Smoothing Using Orthogonal Transformations
di: Gargir, Shahaf, et al.
Pubblicazione: (2025)
di: Gargir, Shahaf, et al.
Pubblicazione: (2025)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
di: Zhao, Zhan, et al.
Pubblicazione: (2026)
di: Zhao, Zhan, et al.
Pubblicazione: (2026)
Faster Vertex Cover Algorithms on GPUs with Component-Aware Parallel Branching
di: Amro, Hussein, et al.
Pubblicazione: (2025)
di: Amro, Hussein, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Parallel state estimation for systems with integrated measurements
di: Yaghoobi, Fatemeh, et al.
Pubblicazione: (2024) -
Temporal Parallelisation of the HJB Equation and Continuous-Time Linear Quadratic Control
di: Särkkä, Simo, et al.
Pubblicazione: (2022) -
Auxiliary MCMC and particle Gibbs samplers for parallelisable inference in latent dynamical systems
di: Corenflos, Adrien, et al.
Pubblicazione: (2023) -
Parallelizing Maximal Clique Enumeration on GPUs
di: Almasri, Mohammad, et al.
Pubblicazione: (2022) -
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
di: Lee, Jin, et al.
Pubblicazione: (2026)