Scalable Dual Coordinate Descent for Kernel Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Shao, Zishan, Devarakonda, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
by: Devarakonda, Aditya, et al.
Published: (2025)
by: Devarakonda, Aditya, et al.
Published: (2025)
A Task Parallel Orthonormalization Multigrid Method For Multiphase Elliptic Problems
by: Toprak, Teoman, et al.
Published: (2025)
by: Toprak, Teoman, et al.
Published: (2025)
Canonicalization of Batched Einstein Summations for Tuning Retrieval
by: Kulkarni, Kaushik, et al.
Published: (2026)
by: Kulkarni, Kaushik, et al.
Published: (2026)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Communication-Efficient and Memory-Aware Parallel Bootstrapping using MPI
by: Zhang, Di
Published: (2025)
by: Zhang, Di
Published: (2025)
Sensor Placement for Tsunami Early Warning via Large-Scale Bayesian Optimal Experimental Design
by: Venkat, Sreeram, et al.
Published: (2026)
by: Venkat, Sreeram, et al.
Published: (2026)
Direct Low-Dose CT Image Reconstruction on GPU using Out-Of-Core: Precision and Quality Study
by: Chillarón, M., et al.
Published: (2024)
by: Chillarón, M., et al.
Published: (2024)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
by: Chillarón, Mónica, et al.
Published: (2024)
by: Chillarón, Mónica, et al.
Published: (2024)
Stencil Computations on AMD and Nvidia Graphics Processors: Performance and Tuning Strategies
by: Pekkilä, Johannes, et al.
Published: (2024)
by: Pekkilä, Johannes, et al.
Published: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
by: Suzuki, Kengo, et al.
Published: (2026)
by: Suzuki, Kengo, et al.
Published: (2026)
Adaptive time step selection for Spectral Deferred Correction
by: Saupe, Thomas, et al.
Published: (2024)
by: Saupe, Thomas, et al.
Published: (2024)
Resilience Against Soft Faults through Adaptivity in Spectral Deferred Correction
by: Saupe, Thomas, et al.
Published: (2024)
by: Saupe, Thomas, et al.
Published: (2024)
Fast GPU Linear Algebra via Compile Time Expression Fusion
by: Curtin, Ryan R., et al.
Published: (2026)
by: Curtin, Ryan R., et al.
Published: (2026)
Armadillo: An Efficient Framework for Numerical Linear Algebra
by: Sanderson, Conrad, et al.
Published: (2025)
by: Sanderson, Conrad, et al.
Published: (2025)
A Virtual Processor brings back the Free Lunch
by: Kutschbach, Haymo
Published: (2026)
by: Kutschbach, Haymo
Published: (2026)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
An Evaluation of Massively Parallel Algorithms for DFA Minimization
by: Martens, Jan, et al.
Published: (2024)
by: Martens, Jan, et al.
Published: (2024)
Scalable s-step Preconditioned Conjugate Gradient with Chebyshev Basis and Gauss-Seidel Gram Solve
by: D'Ambra, Pasqua, et al.
Published: (2026)
by: D'Ambra, Pasqua, et al.
Published: (2026)
Topology-Based Reconstruction Prevention for Decentralised Learning
by: Dekker, Florine W., et al.
Published: (2023)
by: Dekker, Florine W., et al.
Published: (2023)
Fast Evaluation of Truncated Neumann Series by Low-Product Radix Kernels
by: Sao, Piyush
Published: (2026)
by: Sao, Piyush
Published: (2026)
Inexact Gauss Seidel and Coarse Solvers for AMG and s-step CG
by: Thomas, Stephen, et al.
Published: (2025)
by: Thomas, Stephen, et al.
Published: (2025)
Parallel Gauss-Jordan Elimination and System Reduction for Efficient Circuit Simulation
by: Noveski, Filip, et al.
Published: (2026)
by: Noveski, Filip, et al.
Published: (2026)
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations
by: Sedukhin, Stanislav, et al.
Published: (2025)
by: Sedukhin, Stanislav, et al.
Published: (2025)
nuGPR: GPU-Accelerated Gaussian Process Regression with Iterative Algorithms and Low-Rank Approximations
by: Zhao, Ziqi, et al.
Published: (2025)
by: Zhao, Ziqi, et al.
Published: (2025)
Shortest paths search method based on the projective description of unweighted mixed graphs
by: Melent'ev, V. A.
Published: (2023)
by: Melent'ev, V. A.
Published: (2023)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022)
by: Li, Chendi, et al.
Published: (2022)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024)
by: Xu, Yao, et al.
Published: (2024)
Parallelization and scalability analysis of inverse factorization using the Chunks and Tasks programming model
by: Artemov, Anton G., et al.
Published: (2019)
by: Artemov, Anton G., et al.
Published: (2019)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
by: Li, Shigang, et al.
Published: (2020)
by: Li, Shigang, et al.
Published: (2020)
Speeding up an unsteady flow simulation by adaptive BDDC and Krylov subspace recycling
by: Hanek, Martin, et al.
Published: (2024)
by: Hanek, Martin, et al.
Published: (2024)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
by: Kanakagiri, Raghavendra, et al.
Published: (2023)
Similar Items
-
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
by: Devarakonda, Aditya, et al.
Published: (2025) -
A Task Parallel Orthonormalization Multigrid Method For Multiphase Elliptic Problems
by: Toprak, Teoman, et al.
Published: (2025) -
Canonicalization of Batched Einstein Summations for Tuning Retrieval
by: Kulkarni, Kaushik, et al.
Published: (2026) -
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025) -
Communication-Efficient and Memory-Aware Parallel Bootstrapping using MPI
by: Zhang, Di
Published: (2025)