Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Epstein, Elliot L., Dwaraknath, Rajat Vadiraj, Winnicki, John |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SD-KDE: Score-Debiased Kernel Density Estimation
von: Epstein, Elliot L., et al.
Veröffentlicht: (2025)
von: Epstein, Elliot L., et al.
Veröffentlicht: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
Sampling on Metric Graphs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2025)
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2025)
MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization
von: Wang, Zongwu, et al.
Veröffentlicht: (2025)
von: Wang, Zongwu, et al.
Veröffentlicht: (2025)
PANDORA: A Parallel Dendrogram Construction Algorithm for Single Linkage Clustering on GPU
von: Sao, Piyush, et al.
Veröffentlicht: (2024)
von: Sao, Piyush, et al.
Veröffentlicht: (2024)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
von: Shi, Jinliang, et al.
Veröffentlicht: (2024)
von: Shi, Jinliang, et al.
Veröffentlicht: (2024)
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)
A Selective Homomorphic Encryption Approach for Faster Privacy-Preserving Federated Learning
von: Korkmaz, Abdulkadir, et al.
Veröffentlicht: (2025)
von: Korkmaz, Abdulkadir, et al.
Veröffentlicht: (2025)
Self-Stabilizing Weakly Byzantine Perpetual Gathering of Mobile Agents
von: Hirose, Jion, et al.
Veröffentlicht: (2025)
von: Hirose, Jion, et al.
Veröffentlicht: (2025)
Why does Prediction Accuracy Decrease over Time? Uncertain Positive Learning for Cloud Failure Prediction
von: Li, Haozhe, et al.
Veröffentlicht: (2024)
von: Li, Haozhe, et al.
Veröffentlicht: (2024)
HybridFlow: A Flexible and Efficient RLHF Framework
von: Sheng, Guangming, et al.
Veröffentlicht: (2024)
von: Sheng, Guangming, et al.
Veröffentlicht: (2024)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
Aergia: Leveraging Heterogeneity in Federated Learning Systems
von: Cox, Bart, et al.
Veröffentlicht: (2022)
von: Cox, Bart, et al.
Veröffentlicht: (2022)
Towards Optimal Heterogeneous Client Sampling in Multi-Model Federated Learning
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Haoran, et al.
Veröffentlicht: (2025)
Parameterizing Federated Continual Learning for Reproducible Research
von: Cox, Bart, et al.
Veröffentlicht: (2024)
von: Cox, Bart, et al.
Veröffentlicht: (2024)
Asynchronous Byzantine Federated Learning
von: Cox, Bart, et al.
Veröffentlicht: (2024)
von: Cox, Bart, et al.
Veröffentlicht: (2024)
Hyper-parameter Optimization for Federated Learning with Step-wise Adaptive Mechanism
von: Saadati, Yasaman, et al.
Veröffentlicht: (2024)
von: Saadati, Yasaman, et al.
Veröffentlicht: (2024)
Training Diffusion Models with Federated Learning
von: de Goede, Matthijs, et al.
Veröffentlicht: (2024)
von: de Goede, Matthijs, et al.
Veröffentlicht: (2024)
Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
von: Xin, Jihao, et al.
Veröffentlicht: (2023)
von: Xin, Jihao, et al.
Veröffentlicht: (2023)
Asynchronous Multi-Server Federated Learning for Geo-Distributed Clients
von: Zuo, Yuncong, et al.
Veröffentlicht: (2024)
von: Zuo, Yuncong, et al.
Veröffentlicht: (2024)
CloudSim 7G: An Integrated Toolkit for Modeling and Simulation of Future Generation Cloud Computing Environments
von: Andreoli, Remo, et al.
Veröffentlicht: (2024)
von: Andreoli, Remo, et al.
Veröffentlicht: (2024)
Is Flash Attention Stable?
von: Golden, Alicia, et al.
Veröffentlicht: (2024)
von: Golden, Alicia, et al.
Veröffentlicht: (2024)
Fused3S: Fast Sparse Attention on Tensor Cores
von: Li, Zitong, et al.
Veröffentlicht: (2025)
von: Li, Zitong, et al.
Veröffentlicht: (2025)
ADF-LoRA: Alternating Low-Rank Aggregation for Decentralized Federated Fine-Tuning
von: Wang, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoyu, et al.
Veröffentlicht: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
SPARK: Igniting Communication-Efficient Decentralized Learning via Stage-wise Projected NTK and Accelerated Regularization
von: Xia, Li
Veröffentlicht: (2025)
von: Xia, Li
Veröffentlicht: (2025)
On the Effectiveness of the 'Follow-the-Sun' Strategy in Mitigating the Carbon Footprint of AI in Cloud Instances
von: Vergallo, Roberto, et al.
Veröffentlicht: (2025)
von: Vergallo, Roberto, et al.
Veröffentlicht: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
Scalable Machine Learning Training Infrastructure for Online Ads Recommendation and Auction Scoring Modeling at Google
von: Kurian, George, et al.
Veröffentlicht: (2025)
von: Kurian, George, et al.
Veröffentlicht: (2025)
LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval
von: Epstein, Elliot L., et al.
Veröffentlicht: (2025)
von: Epstein, Elliot L., et al.
Veröffentlicht: (2025)
A Full Compression Pipeline for Green Federated Learning in Communication-Constrained Environments
von: Colybes, Elouan, et al.
Veröffentlicht: (2026)
von: Colybes, Elouan, et al.
Veröffentlicht: (2026)
Tempo: Compiled Dynamic Deep Learning with Symbolic Dependence Graphs
von: Silvestre, Pedro F., et al.
Veröffentlicht: (2025)
von: Silvestre, Pedro F., et al.
Veröffentlicht: (2025)
MaxK-GNN: Extremely Fast GPU Kernel Design for Accelerating Graph Neural Networks Training
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
von: Peng, Hongwu, et al.
Veröffentlicht: (2023)
Allocate Marginal Reviews to Borderline Papers Using LLM Comparative Ranking
von: Epstein, Elliot L., et al.
Veröffentlicht: (2026)
von: Epstein, Elliot L., et al.
Veröffentlicht: (2026)
The PetShop Dataset -- Finding Causes of Performance Issues across Microservices
von: Hardt, Michaela, et al.
Veröffentlicht: (2023)
von: Hardt, Michaela, et al.
Veröffentlicht: (2023)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
von: Shu, Zhihao, et al.
Veröffentlicht: (2026)
von: Shu, Zhihao, et al.
Veröffentlicht: (2026)
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
von: Sun, He, et al.
Veröffentlicht: (2026)
von: Sun, He, et al.
Veröffentlicht: (2026)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
Scaling Point-based Differentiable Rendering for Large-scale Reconstruction
von: Zhao, Hexu, et al.
Veröffentlicht: (2025)
von: Zhao, Hexu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SD-KDE: Score-Debiased Kernel Density Estimation
von: Epstein, Elliot L., et al.
Veröffentlicht: (2025) -
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026) -
Sampling on Metric Graphs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2025) -
MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization
von: Wang, Zongwu, et al.
Veröffentlicht: (2025) -
PANDORA: A Parallel Dendrogram Construction Algorithm for Single Linkage Clustering on GPU
von: Sao, Piyush, et al.
Veröffentlicht: (2024)