Efficient GPU implementation of randomized SVD and its applications
Fuente:
arXiv
Guardado en:
| Autores principales: | Struski, Łukasz, Morkisz, Paweł, Spurek, Przemysław, Bernabeu, Samuel Rodriguez, Trzciński, Tomasz |
|---|---|
| Formato: | Preprint |
| Publicado: |
2021
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Bounding Evidence and Estimating Log-Likelihood in VAE
por: Struski, Łukasz, et al.
Publicado: (2022)
por: Struski, Łukasz, et al.
Publicado: (2022)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
por: Shao, Zishan, et al.
Publicado: (2025)
por: Shao, Zishan, et al.
Publicado: (2025)
PixelBrax: Learning Continuous Control from Pixels End-to-End on the GPU
por: McInroe, Trevor, et al.
Publicado: (2025)
por: McInroe, Trevor, et al.
Publicado: (2025)
Stop Marginalizing My Dreams: Model Inversion via Laplace Kernel for Continual Learning
por: Krukowski, Patryk, et al.
Publicado: (2026)
por: Krukowski, Patryk, et al.
Publicado: (2026)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
por: Zhang, Yijia, et al.
Publicado: (2024)
por: Zhang, Yijia, et al.
Publicado: (2024)
Forecasting GPU Performance for Deep Learning Training and Inference
por: Lee, Seonho, et al.
Publicado: (2024)
por: Lee, Seonho, et al.
Publicado: (2024)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
por: Wu, Wenhao, et al.
Publicado: (2026)
por: Wu, Wenhao, et al.
Publicado: (2026)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
por: Yao, Feiyu, et al.
Publicado: (2026)
por: Yao, Feiyu, et al.
Publicado: (2026)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
por: Taneja, Maanas, et al.
Publicado: (2026)
por: Taneja, Maanas, et al.
Publicado: (2026)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
por: Wang, Han, et al.
Publicado: (2026)
por: Wang, Han, et al.
Publicado: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
por: Jaber, Jaber, et al.
Publicado: (2026)
por: Jaber, Jaber, et al.
Publicado: (2026)
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
por: Wang, Liangyu, et al.
Publicado: (2025)
por: Wang, Liangyu, et al.
Publicado: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
por: Ouyang, Anne, et al.
Publicado: (2025)
por: Ouyang, Anne, et al.
Publicado: (2025)
HyperSound: Generating Implicit Neural Representations of Audio Signals with Hypernetworks
por: Szatkowski, Filip, et al.
Publicado: (2022)
por: Szatkowski, Filip, et al.
Publicado: (2022)
Hardware-efficient tractable probabilistic inference for TinyML Neurosymbolic AI applications
por: Leslin, Jelin, et al.
Publicado: (2025)
por: Leslin, Jelin, et al.
Publicado: (2025)
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
por: Wang, Jiaqi, et al.
Publicado: (2026)
por: Wang, Jiaqi, et al.
Publicado: (2026)
Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems
por: Jali, Neharika, et al.
Publicado: (2024)
por: Jali, Neharika, et al.
Publicado: (2024)
ELROND: Exploring and decomposing intrinsic capabilities of diffusion models
por: Skierś, Paweł, et al.
Publicado: (2026)
por: Skierś, Paweł, et al.
Publicado: (2026)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
por: Sunesh, Aman, et al.
Publicado: (2026)
por: Sunesh, Aman, et al.
Publicado: (2026)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
por: Maczan, Jędrzej
Publicado: (2026)
por: Maczan, Jędrzej
Publicado: (2026)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
por: Ziller, Thomas, et al.
Publicado: (2026)
por: Ziller, Thomas, et al.
Publicado: (2026)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
por: Yu, Shan, et al.
Publicado: (2025)
por: Yu, Shan, et al.
Publicado: (2025)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
por: Xue, Leyang, et al.
Publicado: (2024)
por: Xue, Leyang, et al.
Publicado: (2024)
Efficient Graph Knowledge Distillation from GNNs to Kolmogorov--Arnold Networks via Self-Attention Dynamic Sampling
por: Cui, Can, et al.
Publicado: (2025)
por: Cui, Can, et al.
Publicado: (2025)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
por: Wang, Liangyu, et al.
Publicado: (2025)
por: Wang, Liangyu, et al.
Publicado: (2025)
SAfEPaTh: A System-Level Approach for Efficient Power and Thermal Estimation of Convolutional Neural Network Accelerator
por: Chen, Yukai, et al.
Publicado: (2024)
por: Chen, Yukai, et al.
Publicado: (2024)
GPU Cluster Scheduling for Network-Sensitive Deep Learning
por: Sharma, Aakash, et al.
Publicado: (2024)
por: Sharma, Aakash, et al.
Publicado: (2024)
LLMPerf: GPU Performance Modeling meets Large Language Models
por: Nguyen, Khoi N. M., et al.
Publicado: (2025)
por: Nguyen, Khoi N. M., et al.
Publicado: (2025)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
por: Andrews, Martin, et al.
Publicado: (2025)
por: Andrews, Martin, et al.
Publicado: (2025)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
por: Chrapek, Marcin, et al.
Publicado: (2025)
por: Chrapek, Marcin, et al.
Publicado: (2025)
Exchangeability in Neural Network and its Application to Dynamic Pruning
por: Pu, et al.
Publicado: (2025)
por: Pu, et al.
Publicado: (2025)
Parallel Implementations Assessment of a Spatial-Spectral Classifier for Hyperspectral Clinical Applications
por: Lazcano, Raquel, et al.
Publicado: (2024)
por: Lazcano, Raquel, et al.
Publicado: (2024)
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
por: Palaniappan, Kathiravan
Publicado: (2026)
por: Palaniappan, Kathiravan
Publicado: (2026)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
por: Lin, Zhongyi, et al.
Publicado: (2024)
por: Lin, Zhongyi, et al.
Publicado: (2024)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
por: Li, Zhuojin, et al.
Publicado: (2025)
por: Li, Zhuojin, et al.
Publicado: (2025)
Spectral Energy Centroid: a Metric for Improving Performance and Analyzing Spectral Bias in Implicit Neural Representations
por: Dądela, Tomasz, et al.
Publicado: (2026)
por: Dądela, Tomasz, et al.
Publicado: (2026)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
por: Oztop, Beste, et al.
Publicado: (2026)
por: Oztop, Beste, et al.
Publicado: (2026)
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
por: Mozaffari, Mohammad, et al.
Publicado: (2025)
por: Mozaffari, Mohammad, et al.
Publicado: (2025)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
por: Shi, Jiabo, et al.
Publicado: (2025)
por: Shi, Jiabo, et al.
Publicado: (2025)
EcoSpa: Efficient Transformer Training with Coupled Sparsity
por: Xiao, Jinqi, et al.
Publicado: (2025)
por: Xiao, Jinqi, et al.
Publicado: (2025)
Ejemplares similares
-
Bounding Evidence and Estimating Log-Likelihood in VAE
por: Struski, Łukasz, et al.
Publicado: (2022) -
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
por: Shao, Zishan, et al.
Publicado: (2025) -
PixelBrax: Learning Continuous Control from Pixels End-to-End on the GPU
por: McInroe, Trevor, et al.
Publicado: (2025) -
Stop Marginalizing My Dreams: Model Inversion via Laplace Kernel for Continual Learning
por: Krukowski, Patryk, et al.
Publicado: (2026) -
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
por: Zhang, Yijia, et al.
Publicado: (2024)