Flash-KMeans: Fast and Memory-Efficient Exact K-Means
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Shuo, Xi, Haocheng, Zhao, Yilong, Li, Muyang, Fan, Xiaoze, Zhang, Jintao, Cai, Han, Lin, Yujun, Li, Xiuyu, Keutzer, Kurt, Han, Song, Xu, Chenfeng, Stoica, Ion |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pie: Pooling CPU Memory for LLM Inference
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
von: Han, Pengchao, et al.
Veröffentlicht: (2025)
von: Han, Pengchao, et al.
Veröffentlicht: (2025)
MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
von: Liaw, Yong-Cheng, et al.
Veröffentlicht: (2025)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
von: Liu, Xueshen, et al.
Veröffentlicht: (2026)
von: Liu, Xueshen, et al.
Veröffentlicht: (2026)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2025)
AI and Memory Wall
von: Gholami, Amir, et al.
Veröffentlicht: (2024)
von: Gholami, Amir, et al.
Veröffentlicht: (2024)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
von: Xie, Xi, et al.
Veröffentlicht: (2024)
von: Xie, Xi, et al.
Veröffentlicht: (2024)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
von: Xia, Tian, et al.
Veröffentlicht: (2025)
von: Xia, Tian, et al.
Veröffentlicht: (2025)
PUSHtap: PIM-based In-Memory HTAP with Unified Data Storage Format
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
von: Li, Dacheng, et al.
Veröffentlicht: (2023)
von: Li, Dacheng, et al.
Veröffentlicht: (2023)
FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs
von: Zhang, Haijun, et al.
Veröffentlicht: (2025)
von: Zhang, Haijun, et al.
Veröffentlicht: (2025)
GraphFlash: Enabling Fast and Elastic Graph Processing on Serverless Infrastructure
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
The Dawn of Disaggregation and the Coherence Conundrum: A Call for Federated Coherence
von: Hong, Jaewan, et al.
Veröffentlicht: (2025)
von: Hong, Jaewan, et al.
Veröffentlicht: (2025)
Revisiting Cache Freshness for Emerging Real-Time Applications
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
von: Shu, Zhihao, et al.
Veröffentlicht: (2026)
von: Shu, Zhihao, et al.
Veröffentlicht: (2026)
UAT20: Unifying Liquidity Across Rollups
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
von: Cao, Shiyi, et al.
Veröffentlicht: (2024)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
von: Du, Kuntai, et al.
Veröffentlicht: (2025)
von: Du, Kuntai, et al.
Veröffentlicht: (2025)
Is Flash Attention Stable?
von: Golden, Alicia, et al.
Veröffentlicht: (2024)
von: Golden, Alicia, et al.
Veröffentlicht: (2024)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
On Optimizing the Communication of Model Parallelism
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2022)
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2022)
Energy-aware Incremental OTA Update for Flash-based Batteryless IoT Devices
von: Wei, Wei, et al.
Veröffentlicht: (2024)
von: Wei, Wei, et al.
Veröffentlicht: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
von: Xing, Jiarong, et al.
Veröffentlicht: (2025)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
RedunCut: Measurement-Driven Sampling and Accuracy Performance Modeling for Low-Cost Live Video Analytics
von: Sela, Gur-Eyal, et al.
Veröffentlicht: (2025)
von: Sela, Gur-Eyal, et al.
Veröffentlicht: (2025)
Mean field optimal Core Allocation across Malleable jobs
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Pie: Pooling CPU Memory for LLM Inference
von: Xu, Yi, et al.
Veröffentlicht: (2024) -
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
von: Yang, Shuo, et al.
Veröffentlicht: (2025) -
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
von: Zhao, Yilong, et al.
Veröffentlicht: (2026) -
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
von: Sheng, Ying, et al.
Veröffentlicht: (2023) -
Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
von: Han, Pengchao, et al.
Veröffentlicht: (2025)