FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shao, Zishan, Wang, Yixiao, Wang, Qinsi, Jiang, Ting, Du, Zhixu, Ye, Hancheng, Zhuo, Danyang, Chen, Yiran, Li, Hai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
Efficient GPU implementation of randomized SVD and its applications
by: Struski, Łukasz, et al.
Published: (2021)
by: Struski, Łukasz, et al.
Published: (2021)
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
by: Ren, Jie, et al.
Published: (2025)
by: Ren, Jie, et al.
Published: (2025)
FlashFPS: Efficient Farthest Point Sampling for Large-Scale Point Clouds via Pruning and Caching
by: Fu, Yuzhe, et al.
Published: (2026)
by: Fu, Yuzhe, et al.
Published: (2026)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
by: Fu, Zizhuo, et al.
Published: (2025)
by: Fu, Zizhuo, et al.
Published: (2025)
ZEUS: Accelerating Diffusion Models with Only Second-Order Predictor
by: Wang, Yixiao, et al.
Published: (2026)
by: Wang, Yixiao, et al.
Published: (2026)
Updates on the Low-Level Abstraction of Memory Access
by: Gruber, Bernhard Manfred
Published: (2023)
by: Gruber, Bernhard Manfred
Published: (2023)
A Zoned Storage Optimized Flash Cache on ZNS SSDs
by: Yang, Chongzhuo, et al.
Published: (2024)
by: Yang, Chongzhuo, et al.
Published: (2024)
Anatomizing Deep Learning Inference in Web Browsers
by: Wang, Qipeng, et al.
Published: (2024)
by: Wang, Qipeng, et al.
Published: (2024)
An Experimental Study of Low-Latency Video Streaming over 5G
by: Khan, Imran, et al.
Published: (2024)
by: Khan, Imran, et al.
Published: (2024)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
by: Zheng, Wenhao, et al.
Published: (2025)
by: Zheng, Wenhao, et al.
Published: (2025)
Examem: Low-Overhead Memory Instrumentation for Intelligent Memory Systems
by: Poduval, Ashwin, et al.
Published: (2024)
by: Poduval, Ashwin, et al.
Published: (2024)
Scalable Binary CUR Low-Rank Approximation Algorithm
by: Su, Bowen
Published: (2025)
by: Su, Bowen
Published: (2025)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
by: Suo, Jiashun, et al.
Published: (2025)
by: Suo, Jiashun, et al.
Published: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
by: Liu, Songze, et al.
Published: (2025)
by: Liu, Songze, et al.
Published: (2025)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
by: Wang, Liangyu, et al.
Published: (2025)
by: Wang, Liangyu, et al.
Published: (2025)
Efficiently Ranking Software Variants with Minimal Benchmarks
by: Matricon, Théo, et al.
Published: (2025)
by: Matricon, Théo, et al.
Published: (2025)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
by: Metere, Alfredo
Published: (2025)
by: Metere, Alfredo
Published: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
SADA: Stability-guided Adaptive Diffusion Acceleration
by: Jiang, Ting, et al.
Published: (2025)
by: Jiang, Ting, et al.
Published: (2025)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
by: Dong, Ximing, et al.
Published: (2026)
by: Dong, Ximing, et al.
Published: (2026)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025)
by: Fang, Yunhua, et al.
Published: (2025)
CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation
by: Wang, Qinsi, et al.
Published: (2024)
by: Wang, Qinsi, et al.
Published: (2024)
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
Streaming Data in HPC Workflows Using ADIOS
by: Eisenhauer, Greg, et al.
Published: (2024)
by: Eisenhauer, Greg, et al.
Published: (2024)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Iterative Layer Pruning for Efficient Translation Inference
by: Moslem, Yasmin, et al.
Published: (2025)
by: Moslem, Yasmin, et al.
Published: (2025)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
by: Wu, Dan, et al.
Published: (2023)
by: Wu, Dan, et al.
Published: (2023)
Towards Multi-dimensional Elasticity for Pervasive Stream Processing Services
by: Sedlak, Boris, et al.
Published: (2025)
by: Sedlak, Boris, et al.
Published: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
by: Shen, Zixu, et al.
Published: (2025)
by: Shen, Zixu, et al.
Published: (2025)
Unikernels vs. Containers: A Runtime-Level Performance Comparison for Resource-Constrained Edge Workloads
by: Dinh-Tuan, Hai
Published: (2025)
by: Dinh-Tuan, Hai
Published: (2025)
Tuning Fast Memory Size based on Modeling of Page Migration for Tiered Memory
by: Chen, Shangye, et al.
Published: (2024)
by: Chen, Shangye, et al.
Published: (2024)
PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
by: Thapa, Krishu K, et al.
Published: (2025)
by: Thapa, Krishu K, et al.
Published: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
by: An, Zihao, et al.
Published: (2025)
by: An, Zihao, et al.
Published: (2025)
Heterogeneous Memory Pool Tuning
by: Vaverka, Filip, et al.
Published: (2025)
by: Vaverka, Filip, et al.
Published: (2025)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
by: Zhang, Hang, et al.
Published: (2025)
by: Zhang, Hang, et al.
Published: (2025)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023)
by: Gupta, Ahan, et al.
Published: (2023)
Similar Items
-
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026) -
Efficient GPU implementation of randomized SVD and its applications
by: Struski, Łukasz, et al.
Published: (2021) -
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
by: Ren, Jie, et al.
Published: (2025) -
FlashFPS: Efficient Farthest Point Sampling for Large-Scale Point Clouds via Pruning and Caching
by: Fu, Yuzhe, et al.
Published: (2026) -
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
by: Fu, Zizhuo, et al.
Published: (2025)