Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Abillama, Pierre, Lee, Changwoo, Dong, Juechu, Blaauw, David, Sylvester, Dennis, Kim, Hun-Seok |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Differentiable Learning of Generalized Structured Matrices for Efficient Deep Neural Networks
by: Lee, Changwoo, et al.
Published: (2023)
by: Lee, Changwoo, et al.
Published: (2023)
BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network Inference
by: Lee, Changwoo, et al.
Published: (2024)
by: Lee, Changwoo, et al.
Published: (2024)
Quantum Circuit Simulation with Fast Tensor Decision Diagram
by: Zhang, Qirui, et al.
Published: (2024)
by: Zhang, Qirui, et al.
Published: (2024)
MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
by: Yaras, Can, et al.
Published: (2025)
by: Yaras, Can, et al.
Published: (2025)
TCP-SSM: Efficient Vision State Space Models with Token-Conditioned Poles
by: Shoouri, Sara, et al.
Published: (2026)
by: Shoouri, Sara, et al.
Published: (2026)
Toleo: Scaling Freshness to Tera-scale Memory using CXL and PIM
by: Dong, Juechu, et al.
Published: (2024)
by: Dong, Juechu, et al.
Published: (2024)
Memory Allocation in Resource-Constrained Reinforcement Learning
by: Tamborski, Massimiliano, et al.
Published: (2025)
by: Tamborski, Massimiliano, et al.
Published: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
by: Kim, Hyunwoo, et al.
Published: (2025)
by: Kim, Hyunwoo, et al.
Published: (2025)
Less is More: Resource-Efficient Low-Rank Adaptation
by: Tian, Chunlin, et al.
Published: (2025)
by: Tian, Chunlin, et al.
Published: (2025)
Batched Low-Rank Adaptation of Foundation Models
by: Wen, Yeming, et al.
Published: (2023)
by: Wen, Yeming, et al.
Published: (2023)
Logarithmic Memory Networks (LMNs): Efficient Long-Range Sequence Modeling for Resource-Constrained Environments
by: Taha, Mohamed A.
Published: (2025)
by: Taha, Mohamed A.
Published: (2025)
Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
by: Shi, Jiang-Xin, et al.
Published: (2025)
by: Shi, Jiang-Xin, et al.
Published: (2025)
SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
by: Fan, Zichen, et al.
Published: (2025)
by: Fan, Zichen, et al.
Published: (2025)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
by: Metere, Alfredo
Published: (2025)
by: Metere, Alfredo
Published: (2025)
Accelerating AI Performance using Anderson Extrapolation on GPUs
by: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Published: (2024)
by: Dajani, Saleem Abdul Fattah Ahmed Al, et al.
Published: (2024)
BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation
by: Liu, Zelin, et al.
Published: (2025)
by: Liu, Zelin, et al.
Published: (2025)
Low-Rank Adaptation of Time Series Foundational Models for Out-of-Domain Modality Forecasting
by: Gupta, Divij, et al.
Published: (2024)
by: Gupta, Divij, et al.
Published: (2024)
Low-Rank Adaptation for Foundation Models: A Comprehensive Review
by: Yang, Menglin, et al.
Published: (2024)
by: Yang, Menglin, et al.
Published: (2024)
MeKi: Memory-based Expert Knowledge Injection for Efficient LLM Scaling
by: Ding, Ning, et al.
Published: (2026)
by: Ding, Ning, et al.
Published: (2026)
SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
by: Lu, Yu-Chen, et al.
Published: (2025)
by: Lu, Yu-Chen, et al.
Published: (2025)
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
by: Yu, Seungmin, et al.
Published: (2024)
by: Yu, Seungmin, et al.
Published: (2024)
Easy Adaptation: An Efficient Task-Specific Knowledge Injection Method for Large Models in Resource-Constrained Environments
by: Chen, Dong, et al.
Published: (2025)
by: Chen, Dong, et al.
Published: (2025)
A Stronger Mixture of Low-Rank Experts for Fine-Tuning Foundation Models
by: Sun, Mengyang, et al.
Published: (2025)
by: Sun, Mengyang, et al.
Published: (2025)
On Accelerating Edge AI: Optimizing Resource-Constrained Environments
by: Sander, Jacob, et al.
Published: (2025)
by: Sander, Jacob, et al.
Published: (2025)
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
by: Zhang, Jiyuan, et al.
Published: (2026)
by: Zhang, Jiyuan, et al.
Published: (2026)
GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
by: Jung, Yeonjoon, et al.
Published: (2025)
by: Jung, Yeonjoon, et al.
Published: (2025)
Prompt Tuning Strikes Back: Customizing Foundation Models with Low-Rank Prompt Adaptation
by: Jain, Abhinav, et al.
Published: (2024)
by: Jain, Abhinav, et al.
Published: (2024)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
by: Shakerdargah, Mohammadali, et al.
Published: (2024)
by: Shakerdargah, Mohammadali, et al.
Published: (2024)
Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
by: Liu, Yi
Published: (2026)
by: Liu, Yi
Published: (2026)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
by: Lee, Seoungsub, et al.
Published: (2026)
by: Lee, Seoungsub, et al.
Published: (2026)
ColPackAgent: Agent-Skill-Guided Hard-Particle Monte Carlo Workflows for Colloidal Packing
by: Ding, Lijie, et al.
Published: (2026)
by: Ding, Lijie, et al.
Published: (2026)
FOAM: Blocked State Folding for Memory-Efficient LLM Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
by: Hajimolahoseini, Habib, et al.
Published: (2023)
by: Hajimolahoseini, Habib, et al.
Published: (2023)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
Breaking the Blocks: Continuous Low-Rank Decomposed Scaling for Unified LLM Quantization and Adaptation
by: Tang, Pingzhi, et al.
Published: (2026)
by: Tang, Pingzhi, et al.
Published: (2026)
SasAgent: Multi-Agent AI System for Small-Angle Scattering Data Analysis
by: Ding, Lijie, et al.
Published: (2025)
by: Ding, Lijie, et al.
Published: (2025)
VEPO: Variable Entropy Policy Optimization for Low-Resource Language Foundation Models
by: Liu, Chonghan, et al.
Published: (2026)
by: Liu, Chonghan, et al.
Published: (2026)
CoRAST: Towards Foundation Model-Powered Correlated Data Analysis in Resource-Constrained CPS and IoT
by: Hu, Yi, et al.
Published: (2024)
by: Hu, Yi, et al.
Published: (2024)
Accurate and Efficient Low-Rank Model Merging in Core Space
by: Panariello, Aniello, et al.
Published: (2025)
by: Panariello, Aniello, et al.
Published: (2025)
Similar Items
-
Differentiable Learning of Generalized Structured Matrices for Efficient Deep Neural Networks
by: Lee, Changwoo, et al.
Published: (2023) -
BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network Inference
by: Lee, Changwoo, et al.
Published: (2024) -
Quantum Circuit Simulation with Fast Tensor Decision Diagram
by: Zhang, Qirui, et al.
Published: (2024) -
MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
by: Yaras, Can, et al.
Published: (2025) -
TCP-SSM: Efficient Vision State Space Models with Token-Conditioned Poles
by: Shoouri, Sara, et al.
Published: (2026)