Optimal Brain Decomposition for Accurate LLM Low-Rank Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuhang, Lee, Donghyun, Yin, Ruokai, Panda, Priyadarshini |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
Memba: Membrane-driven Parameter-Efficient Fine-Tuning for Mamba
by: Lee, Donghyun, et al.
Published: (2025)
by: Lee, Donghyun, et al.
Published: (2025)
GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
by: Li, Yuhang, et al.
Published: (2025)
by: Li, Yuhang, et al.
Published: (2025)
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
TT-SNN: Tensor Train Decomposition for Efficient Spiking Neural Network Training
by: Lee, Donghyun, et al.
Published: (2024)
by: Lee, Donghyun, et al.
Published: (2024)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
ReSpike: Residual Frames-based Hybrid Spiking Neural Networks for Efficient Action Recognition
by: Xiao, Shiting, et al.
Published: (2024)
by: Xiao, Shiting, et al.
Published: (2024)
MD-SNN: Membrane Potential-aware Distillation on Quantized Spiking Neural Network
by: Lee, Donghyun, et al.
Published: (2025)
by: Lee, Donghyun, et al.
Published: (2025)
Structured Variational $D$-Decomposition for Accurate and Stable Low-Rank Approximation
by: Katende, Ronald
Published: (2025)
by: Katende, Ronald
Published: (2025)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
by: Li, Yuhang, et al.
Published: (2026)
by: Li, Yuhang, et al.
Published: (2026)
Do We Really Need a Large Number of Visual Prompts?
by: Kim, Youngeun, et al.
Published: (2023)
by: Kim, Youngeun, et al.
Published: (2023)
GenQ: Quantization in Low Data Regimes with Generative Synthetic Data
by: Li, Yuhang, et al.
Published: (2023)
by: Li, Yuhang, et al.
Published: (2023)
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
by: Cho, Yoonjun, et al.
Published: (2025)
by: Cho, Yoonjun, et al.
Published: (2025)
Optimal Policy Sparsification and Low Rank Decomposition for Deep Reinforcement Learning
by: Goddla, Vikram
Published: (2024)
by: Goddla, Vikram
Published: (2024)
LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks
by: Yin, Ruokai, et al.
Published: (2024)
by: Yin, Ruokai, et al.
Published: (2024)
When In-memory Computing Meets Spiking Neural Networks -- A Perspective on Device-Circuit-System-and-Algorithm Co-design
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
ClipFormer: Key-Value Clipping of Transformers on Memristive Crossbars for Write Noise Mitigation
by: Bhattacharjee, Abhiroop, et al.
Published: (2024)
by: Bhattacharjee, Abhiroop, et al.
Published: (2024)
Efficient Alternating Minimization with Applications to Weighted Low Rank Approximation
by: Song, Zhao, et al.
Published: (2023)
by: Song, Zhao, et al.
Published: (2023)
FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
by: Zaccone, Riccardo, et al.
Published: (2026)
by: Zaccone, Riccardo, et al.
Published: (2026)
Importance-Guided Basis Selection for Low-Rank Decomposition of Large Language Models
by: Asante, Daniel Agyei, et al.
Published: (2026)
by: Asante, Daniel Agyei, et al.
Published: (2026)
Maestro: Uncovering Low-Rank Structures via Trainable Decomposition
by: Horvath, Samuel, et al.
Published: (2023)
by: Horvath, Samuel, et al.
Published: (2023)
Spiking Transformer with Spatial-Temporal Attention
by: Lee, Donghyun, et al.
Published: (2024)
by: Lee, Donghyun, et al.
Published: (2024)
SpikePool: Event-driven Spiking Transformer with Pooling Attention
by: Lee, Donghyun, et al.
Published: (2025)
by: Lee, Donghyun, et al.
Published: (2025)
Low-Rank Curvature for Zeroth-Order Optimization in LLM Fine-Tuning
by: Seung, Hyunseok, et al.
Published: (2025)
by: Seung, Hyunseok, et al.
Published: (2025)
The Rank-Reduced Kalman Filter: Approximate Dynamical-Low-Rank Filtering In High Dimensions
by: Schmidt, Jonathan, et al.
Published: (2023)
by: Schmidt, Jonathan, et al.
Published: (2023)
Low-Rank Tensor Decompositions for the Theory of Neural Networks
by: Borsoi, Ricardo, et al.
Published: (2025)
by: Borsoi, Ricardo, et al.
Published: (2025)
MAP: Revisiting Weight Decomposition for Low-Rank Adaptation
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
LoRA-GA: Low-Rank Adaptation with Gradient Approximation
by: Wang, Shaowen, et al.
Published: (2024)
by: Wang, Shaowen, et al.
Published: (2024)
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
by: Moitra, Abhishek, et al.
Published: (2025)
by: Moitra, Abhishek, et al.
Published: (2025)
Approximately Optimal Core Shapes for Tensor Decompositions
by: Ghadiri, Mehrdad, et al.
Published: (2023)
by: Ghadiri, Mehrdad, et al.
Published: (2023)
Workload-Balanced Pruning for Sparse Spiking Neural Networks
by: Yin, Ruokai, et al.
Published: (2023)
by: Yin, Ruokai, et al.
Published: (2023)
Low-Rank Approximation of Structural Redundancy for Self-Supervised Learning
by: Du, Kang, et al.
Published: (2024)
by: Du, Kang, et al.
Published: (2024)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
by: Lee, Seoungsub, et al.
Published: (2026)
by: Lee, Seoungsub, et al.
Published: (2026)
Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
Structure-Preserving Margin Distribution Learning for High-Order Tensor Data with Low-Rank Decomposition
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Fast Learnings of Coupled Nonnegative Tensor Decomposition Using Optimal Gradient and Low-rank Approximation
by: Wang, Xiulin, et al.
Published: (2023)
by: Wang, Xiulin, et al.
Published: (2023)
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
by: Ghosh, Arkapravo, et al.
Published: (2025)
by: Ghosh, Arkapravo, et al.
Published: (2025)
Stabilizing Native Low-Rank LLM Pretraining
by: Janson, Paul, et al.
Published: (2026)
by: Janson, Paul, et al.
Published: (2026)
Similar Items
-
DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
by: Yin, Ruokai, et al.
Published: (2025) -
Memba: Membrane-driven Parameter-Efficient Fine-Tuning for Mamba
by: Lee, Donghyun, et al.
Published: (2025) -
GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
by: Li, Yuhang, et al.
Published: (2025) -
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
by: Li, Yuhang, et al.
Published: (2024) -
TT-SNN: Tensor Train Decomposition for Efficient Spiking Neural Network Training
by: Lee, Donghyun, et al.
Published: (2024)