NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Dinghong, Xu, Jierui, Yang, Weichu, Su, Pengfei, Li, Dong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HLAT: High-quality Large Language Model Pre-trained on AWS Trainium
by: Fan, Haozheng, et al.
Published: (2024)
by: Fan, Haozheng, et al.
Published: (2024)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
by: Song, Dinghong, et al.
Published: (2025)
by: Song, Dinghong, et al.
Published: (2025)
NinjaLLM: Fast, Scalable and Cost-effective RAG using Amazon SageMaker and AWS Trainium and Inferentia2
by: Xue, Tengfei, et al.
Published: (2024)
by: Xue, Tengfei, et al.
Published: (2024)
ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
by: Yuan, Zhihang, et al.
Published: (2023)
by: Yuan, Zhihang, et al.
Published: (2023)
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Disentangling MLP Neuron Weights in Vocabulary Space
by: Avrahamy, Asaf, et al.
Published: (2026)
by: Avrahamy, Asaf, et al.
Published: (2026)
Adversarial Contrastive Learning for LLM Quantization Attacks
by: Song, Dinghong, et al.
Published: (2026)
by: Song, Dinghong, et al.
Published: (2026)
EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition
by: Nasiri, Hamid, et al.
Published: (2025)
by: Nasiri, Hamid, et al.
Published: (2025)
Distilling Algorithmic Reasoning from LLMs via Explaining Solution Programs
by: Li, Jierui, et al.
Published: (2024)
by: Li, Jierui, et al.
Published: (2024)
SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Identifying and Transferring Reasoning-Critical Neurons: Improving LLM Inference Reliability via Activation Steering
by: Dong, Fangan, et al.
Published: (2026)
by: Dong, Fangan, et al.
Published: (2026)
Dual Decomposition of Weights and Singular Value Low Rank Adaptation
by: Han, Jialong, et al.
Published: (2025)
by: Han, Jialong, et al.
Published: (2025)
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
SORSA: Singular Values and Orthonormal Regularized Singular Vectors Adaptation of Large Language Models
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Efficient and Effective Prompt Tuning via Prompt Decomposition and Compressed Outer Product
by: Lan, Pengxiang, et al.
Published: (2025)
by: Lan, Pengxiang, et al.
Published: (2025)
Understanding How Value Neurons Shape the Generation of Specified Values in LLMs
by: Su, Yi, et al.
Published: (2025)
by: Su, Yi, et al.
Published: (2025)
AdaSVD: Adaptive Singular Value Decomposition for Large Language Models
by: Li, Zhiteng, et al.
Published: (2025)
by: Li, Zhiteng, et al.
Published: (2025)
A Mechanism and Optimization Study on the Impact of Information Density on User-Generated Content Named Entity Recognition
by: Xiaobo, Jiang, et al.
Published: (2026)
by: Xiaobo, Jiang, et al.
Published: (2026)
AlgoSimBench: Identifying Algorithmically Similar Problems for Competitive Programming
by: Li, Jierui, et al.
Published: (2025)
by: Li, Jierui, et al.
Published: (2025)
NeuronMoE: Neuron-Guided Mixture-of-Experts for Efficient Multilingual LLM Extension
by: Li, Rongzhi, et al.
Published: (2026)
by: Li, Rongzhi, et al.
Published: (2026)
Optimal Singular Damage: Efficient LLM Inference in Low Storage Regimes
by: Alipour, Mohammadsajad, et al.
Published: (2025)
by: Alipour, Mohammadsajad, et al.
Published: (2025)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
by: Yang, Dongjie, et al.
Published: (2024)
by: Yang, Dongjie, et al.
Published: (2024)
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
by: Song, Jiwon, et al.
Published: (2025)
by: Song, Jiwon, et al.
Published: (2025)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
by: Ma, Da, et al.
Published: (2024)
by: Ma, Da, et al.
Published: (2024)
Optimizing Singular Spectrum for Large Language Model Compression
by: Li, Dengjie, et al.
Published: (2025)
by: Li, Dengjie, et al.
Published: (2025)
Value-Spectrum: Quantifying Preferences of Vision-Language Models via Value Decomposition in Social Media Contexts
by: Li, Jingxuan, et al.
Published: (2024)
by: Li, Jingxuan, et al.
Published: (2024)
Efficient LLM-based Advertising via Model Compression and Parallel Verification
by: Dong, Wenxin, et al.
Published: (2026)
by: Dong, Wenxin, et al.
Published: (2026)
PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel
by: Yi, Jinjun, et al.
Published: (2025)
by: Yi, Jinjun, et al.
Published: (2025)
AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
by: Ye, Jiancai, et al.
Published: (2026)
by: Ye, Jiancai, et al.
Published: (2026)
CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values
by: Koleilat, Taha, et al.
Published: (2025)
by: Koleilat, Taha, et al.
Published: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
by: Tian, Yuxuan, et al.
Published: (2025)
by: Tian, Yuxuan, et al.
Published: (2025)
Image Compression Using Singular Value Decomposition
by: Jiang, Justin
Published: (2025)
by: Jiang, Justin
Published: (2025)
GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compression
by: Liu, Kainan, et al.
Published: (2024)
by: Liu, Kainan, et al.
Published: (2024)
ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models
by: Li, Jierui, et al.
Published: (2023)
by: Li, Jierui, et al.
Published: (2023)
Jailbreak Instruction-Tuned LLMs via end-of-sentence MLP Re-weighting
by: Luo, Yifan, et al.
Published: (2024)
by: Luo, Yifan, et al.
Published: (2024)
YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference
by: Wu, You, et al.
Published: (2026)
by: Wu, You, et al.
Published: (2026)
Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering
by: Zheng, Li, et al.
Published: (2026)
by: Zheng, Li, et al.
Published: (2026)
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
by: Zhu, Kejian, et al.
Published: (2025)
by: Zhu, Kejian, et al.
Published: (2025)
Similar Items
-
HLAT: High-quality Large Language Model Pre-trained on AWS Trainium
by: Fan, Haozheng, et al.
Published: (2024) -
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
by: Song, Dinghong, et al.
Published: (2025) -
NinjaLLM: Fast, Scalable and Cost-effective RAG using Amazon SageMaker and AWS Trainium and Inferentia2
by: Xue, Tengfei, et al.
Published: (2024) -
ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
by: Yuan, Zhihang, et al.
Published: (2023) -
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
by: Wang, Xin, et al.
Published: (2024)