FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Yu-Chen, Chen, Chong-Yan, Chang, Chi-Chih, Hu, Yu-Fang, Wu, Kai-Chiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
von: Lu, Yu-Chen, et al.
Veröffentlicht: (2025)
von: Lu, Yu-Chen, et al.
Veröffentlicht: (2025)
Palu: Compressing KV-Cache with Low-Rank Projection
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
TsqLoRA: Towards Sensitivity and Quality Low-Rank Adaptation for Efficient Fine-Tuning
von: Chen, Yu, et al.
Veröffentlicht: (2025)
von: Chen, Yu, et al.
Veröffentlicht: (2025)
EoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2024)
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2024)
MSPLoRA: A Multi-Scale Pyramid Low-Rank Adaptation for Efficient Model Fine-Tuning
von: Zhao, Jiancheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jiancheng, et al.
Veröffentlicht: (2025)
LoR2C : Low-Rank Residual Connection Adaptation for Parameter-Efficient Fine-Tuning
von: Zhao, Jiancheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jiancheng, et al.
Veröffentlicht: (2025)
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
von: Perin, Gabriel J., et al.
Veröffentlicht: (2025)
von: Perin, Gabriel J., et al.
Veröffentlicht: (2025)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
von: Hu, Xing, et al.
Veröffentlicht: (2024)
von: Hu, Xing, et al.
Veröffentlicht: (2024)
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
Penrose Tiled Low-Rank Compression and Section-Wise Q&A Fine-Tuning: A General Framework for Domain-Specific Large Language Model Adaptation
von: Kuo, Chuan-Wei, et al.
Veröffentlicht: (2025)
von: Kuo, Chuan-Wei, et al.
Veröffentlicht: (2025)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
Low-Resource Fine-Tuning for Multi-Task Structured Information Extraction with a Billion-Parameter Instruction-Tuned Model
von: Chih, Yu Cheng, et al.
Veröffentlicht: (2025)
von: Chih, Yu Cheng, et al.
Veröffentlicht: (2025)
A New Pipeline For Generating Instruction Dataset via RAG and Self Fine-Tuning
von: Song, Chih-Wei, et al.
Veröffentlicht: (2024)
von: Song, Chih-Wei, et al.
Veröffentlicht: (2024)
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
von: Zhang, Ziqian, et al.
Veröffentlicht: (2026)
von: Zhang, Ziqian, et al.
Veröffentlicht: (2026)
When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models
von: Zheng, Yingming, et al.
Veröffentlicht: (2025)
von: Zheng, Yingming, et al.
Veröffentlicht: (2025)
PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
von: Tsai, Yu-Che, et al.
Veröffentlicht: (2026)
von: Tsai, Yu-Che, et al.
Veröffentlicht: (2026)
Enhancing Low-Resource Minority Language Translation with LLMs and Retrieval-Augmented Generation for Cultural Nuances
von: Chang, Chen-Chi, et al.
Veröffentlicht: (2025)
von: Chang, Chen-Chi, et al.
Veröffentlicht: (2025)
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
LoRA+: Efficient Low Rank Adaptation of Large Models
von: Hayou, Soufiane, et al.
Veröffentlicht: (2024)
von: Hayou, Soufiane, et al.
Veröffentlicht: (2024)
Team Trifecta at Factify5WQA: Setting the Standard in Fact Verification with Fine-Tuning
von: Chiang, Shang-Hsuan, et al.
Veröffentlicht: (2024)
von: Chiang, Shang-Hsuan, et al.
Veröffentlicht: (2024)
Fine-grained Stateful Knowledge Exploration: Effective and Efficient Graph Retrieval with Large Language Models
von: Tao, Dehao, et al.
Veröffentlicht: (2024)
von: Tao, Dehao, et al.
Veröffentlicht: (2024)
LLM-RankFusion: Mitigating Intrinsic Inconsistency in LLM-based Ranking
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2024)
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2024)
Localized LoRA: A Structured Low-Rank Approximation for Efficient Fine-Tuning
von: Barazandeh, Babak, et al.
Veröffentlicht: (2025)
von: Barazandeh, Babak, et al.
Veröffentlicht: (2025)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
Semi-Supervised Learning from Small Annotated Data and Large Unlabeled Data for Fine-grained PICO Entity Recognition
von: Chen, Fangyi, et al.
Veröffentlicht: (2024)
von: Chen, Fangyi, et al.
Veröffentlicht: (2024)
TriAdaptLoRA: Brain-Inspired Triangular Adaptive Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
von: Liang, Yao, et al.
Veröffentlicht: (2025)
von: Liang, Yao, et al.
Veröffentlicht: (2025)
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
von: Chang, Chen-Chi, et al.
Veröffentlicht: (2024)
von: Chang, Chen-Chi, et al.
Veröffentlicht: (2024)
V"Mean"ba: Visual State Space Models only need 1 hidden dimension
von: Chi, Tien-Yu, et al.
Veröffentlicht: (2024)
von: Chi, Tien-Yu, et al.
Veröffentlicht: (2024)
ChameleonLLM: Batch-Aware Dynamic Low-Rank Adaptation via Inference-Time Clusters
von: Yuksel, Kamer Ali, et al.
Veröffentlicht: (2025)
von: Yuksel, Kamer Ali, et al.
Veröffentlicht: (2025)
Modeling LLM Agent Reviewer Dynamics in Elo-Ranked Review System
von: Huang, Hsiang-Wei, et al.
Veröffentlicht: (2026)
von: Huang, Hsiang-Wei, et al.
Veröffentlicht: (2026)
Towards Efficient LLM-aware Heterogeneous Graph Learning
von: Li, Wenda, et al.
Veröffentlicht: (2025)
von: Li, Wenda, et al.
Veröffentlicht: (2025)
DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
RankAdaptor: Hierarchical Rank Allocation for Efficient Fine-Tuning Pruned LLMs via Performance Model
von: Zhou, Changhai, et al.
Veröffentlicht: (2024)
von: Zhou, Changhai, et al.
Veröffentlicht: (2024)
Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows
von: Ayala, Orlando Marquez, et al.
Veröffentlicht: (2025)
von: Ayala, Orlando Marquez, et al.
Veröffentlicht: (2025)
ShareLoRA: Parameter Efficient and Robust Large Language Model Fine-tuning via Shared Low-Rank Adaptation
von: Song, Yurun, et al.
Veröffentlicht: (2024)
von: Song, Yurun, et al.
Veröffentlicht: (2024)
Matrix-Transformation Based Low-Rank Adaptation (MTLoRA): A Brain-Inspired Method for Parameter-Efficient Fine-Tuning
von: Liang, Yao, et al.
Veröffentlicht: (2024)
von: Liang, Yao, et al.
Veröffentlicht: (2024)
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
Less is More: Resource-Efficient Low-Rank Adaptation
von: Tian, Chunlin, et al.
Veröffentlicht: (2025)
von: Tian, Chunlin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
von: Lu, Yu-Chen, et al.
Veröffentlicht: (2025) -
Palu: Compressing KV-Cache with Low-Rank Projection
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024) -
TsqLoRA: Towards Sensitivity and Quality Low-Rank Adaptation for Efficient Fine-Tuning
von: Chen, Yu, et al.
Veröffentlicht: (2025) -
EoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2024) -
MSPLoRA: A Multi-Scale Pyramid Low-Rank Adaptation for Efficient Model Fine-Tuning
von: Zhao, Jiancheng, et al.
Veröffentlicht: (2025)