Enhancing Learned Knowledge in LoRA Adapters Through Efficient Contrastive Decoding on Ascend NPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Heisler, Morgan Lindsay, Xing, Linzi, Shi, Ge, Sadri, Hanieh, Singh, Gursimran, Zhang, Weiwei, Ye, Tao, Xiong, Ying, Zhang, Yong, Fan, Zhenan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service
by: Yu, Timothy Tin Long, et al.
Published: (2026)
by: Yu, Timothy Tin Long, et al.
Published: (2026)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
by: Shi, Ge, et al.
Published: (2025)
by: Shi, Ge, et al.
Published: (2025)
DECKBench: Benchmarking Multi-Agent Frameworks for Academic Slide Generation and Editing
by: Jang, Daesik, et al.
Published: (2026)
by: Jang, Daesik, et al.
Published: (2026)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
by: Singh, Gursimran, et al.
Published: (2025)
by: Singh, Gursimran, et al.
Published: (2025)
LoRACode: LoRA Adapters for Code Embeddings
by: Chaturvedi, Saumya, et al.
Published: (2025)
by: Chaturvedi, Saumya, et al.
Published: (2025)
Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
by: Shen, Yixin
Published: (2025)
by: Shen, Yixin
Published: (2025)
AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters
by: Shi, Fangming, et al.
Published: (2025)
by: Shi, Fangming, et al.
Published: (2025)
LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference
by: Rajabzadeh, Hossein, et al.
Published: (2026)
by: Rajabzadeh, Hossein, et al.
Published: (2026)
NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion
by: Chen, Chuheng, et al.
Published: (2025)
by: Chen, Chuheng, et al.
Published: (2025)
Put the Space of LoRA Initialization to the Extreme to Preserve Pre-trained Knowledge
by: Tang, Pengwei, et al.
Published: (2025)
by: Tang, Pengwei, et al.
Published: (2025)
Weight space Detection of Backdoors in LoRA Adapters
by: Merenciano, David Puertolas, et al.
Published: (2026)
by: Merenciano, David Puertolas, et al.
Published: (2026)
Do LLMs Align with My Task? Evaluating Text-to-SQL via Dataset Alignment
by: Rafiei, Davood, et al.
Published: (2025)
by: Rafiei, Davood, et al.
Published: (2025)
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
by: Li, Wenbing, et al.
Published: (2025)
by: Li, Wenbing, et al.
Published: (2025)
EAGLE-Pangu: Accelerator-Safe Tree Speculative Decoding on Ascend NPUs
by: Han, Chang, et al.
Published: (2026)
by: Han, Chang, et al.
Published: (2026)
AutoRAG-LoRA: Hallucination-Triggered Knowledge Retuning via Lightweight Adapters
by: Dwivedi, Kaushik, et al.
Published: (2025)
by: Dwivedi, Kaushik, et al.
Published: (2025)
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
Effective LoRA Adapter Routing using Task Representations
by: Dhasade, Akash, et al.
Published: (2026)
by: Dhasade, Akash, et al.
Published: (2026)
SLAD : Shared LoRA Adapters for Task Specific Distillation
by: Bensaid, Reda, et al.
Published: (2026)
by: Bensaid, Reda, et al.
Published: (2026)
Parametric Retrieval-Augmented Generation using Latent Routing of LoRA Adapters
by: Su, Zhan, et al.
Published: (2025)
by: Su, Zhan, et al.
Published: (2025)
FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
by: Zhang, Haoran, et al.
Published: (2026)
by: Zhang, Haoran, et al.
Published: (2026)
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs
by: Ye, Zhengmao, et al.
Published: (2023)
by: Ye, Zhengmao, et al.
Published: (2023)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
LoRA-FAIR: Federated LoRA Fine-Tuning with Aggregation and Initialization Refinement
by: Bian, Jieming, et al.
Published: (2024)
by: Bian, Jieming, et al.
Published: (2024)
LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters
by: Ahn, Beomjin, et al.
Published: (2026)
by: Ahn, Beomjin, et al.
Published: (2026)
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
Efficiently Serving Large Multimodal Models Using EPD Disaggregation
by: Singh, Gursimran, et al.
Published: (2024)
by: Singh, Gursimran, et al.
Published: (2024)
HypeLoRA: Hyper-Network-Generated LoRA Adapters for Calibrated Language Model Fine-Tuning
by: Trojan, Bartosz, et al.
Published: (2026)
by: Trojan, Bartosz, et al.
Published: (2026)
DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion Model
by: Xia, Siwei, et al.
Published: (2025)
by: Xia, Siwei, et al.
Published: (2025)
KD-LoRA: A Hybrid Approach to Efficient Fine-Tuning with LoRA and Knowledge Distillation
by: Azimi, Rambod, et al.
Published: (2024)
by: Azimi, Rambod, et al.
Published: (2024)
SemAug: Semantically Meaningful Image Augmentations for Object Detection Through Language Grounding
by: Heisler, Morgan, et al.
Published: (2022)
by: Heisler, Morgan, et al.
Published: (2022)
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
by: Pletenev, Sergey, et al.
Published: (2025)
by: Pletenev, Sergey, et al.
Published: (2025)
SCALE-LoRA: Auditing Post-Retrieval LoRA Composition with Residual Merging and View Reliability
by: Zhou, Shuaipeng, et al.
Published: (2026)
by: Zhou, Shuaipeng, et al.
Published: (2026)
SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA
by: Luo, Minrui, et al.
Published: (2025)
by: Luo, Minrui, et al.
Published: (2025)
Run LoRA Run: Faster and Lighter LoRA Implementations
by: Cherniuk, Daria, et al.
Published: (2023)
by: Cherniuk, Daria, et al.
Published: (2023)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
LoRA as Oracle
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
Vision as LoRA
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Rethinking Inter-LoRA Orthogonality in Adapter Merging: Insights from Orthogonal Monte Carlo Dropout
by: Zhang, Andi, et al.
Published: (2025)
by: Zhang, Andi, et al.
Published: (2025)
Similar Items
-
MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service
by: Yu, Timothy Tin Long, et al.
Published: (2026) -
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
by: Shi, Ge, et al.
Published: (2025) -
DECKBench: Benchmarking Multi-Agent Frameworks for Academic Slide Generation and Editing
by: Jang, Daesik, et al.
Published: (2026) -
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
by: Sheng, Ying, et al.
Published: (2023) -
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
by: Singh, Gursimran, et al.
Published: (2025)