Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
Fuente:
arXiv
Guardado en:
| Autor principal: | Steele, Brady |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
por: Steele, Brady
Publicado: (2026)
por: Steele, Brady
Publicado: (2026)
On the Limits of Learned Importance Scoring for KV Cache Compression
por: Steele, Brady
Publicado: (2026)
por: Steele, Brady
Publicado: (2026)
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
por: Jaber, Jaber, et al.
Publicado: (2026)
por: Jaber, Jaber, et al.
Publicado: (2026)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
por: Steele, Brady, et al.
Publicado: (2026)
por: Steele, Brady, et al.
Publicado: (2026)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
por: Ye, Hua, et al.
Publicado: (2025)
por: Ye, Hua, et al.
Publicado: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
Planning vs Reasoning: Ablations to Test Capabilities of LoRA layers
por: Redkar, Neel
Publicado: (2024)
por: Redkar, Neel
Publicado: (2024)
Targeted Lexical Injection: Unlocking Latent Cross-Lingual Alignment in Lugha-Llama via Early-Layer LoRA Fine-Tuning
por: Ngugi, Stanley
Publicado: (2025)
por: Ngugi, Stanley
Publicado: (2025)
On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning
por: Deshmukh, Pratik, et al.
Publicado: (2026)
por: Deshmukh, Pratik, et al.
Publicado: (2026)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
por: Marmoret, Axel, et al.
Publicado: (2025)
por: Marmoret, Axel, et al.
Publicado: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
por: Schesch, Benedikt, et al.
Publicado: (2026)
por: Schesch, Benedikt, et al.
Publicado: (2026)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
por: Shravan, Rohan
Publicado: (2026)
por: Shravan, Rohan
Publicado: (2026)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
por: Easley, Eric, et al.
Publicado: (2026)
por: Easley, Eric, et al.
Publicado: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
por: McCann, Jordan F.
Publicado: (2026)
por: McCann, Jordan F.
Publicado: (2026)
Super Apriel: One Checkpoint, Many Speeds
por: Labs, SLAM, et al.
Publicado: (2026)
por: Labs, SLAM, et al.
Publicado: (2026)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
por: Rozonoyer, Benjamin, et al.
Publicado: (2026)
por: Rozonoyer, Benjamin, et al.
Publicado: (2026)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
por: Salfati, Samuel
Publicado: (2026)
por: Salfati, Samuel
Publicado: (2026)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
por: Atad, Ido Andrew, et al.
Publicado: (2026)
por: Atad, Ido Andrew, et al.
Publicado: (2026)
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
por: Moghadasi, Mahdi Naser, et al.
Publicado: (2026)
por: Moghadasi, Mahdi Naser, et al.
Publicado: (2026)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
por: Saxena, Udit
Publicado: (2025)
por: Saxena, Udit
Publicado: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
por: Bakish, Yarden, et al.
Publicado: (2025)
por: Bakish, Yarden, et al.
Publicado: (2025)
QuAnTS: Question Answering on Time Series
por: Divo, Felix, et al.
Publicado: (2025)
por: Divo, Felix, et al.
Publicado: (2025)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
por: Heyman, Alex, et al.
Publicado: (2025)
por: Heyman, Alex, et al.
Publicado: (2025)
The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It
por: Garcia, Gabriel
Publicado: (2026)
por: Garcia, Gabriel
Publicado: (2026)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
por: Lelle, Travis
Publicado: (2026)
por: Lelle, Travis
Publicado: (2026)
Synthius-Mem: Brain-Inspired Hallucination-Resistant Persona Memory Achieving 94.4% Memory Accuracy and 99.6% Adversarial Robustness on LoCoMo
por: Gadzhiev, Artem, et al.
Publicado: (2026)
por: Gadzhiev, Artem, et al.
Publicado: (2026)
Quantization-Robust LLM Unlearning via Low-Rank Adaptation
por: Abitante, João Vitor Boer, et al.
Publicado: (2026)
por: Abitante, João Vitor Boer, et al.
Publicado: (2026)
OFMU: Optimization-Driven Framework for Machine Unlearning
por: Asif, Sadia, et al.
Publicado: (2025)
por: Asif, Sadia, et al.
Publicado: (2025)
Relating Misfit to Gain in Weak-to-Strong Generalization Beyond the Squared Loss
por: Mulgund, Abhijeet, et al.
Publicado: (2025)
por: Mulgund, Abhijeet, et al.
Publicado: (2025)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
por: Gu, Hao, et al.
Publicado: (2025)
por: Gu, Hao, et al.
Publicado: (2025)
Latent Cache Flow: Model-to-Model Communication Without Text
por: Rossi, Maximillian, et al.
Publicado: (2026)
por: Rossi, Maximillian, et al.
Publicado: (2026)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
por: Yuan, Xin, et al.
Publicado: (2025)
por: Yuan, Xin, et al.
Publicado: (2025)
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
por: Saghir, Hamidreza
Publicado: (2026)
por: Saghir, Hamidreza
Publicado: (2026)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
por: Bianchessi, Arthur S., et al.
Publicado: (2025)
por: Bianchessi, Arthur S., et al.
Publicado: (2025)
Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
por: Pareja, Aldo, et al.
Publicado: (2024)
por: Pareja, Aldo, et al.
Publicado: (2024)
Robustness of Spatio-temporal Graph Neural Networks for Fault Location in Partially Observable Distribution Grids
por: Karabulut, Burak, et al.
Publicado: (2026)
por: Karabulut, Burak, et al.
Publicado: (2026)
From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models
por: Li, Zehao, et al.
Publicado: (2026)
por: Li, Zehao, et al.
Publicado: (2026)
Suppressing Domain-Specific Hallucination in Construction LLMs: A Knowledge Graph Foundation for GraphRAG and QLoRA on River and Sediment Control Technical Standards
por: Yasuno, Takato
Publicado: (2026)
por: Yasuno, Takato
Publicado: (2026)
Ejemplares similares
-
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
por: Steele, Brady
Publicado: (2026) -
On the Limits of Learned Importance Scoring for KV Cache Compression
por: Steele, Brady
Publicado: (2026) -
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
por: Jaber, Jaber, et al.
Publicado: (2026) -
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
por: Steele, Brady, et al.
Publicado: (2026) -
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
por: Ye, Hua, et al.
Publicado: (2025)