Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
Fuente:
arXiv
Salvato in:
| Autore principale: | Steele, Brady |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
di: Steele, Brady
Pubblicazione: (2026)
di: Steele, Brady
Pubblicazione: (2026)
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
di: Fadli, Samih
Pubblicazione: (2025)
di: Fadli, Samih
Pubblicazione: (2025)
Targeted Lexical Injection: Unlocking Latent Cross-Lingual Alignment in Lugha-Llama via Early-Layer LoRA Fine-Tuning
di: Ngugi, Stanley
Pubblicazione: (2025)
di: Ngugi, Stanley
Pubblicazione: (2025)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
di: Walker, Nicholas
Pubblicazione: (2024)
di: Walker, Nicholas
Pubblicazione: (2024)
On the Limits of Learned Importance Scoring for KV Cache Compression
di: Steele, Brady
Pubblicazione: (2026)
di: Steele, Brady
Pubblicazione: (2026)
Continuous-Depth Transformers with Learned Control Dynamics
di: Jemley, Peter
Pubblicazione: (2026)
di: Jemley, Peter
Pubblicazione: (2026)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
di: Steele, Brady, et al.
Pubblicazione: (2026)
di: Steele, Brady, et al.
Pubblicazione: (2026)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
di: Szilvasy, Gergely, et al.
Pubblicazione: (2026)
di: Szilvasy, Gergely, et al.
Pubblicazione: (2026)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
di: Lelle, Travis
Pubblicazione: (2026)
di: Lelle, Travis
Pubblicazione: (2026)
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
di: Garg, Aashna, et al.
Pubblicazione: (2026)
di: Garg, Aashna, et al.
Pubblicazione: (2026)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
di: Ye, Hua, et al.
Pubblicazione: (2025)
di: Ye, Hua, et al.
Pubblicazione: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
Planning vs Reasoning: Ablations to Test Capabilities of LoRA layers
di: Redkar, Neel
Pubblicazione: (2024)
di: Redkar, Neel
Pubblicazione: (2024)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
di: Feng, Qi, et al.
Pubblicazione: (2025)
di: Feng, Qi, et al.
Pubblicazione: (2025)
Quantization-Robust LLM Unlearning via Low-Rank Adaptation
di: Abitante, João Vitor Boer, et al.
Pubblicazione: (2026)
di: Abitante, João Vitor Boer, et al.
Pubblicazione: (2026)
OCRR: A Benchmark for Online Correction Recovery under Distribution Shift
di: Grassi, Adrian
Pubblicazione: (2026)
di: Grassi, Adrian
Pubblicazione: (2026)
Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
di: Naser-Moghadasi, Mahdi, et al.
Pubblicazione: (2026)
di: Naser-Moghadasi, Mahdi, et al.
Pubblicazione: (2026)
Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure
di: Mahmood, Syed Naveed, et al.
Pubblicazione: (2026)
di: Mahmood, Syed Naveed, et al.
Pubblicazione: (2026)
TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning
di: Das, Susmit
Pubblicazione: (2026)
di: Das, Susmit
Pubblicazione: (2026)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
di: Patel, Het, et al.
Pubblicazione: (2026)
di: Patel, Het, et al.
Pubblicazione: (2026)
Memory Bank Compression for Continual Adaptation of Large Language Models
di: Katraouras, Thomas, et al.
Pubblicazione: (2026)
di: Katraouras, Thomas, et al.
Pubblicazione: (2026)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
di: Shravan, Rohan
Pubblicazione: (2026)
di: Shravan, Rohan
Pubblicazione: (2026)
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
di: Saghir, Hamidreza
Pubblicazione: (2026)
di: Saghir, Hamidreza
Pubblicazione: (2026)
Stratified Hazard Sampling: Minimal-Variance Event Scheduling for CTMC/DTMC Discrete Diffusion and Flow Models
di: Jang, Seunghwan, et al.
Pubblicazione: (2026)
di: Jang, Seunghwan, et al.
Pubblicazione: (2026)
Mean-Pooled Cosine Similarity is Not Length-Invariant: Theory and Cross-Domain Evidence for a Length-Invariant Alternative
di: Mitra, Sibayan, et al.
Pubblicazione: (2026)
di: Mitra, Sibayan, et al.
Pubblicazione: (2026)
The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It
di: Garcia, Gabriel
Pubblicazione: (2026)
di: Garcia, Gabriel
Pubblicazione: (2026)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
di: Gu, Hao, et al.
Pubblicazione: (2025)
di: Gu, Hao, et al.
Pubblicazione: (2025)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
di: Young, Richard J., et al.
Pubblicazione: (2025)
di: Young, Richard J., et al.
Pubblicazione: (2025)
PersonalLLM: Tailoring LLMs to Individual Preferences
di: Zollo, Thomas P., et al.
Pubblicazione: (2024)
di: Zollo, Thomas P., et al.
Pubblicazione: (2024)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
di: Bianchessi, Arthur S., et al.
Pubblicazione: (2025)
di: Bianchessi, Arthur S., et al.
Pubblicazione: (2025)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
di: Zhu, Jiajun, et al.
Pubblicazione: (2025)
di: Zhu, Jiajun, et al.
Pubblicazione: (2025)
Sarcasm Detection in a Less-Resourced Language
di: Đoković, Lazar, et al.
Pubblicazione: (2024)
di: Đoković, Lazar, et al.
Pubblicazione: (2024)
Thread Detection and Response Generation using Transformers with Prompt Optimisation
di: T, Kevin Joshua, et al.
Pubblicazione: (2024)
di: T, Kevin Joshua, et al.
Pubblicazione: (2024)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
di: Ponnock, Jesse
Pubblicazione: (2025)
di: Ponnock, Jesse
Pubblicazione: (2025)
Combining Language and Topic Models for Hierarchical Text Classification
di: Toit, Jaco du, et al.
Pubblicazione: (2025)
di: Toit, Jaco du, et al.
Pubblicazione: (2025)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
di: He, Yanjin, et al.
Pubblicazione: (2025)
di: He, Yanjin, et al.
Pubblicazione: (2025)
Language Models Are Implicitly Continuous
di: Marro, Samuele, et al.
Pubblicazione: (2025)
di: Marro, Samuele, et al.
Pubblicazione: (2025)
LLM Vocabulary Compression for Low-Compute Environments
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
di: Saini, Mayank, et al.
Pubblicazione: (2025)
di: Saini, Mayank, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
di: Steele, Brady
Pubblicazione: (2026) -
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
di: Jaber, Jaber, et al.
Pubblicazione: (2026) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
di: Fadli, Samih
Pubblicazione: (2025) -
Targeted Lexical Injection: Unlocking Latent Cross-Lingual Alignment in Lugha-Llama via Early-Layer LoRA Fine-Tuning
di: Ngugi, Stanley
Pubblicazione: (2025) -
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
di: Walker, Nicholas
Pubblicazione: (2024)