Fine-Tuning without Performance Degradation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Han, White, Adam, White, Martha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
von: Patterson, Andrew, et al.
Veröffentlicht: (2021)
von: Patterson, Andrew, et al.
Veröffentlicht: (2021)
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
von: Elelimy, Esraa, et al.
Veröffentlicht: (2024)
von: Elelimy, Esraa, et al.
Veröffentlicht: (2024)
Empirical Design in Reinforcement Learning
von: Patterson, Andrew, et al.
Veröffentlicht: (2023)
von: Patterson, Andrew, et al.
Veröffentlicht: (2023)
Investigating the Interplay of Prioritized Replay and Generalization
von: Panahi, Parham Mohammad, et al.
Veröffentlicht: (2024)
von: Panahi, Parham Mohammad, et al.
Veröffentlicht: (2024)
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
von: He, Jiamin, et al.
Veröffentlicht: (2026)
von: He, Jiamin, et al.
Veröffentlicht: (2026)
A New View on Planning in Online Reinforcement Learning
von: Roice, Kevin, et al.
Veröffentlicht: (2024)
von: Roice, Kevin, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning with Gradient Eligibility Traces
von: Elelimy, Esraa, et al.
Veröffentlicht: (2025)
von: Elelimy, Esraa, et al.
Veröffentlicht: (2025)
Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning
von: Wahab, Abdul, et al.
Veröffentlicht: (2026)
von: Wahab, Abdul, et al.
Veröffentlicht: (2026)
Goal-Space Planning with Subgoal Models
von: Lo, Chunlok, et al.
Veröffentlicht: (2022)
von: Lo, Chunlok, et al.
Veröffentlicht: (2022)
Gradient Iterated Temporal-Difference Learning
von: Vincent, Théo, et al.
Veröffentlicht: (2026)
von: Vincent, Théo, et al.
Veröffentlicht: (2026)
Deep Double Q-learning
von: Nagarajan, Prabhat, et al.
Veröffentlicht: (2025)
von: Nagarajan, Prabhat, et al.
Veröffentlicht: (2025)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
von: He, Jiamin, et al.
Veröffentlicht: (2025)
von: He, Jiamin, et al.
Veröffentlicht: (2025)
Demystifying the Recency Heuristic in Temporal-Difference Learning
von: Daley, Brett, et al.
Veröffentlicht: (2024)
von: Daley, Brett, et al.
Veröffentlicht: (2024)
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
von: Han, Bing, et al.
Veröffentlicht: (2025)
von: Han, Bing, et al.
Veröffentlicht: (2025)
A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning
von: Adkins, Jacob, et al.
Veröffentlicht: (2024)
von: Adkins, Jacob, et al.
Veröffentlicht: (2024)
Rethinking the Foundations for Continual Reinforcement Learning
von: Elelimy, Esraa, et al.
Veröffentlicht: (2025)
von: Elelimy, Esraa, et al.
Veröffentlicht: (2025)
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
von: Zhu, Lingwei, et al.
Veröffentlicht: (2023)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2023)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
von: Liu, Vincent, et al.
Veröffentlicht: (2023)
von: Liu, Vincent, et al.
Veröffentlicht: (2023)
Harnessing Discrete Representations For Continual Reinforcement Learning
von: Meyer, Edan, et al.
Veröffentlicht: (2023)
von: Meyer, Edan, et al.
Veröffentlicht: (2023)
Forager: a lightweight testbed for continual learning with partial observability in RL
von: Tang, Steven, et al.
Veröffentlicht: (2026)
von: Tang, Steven, et al.
Veröffentlicht: (2026)
An Analysis of Action-Value Temporal-Difference Methods That Learn State Values
von: Daley, Brett, et al.
Veröffentlicht: (2025)
von: Daley, Brett, et al.
Veröffentlicht: (2025)
Toward Data Efficient Model Merging between Different Datasets without Performance Degradation
von: Yamada, Masanori, et al.
Veröffentlicht: (2023)
von: Yamada, Masanori, et al.
Veröffentlicht: (2023)
Learning Evolving Latent Strategies for Multi-Agent Language Systems without Model Fine-Tuning
von: Tang, Wenlong
Veröffentlicht: (2025)
von: Tang, Wenlong
Veröffentlicht: (2025)
Symmetric Behavior Regularized Policy Optimization
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
AGaLiTe: Approximate Gated Linear Transformers for Online Reinforcement Learning
von: Pramanik, Subhojeet, et al.
Veröffentlicht: (2023)
von: Pramanik, Subhojeet, et al.
Veröffentlicht: (2023)
Investigating the Histogram Loss in Regression
von: Imani, Ehsan, et al.
Veröffentlicht: (2024)
von: Imani, Ehsan, et al.
Veröffentlicht: (2024)
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning
von: Zheng, Han, et al.
Veröffentlicht: (2026)
von: Zheng, Han, et al.
Veröffentlicht: (2026)
KALIE: Fine-Tuning Vision-Language Models for Open-World Manipulation without Robot Data
von: Tang, Grace, et al.
Veröffentlicht: (2024)
von: Tang, Grace, et al.
Veröffentlicht: (2024)
Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
von: Aminmansour, Farzane, et al.
Veröffentlicht: (2020)
von: Aminmansour, Farzane, et al.
Veröffentlicht: (2020)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
von: Xia, Yuchen, et al.
Veröffentlicht: (2024)
von: Xia, Yuchen, et al.
Veröffentlicht: (2024)
Performance Optimization of Ratings-Based Reinforcement Learning
von: Rose, Evelyn, et al.
Veröffentlicht: (2025)
von: Rose, Evelyn, et al.
Veröffentlicht: (2025)
AutoMixQ: Self-Adjusting Quantization for High Performance Memory-Efficient Fine-Tuning
von: Zhou, Changhai, et al.
Veröffentlicht: (2024)
von: Zhou, Changhai, et al.
Veröffentlicht: (2024)
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
Alignment Dynamics in LLM Fine-Tuning
von: Huang, Yuhan, et al.
Veröffentlicht: (2026)
von: Huang, Yuhan, et al.
Veröffentlicht: (2026)
Active Few-Shot Fine-Tuning
von: Hübotter, Jonas, et al.
Veröffentlicht: (2024)
von: Hübotter, Jonas, et al.
Veröffentlicht: (2024)
Rotation-Preserving Supervised Fine-Tuning
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
Silent Inconsistency in Data-Parallel Full Fine-Tuning: Diagnosing Worker-Level Optimization Misalignment
von: Li, Hong, et al.
Veröffentlicht: (2026)
von: Li, Hong, et al.
Veröffentlicht: (2026)
LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection
von: Zeng, Xinyue, et al.
Veröffentlicht: (2025)
von: Zeng, Xinyue, et al.
Veröffentlicht: (2025)
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
von: Yu, Ziming, et al.
Veröffentlicht: (2024)
von: Yu, Ziming, et al.
Veröffentlicht: (2024)
RPO:Reinforcement Fine-Tuning with Partial Reasoning Optimization
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
von: Patterson, Andrew, et al.
Veröffentlicht: (2021) -
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
von: Elelimy, Esraa, et al.
Veröffentlicht: (2024) -
Empirical Design in Reinforcement Learning
von: Patterson, Andrew, et al.
Veröffentlicht: (2023) -
Investigating the Interplay of Prioritized Replay and Generalization
von: Panahi, Parham Mohammad, et al.
Veröffentlicht: (2024) -
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
von: He, Jiamin, et al.
Veröffentlicht: (2026)