Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hahm, Dongyoon, Min, Taywon, Jin, Woogyeol, Lee, Kimin |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Enhancing LLM Agent Safety via Causal Influence Prompting
par: Hahm, Dongyoon, et autres
Publié: (2025)
par: Hahm, Dongyoon, et autres
Publié: (2025)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
par: Hahm, Dongyoon, et autres
Publié: (2026)
par: Hahm, Dongyoon, et autres
Publié: (2026)
Benchmarking Mobile Device Control Agents across Diverse Configurations
par: Lee, Juyong, et autres
Publié: (2024)
par: Lee, Juyong, et autres
Publié: (2024)
Understanding Impact of Human Feedback via Influence Functions
par: Min, Taywon, et autres
Publié: (2025)
par: Min, Taywon, et autres
Publié: (2025)
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
par: Szep, Marton, et autres
Publié: (2026)
par: Szep, Marton, et autres
Publié: (2026)
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
par: Lee, Juyong, et autres
Publié: (2024)
par: Lee, Juyong, et autres
Publié: (2024)
Mitigating Unintended Memorization with LoRA in Federated Learning for LLMs
par: Bossy, Thierry, et autres
Publié: (2025)
par: Bossy, Thierry, et autres
Publié: (2025)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
par: Giordani, Jeremiah
Publié: (2025)
par: Giordani, Jeremiah
Publié: (2025)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
par: Liang, Kaiqu, et autres
Publié: (2025)
par: Liang, Kaiqu, et autres
Publié: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
par: Choi, Sooyung, et autres
Publié: (2025)
par: Choi, Sooyung, et autres
Publié: (2025)
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
par: Lai, Song, et autres
Publié: (2025)
par: Lai, Song, et autres
Publié: (2025)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
par: Diao, Muxi, et autres
Publié: (2026)
par: Diao, Muxi, et autres
Publié: (2026)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
par: Kim, Dongyoung, et autres
Publié: (2024)
par: Kim, Dongyoung, et autres
Publié: (2024)
State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
par: Choi, Juheon, et autres
Publié: (2025)
par: Choi, Juheon, et autres
Publié: (2025)
CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting Mitigation
par: Fawi, Muhammad
Publié: (2024)
par: Fawi, Muhammad
Publié: (2024)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
par: Ju, Yiming, et autres
Publié: (2024)
par: Ju, Yiming, et autres
Publié: (2024)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
par: Kabra, Sanchit, et autres
Publié: (2025)
par: Kabra, Sanchit, et autres
Publié: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
par: Pandey, Punya Syon, et autres
Publié: (2025)
par: Pandey, Punya Syon, et autres
Publié: (2025)
Proximal Supervised Fine-Tuning
par: Zhu, Wenhong, et autres
Publié: (2025)
par: Zhu, Wenhong, et autres
Publié: (2025)
Instruction Tuning with Human Curriculum
par: Lee, Bruce W., et autres
Publié: (2023)
par: Lee, Bruce W., et autres
Publié: (2023)
Order-Independence Without Fine Tuning
par: McIlroy-Young, Reid, et autres
Publié: (2024)
par: McIlroy-Young, Reid, et autres
Publié: (2024)
Model Editing by Standard Fine-Tuning
par: Gangadhar, Govind, et autres
Publié: (2024)
par: Gangadhar, Govind, et autres
Publié: (2024)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
par: Chen, Xiaoyi, et autres
Publié: (2023)
par: Chen, Xiaoyi, et autres
Publié: (2023)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
par: Goel, Raghavv, et autres
Publié: (2024)
par: Goel, Raghavv, et autres
Publié: (2024)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
par: Wang, Zige, et autres
Publié: (2025)
par: Wang, Zige, et autres
Publié: (2025)
Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer
par: Du, Guodong, et autres
Publié: (2025)
par: Du, Guodong, et autres
Publié: (2025)
Parameter-Efficient Fine-Tuning for Foundation Models
par: Zhang, Dan, et autres
Publié: (2025)
par: Zhang, Dan, et autres
Publié: (2025)
Supervised Fine-Tuning as Inverse Reinforcement Learning
par: Sun, Hao
Publié: (2024)
par: Sun, Hao
Publié: (2024)
Boosting Large Language Models with Mask Fine-Tuning
par: Zhang, Mingyuan, et autres
Publié: (2025)
par: Zhang, Mingyuan, et autres
Publié: (2025)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
par: Huang, Zeyu, et autres
Publié: (2025)
par: Huang, Zeyu, et autres
Publié: (2025)
Teaching LLMs How to Learn with Contextual Fine-Tuning
par: Choi, Younwoo, et autres
Publié: (2025)
par: Choi, Younwoo, et autres
Publié: (2025)
Mitigate Negative Transfer with Similarity Heuristic Lifelong Prompt Tuning
par: Wu, Chenyuan, et autres
Publié: (2024)
par: Wu, Chenyuan, et autres
Publié: (2024)
Fine-Tuning and Evaluating Conversational AI for Agricultural Advisory
par: Singh, Sanyam, et autres
Publié: (2026)
par: Singh, Sanyam, et autres
Publié: (2026)
Parameter-Efficient Fine-Tuning with Discrete Fourier Transform
par: Gao, Ziqi, et autres
Publié: (2024)
par: Gao, Ziqi, et autres
Publié: (2024)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
par: Hameed, Marawan Gamal Abdel, et autres
Publié: (2024)
par: Hameed, Marawan Gamal Abdel, et autres
Publié: (2024)
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
par: Ovadia, Oded, et autres
Publié: (2023)
par: Ovadia, Oded, et autres
Publié: (2023)
Fine-Tuning Language Models with Reward Learning on Policy
par: Lang, Hao, et autres
Publié: (2024)
par: Lang, Hao, et autres
Publié: (2024)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
par: Xia, Yuchen, et autres
Publié: (2024)
par: Xia, Yuchen, et autres
Publié: (2024)
Scaling Sparse Fine-Tuning to Large Language Models
par: Ansell, Alan, et autres
Publié: (2024)
par: Ansell, Alan, et autres
Publié: (2024)
Instruction Fine-Tuning: Does Prompt Loss Matter?
par: Huerta-Enochian, Mathew, et autres
Publié: (2024)
par: Huerta-Enochian, Mathew, et autres
Publié: (2024)
Documents similaires
-
Enhancing LLM Agent Safety via Causal Influence Prompting
par: Hahm, Dongyoon, et autres
Publié: (2025) -
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
par: Hahm, Dongyoon, et autres
Publié: (2026) -
Benchmarking Mobile Device Control Agents across Diverse Configurations
par: Lee, Juyong, et autres
Publié: (2024) -
Understanding Impact of Human Feedback via Influence Functions
par: Min, Taywon, et autres
Publié: (2025) -
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
par: Szep, Marton, et autres
Publié: (2026)