Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
Fuente:
arXiv
Salvato in:
| Autore principale: | Sahoo, Subramanyam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
di: Krishnan, Ranganath, et al.
Pubblicazione: (2024)
di: Krishnan, Ranganath, et al.
Pubblicazione: (2024)
Teaching LLMs How to Learn with Contextual Fine-Tuning
di: Choi, Younwoo, et al.
Pubblicazione: (2025)
di: Choi, Younwoo, et al.
Pubblicazione: (2025)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
di: Nikitin, Alexander, et al.
Pubblicazione: (2024)
di: Nikitin, Alexander, et al.
Pubblicazione: (2024)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
di: Wang, Ziyu, et al.
Pubblicazione: (2024)
di: Wang, Ziyu, et al.
Pubblicazione: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
di: Chen, Lichang, et al.
Pubblicazione: (2024)
di: Chen, Lichang, et al.
Pubblicazione: (2024)
Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
di: Ling, Zhenqing, et al.
Pubblicazione: (2025)
di: Ling, Zhenqing, et al.
Pubblicazione: (2025)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
di: Pan, Alexander, et al.
Pubblicazione: (2024)
di: Pan, Alexander, et al.
Pubblicazione: (2024)
Fine-Tuning Language Models with Reward Learning on Policy
di: Lang, Hao, et al.
Pubblicazione: (2024)
di: Lang, Hao, et al.
Pubblicazione: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
di: Ackermann, Johannes, et al.
Pubblicazione: (2026)
di: Ackermann, Johannes, et al.
Pubblicazione: (2026)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
di: Giordani, Jeremiah
Pubblicazione: (2025)
di: Giordani, Jeremiah
Pubblicazione: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
di: Wang, Ye, et al.
Pubblicazione: (2026)
di: Wang, Ye, et al.
Pubblicazione: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
di: Ovadia, Oded, et al.
Pubblicazione: (2023)
di: Ovadia, Oded, et al.
Pubblicazione: (2023)
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
di: Bouchard, Dylan, et al.
Pubblicazione: (2026)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
di: Dai, Hui, et al.
Pubblicazione: (2026)
di: Dai, Hui, et al.
Pubblicazione: (2026)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
di: Guo, Song, et al.
Pubblicazione: (2024)
di: Guo, Song, et al.
Pubblicazione: (2024)
iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use
di: Zeng, Yirong, et al.
Pubblicazione: (2025)
di: Zeng, Yirong, et al.
Pubblicazione: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
di: Sanyal, Sunny, et al.
Pubblicazione: (2024)
di: Sanyal, Sunny, et al.
Pubblicazione: (2024)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
Fine-Tuning LLMs for Report Summarization: Analysis on Supervised and Unsupervised Data
di: Rallapalli, Swati, et al.
Pubblicazione: (2025)
di: Rallapalli, Swati, et al.
Pubblicazione: (2025)
Efficient Differentially Private Fine-Tuning of LLMs via Reinforcement Learning
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
Hallucination Detection in LLMs: Fast and Memory-Efficient Fine-Tuned Models
di: Arteaga, Gabriel Y., et al.
Pubblicazione: (2024)
di: Arteaga, Gabriel Y., et al.
Pubblicazione: (2024)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
di: Liu, Liangxin, et al.
Pubblicazione: (2024)
di: Liu, Liangxin, et al.
Pubblicazione: (2024)
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
di: Misra, Dipendra, et al.
Pubblicazione: (2026)
di: Misra, Dipendra, et al.
Pubblicazione: (2026)
The Price of Format: Diversity Collapse in LLMs
di: Yun, Longfei, et al.
Pubblicazione: (2025)
di: Yun, Longfei, et al.
Pubblicazione: (2025)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
di: Kim, Taeho, et al.
Pubblicazione: (2024)
di: Kim, Taeho, et al.
Pubblicazione: (2024)
ALKAFI-LLAMA3: Fine-Tuning LLMs for Precise Legal Understanding in Palestine
di: Qasem, Rabee, et al.
Pubblicazione: (2024)
di: Qasem, Rabee, et al.
Pubblicazione: (2024)
Prompting and Fine-Tuning of Small LLMs for Length-Controllable Telephone Call Summarization
di: Thulke, David, et al.
Pubblicazione: (2024)
di: Thulke, David, et al.
Pubblicazione: (2024)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
di: Yang, Daniel, et al.
Pubblicazione: (2026)
di: Yang, Daniel, et al.
Pubblicazione: (2026)
Show Me How It's Done: The Role of Explanations in Fine-Tuning Language Models
di: Ballout, Mohamad, et al.
Pubblicazione: (2024)
di: Ballout, Mohamad, et al.
Pubblicazione: (2024)
Deconfounded Causality-aware Parameter-Efficient Fine-Tuning for Problem-Solving Improvement of LLMs
di: Wang, Ruoyu, et al.
Pubblicazione: (2024)
di: Wang, Ruoyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025) -
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
di: Krishnan, Ranganath, et al.
Pubblicazione: (2024) -
Teaching LLMs How to Learn with Contextual Fine-Tuning
di: Choi, Younwoo, et al.
Pubblicazione: (2025) -
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
di: Nikitin, Alexander, et al.
Pubblicazione: (2024) -
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
di: Wang, Ziyu, et al.
Pubblicazione: (2024)