Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Beigi, Mohammad, Shen, Ying, Shojaee, Parshin, Wang, Qifan, Wang, Zichao, Reddy, Chandan, Jin, Ming, Huang, Lifu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges
von: Reddy, Chandan K, et al.
Veröffentlicht: (2024)
von: Reddy, Chandan K, et al.
Veröffentlicht: (2024)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
von: Abhyankar, Nikhil, et al.
Veröffentlicht: (2025)
von: Abhyankar, Nikhil, et al.
Veröffentlicht: (2025)
SURFACEBENCH: A Geometry-Aware Benchmark for Symbolic Surface Discovery
von: Kabra, Sanchit, et al.
Veröffentlicht: (2025)
von: Kabra, Sanchit, et al.
Veröffentlicht: (2025)
SNIP: Bridging Mathematical Symbolic and Numeric Realms with Unified Pre-training
von: Meidani, Kazem, et al.
Veröffentlicht: (2023)
von: Meidani, Kazem, et al.
Veröffentlicht: (2023)
Discovering Heuristics with Large Language Models (LLMs) for Mixed-Integer Programs: Single-Machine Scheduling
von: Çetinkaya, İbrahim Oğuz, et al.
Veröffentlicht: (2025)
von: Çetinkaya, İbrahim Oğuz, et al.
Veröffentlicht: (2025)
LLM-SR: Scientific Equation Discovery via Programming with Large Language Models
von: Shojaee, Parshin, et al.
Veröffentlicht: (2024)
von: Shojaee, Parshin, et al.
Veröffentlicht: (2024)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road
von: Nguyen, Ngoc-Hieu, et al.
Veröffentlicht: (2026)
von: Nguyen, Ngoc-Hieu, et al.
Veröffentlicht: (2026)
InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
von: Beigi, Mohammad, et al.
Veröffentlicht: (2024)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2024)
LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
von: Shojaee, Parshin, et al.
Veröffentlicht: (2025)
von: Shojaee, Parshin, et al.
Veröffentlicht: (2025)
Rethinking the Uncertainty: A Critical Review and Analysis in the Era of Large Language Models
von: Beigi, Mohammad, et al.
Veröffentlicht: (2024)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2024)
Navigating the Dual Facets: A Comprehensive Evaluation of Sequential Memory Editing in Large Language Models
von: Lin, Zihao, et al.
Veröffentlicht: (2024)
von: Lin, Zihao, et al.
Veröffentlicht: (2024)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
von: Shojaee, Parshin, et al.
Veröffentlicht: (2025)
von: Shojaee, Parshin, et al.
Veröffentlicht: (2025)
Error-driven Data-efficient Large Multimodal Model Tuning
von: Yao, Barry Menglong, et al.
Veröffentlicht: (2024)
von: Yao, Barry Menglong, et al.
Veröffentlicht: (2024)
Multimodal Instruction Tuning with Conditional Mixture of LoRA
von: Shen, Ying, et al.
Veröffentlicht: (2024)
von: Shen, Ying, et al.
Veröffentlicht: (2024)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
von: Kabra, Sanchit, et al.
Veröffentlicht: (2025)
von: Kabra, Sanchit, et al.
Veröffentlicht: (2025)
Reinforcement Learning-based Knowledge Distillation with LLM-as-a-Judge
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
LLM Braces: Straightening Out LLM Predictions with Relevant Sub-Updates
von: Shen, Ying, et al.
Veröffentlicht: (2025)
von: Shen, Ying, et al.
Veröffentlicht: (2025)
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
von: Qi, Jingyuan, et al.
Veröffentlicht: (2025)
von: Qi, Jingyuan, et al.
Veröffentlicht: (2025)
Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization
von: Gao, Shiping, et al.
Veröffentlicht: (2026)
von: Gao, Shiping, et al.
Veröffentlicht: (2026)
Maintaining User Trust Through Multistage Uncertainty Aware Inference
von: Agrawal, Chandan, et al.
Veröffentlicht: (2023)
von: Agrawal, Chandan, et al.
Veröffentlicht: (2023)
Debate as Optimization: Adaptive Conformal Prediction and Diverse Retrieval for Event Extraction
von: Wang, Sijia, et al.
Veröffentlicht: (2024)
von: Wang, Sijia, et al.
Veröffentlicht: (2024)
Federated Retrieval Augmented Generation for Multi-Product Question Answering
von: Shojaee, Parshin, et al.
Veröffentlicht: (2025)
von: Shojaee, Parshin, et al.
Veröffentlicht: (2025)
Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention
von: Wang, Libo
Veröffentlicht: (2024)
von: Wang, Libo
Veröffentlicht: (2024)
StructCoder: Structure-Aware Transformer for Code Generation
von: Tipirneni, Sindhu, et al.
Veröffentlicht: (2022)
von: Tipirneni, Sindhu, et al.
Veröffentlicht: (2022)
Beyond Human Preferences: Exploring Reinforcement Learning Trajectory Evaluation and Improvement through LLMs
von: Shen, Zichao, et al.
Veröffentlicht: (2024)
von: Shen, Zichao, et al.
Veröffentlicht: (2024)
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
von: Xu, Juangui, et al.
Veröffentlicht: (2025)
von: Xu, Juangui, et al.
Veröffentlicht: (2025)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning
von: Ni, Juntong, et al.
Veröffentlicht: (2026)
von: Ni, Juntong, et al.
Veröffentlicht: (2026)
Sycophancy in Large Language Models: Causes and Mitigations
von: Malmqvist, Lars
Veröffentlicht: (2024)
von: Malmqvist, Lars
Veröffentlicht: (2024)
Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
von: Shen, Si, et al.
Veröffentlicht: (2025)
von: Shen, Si, et al.
Veröffentlicht: (2025)
Accounting for Sycophancy in Language Model Uncertainty Estimation
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables
von: Abhyankar, Nikhil, et al.
Veröffentlicht: (2024)
von: Abhyankar, Nikhil, et al.
Veröffentlicht: (2024)
Distribution Shift Aware Neural Tabular Learning
von: Ying, Wangyang, et al.
Veröffentlicht: (2025)
von: Ying, Wangyang, et al.
Veröffentlicht: (2025)
FinHEAR: Human Expertise and Adaptive Risk-Aware Temporal Reasoning for Financial Decision-Making
von: Chen, Jiaxiang, et al.
Veröffentlicht: (2025)
von: Chen, Jiaxiang, et al.
Veröffentlicht: (2025)
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
von: Feng, Zhaoxin, et al.
Veröffentlicht: (2026)
von: Feng, Zhaoxin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges
von: Reddy, Chandan K, et al.
Veröffentlicht: (2024) -
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026) -
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026) -
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
von: Abhyankar, Nikhil, et al.
Veröffentlicht: (2025) -
SURFACEBENCH: A Geometry-Aware Benchmark for Symbolic Surface Discovery
von: Kabra, Sanchit, et al.
Veröffentlicht: (2025)