Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mohamadi, Mohamad Amin, Wang, Tianhao, Li, Zhiyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition
von: Mohamadi, Mohamad Amin, et al.
Veröffentlicht: (2024)
von: Mohamadi, Mohamad Amin, et al.
Veröffentlicht: (2024)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
von: Ren, Richard, et al.
Veröffentlicht: (2025)
von: Ren, Richard, et al.
Veröffentlicht: (2025)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
Training LLMs for Honesty via Confessions
von: Joglekar, Manas, et al.
Veröffentlicht: (2025)
von: Joglekar, Manas, et al.
Veröffentlicht: (2025)
EvidenceRL: Reinforcing Evidence Consistency for Trustworthy Language Models
von: Tamo, J. Ben, et al.
Veröffentlicht: (2026)
von: Tamo, J. Ben, et al.
Veröffentlicht: (2026)
How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
Honesty in Causal Forests: When It Helps and When It Hurts
von: Hou, Yanfang, et al.
Veröffentlicht: (2025)
von: Hou, Yanfang, et al.
Veröffentlicht: (2025)
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
von: Yadav, Robin, et al.
Veröffentlicht: (2025)
von: Yadav, Robin, et al.
Veröffentlicht: (2025)
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
von: Xie, Shuo, et al.
Veröffentlicht: (2025)
von: Xie, Shuo, et al.
Veröffentlicht: (2025)
Relative Kinetic Utility for Reasoning-Aware Structural Pruning in Large Language Models
von: Qian, Tianhao
Veröffentlicht: (2026)
von: Qian, Tianhao
Veröffentlicht: (2026)
Know your Trajectory -- Trustworthy Reinforcement Learning deployment through Importance-Based Trajectory Analysis
von: F, Clifford, et al.
Veröffentlicht: (2025)
von: F, Clifford, et al.
Veröffentlicht: (2025)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
von: Cundy, Chris, et al.
Veröffentlicht: (2025)
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
Structured Preconditioners in Adaptive Optimization: A Unified Analysis
von: Xie, Shuo, et al.
Veröffentlicht: (2025)
von: Xie, Shuo, et al.
Veröffentlicht: (2025)
Incentivizing Honesty among Competitors in Collaborative Learning and Optimization
von: Dorner, Florian E., et al.
Veröffentlicht: (2023)
von: Dorner, Florian E., et al.
Veröffentlicht: (2023)
AntiPaSTO: Self-Supervised Honesty Steering via Anti-Parallel Representations
von: Clark, Michael J.
Veröffentlicht: (2026)
von: Clark, Michael J.
Veröffentlicht: (2026)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
von: Taufeeque, Mohammad, et al.
Veröffentlicht: (2026)
TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models
von: Mo, Yichuan, et al.
Veröffentlicht: (2026)
von: Mo, Yichuan, et al.
Veröffentlicht: (2026)
The Marginal Value of Momentum for Small Learning Rate SGD
von: Wang, Runzhe, et al.
Veröffentlicht: (2023)
von: Wang, Runzhe, et al.
Veröffentlicht: (2023)
Enhanced High-Dimensional Data Visualization through Adaptive Multi-Scale Manifold Embedding
von: Ni, Tianhao, et al.
Veröffentlicht: (2025)
von: Ni, Tianhao, et al.
Veröffentlicht: (2025)
Think Before You Lie: How Reasoning Leads to Honesty
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
Your Offline Policy is Not Trustworthy: Bilevel Reinforcement Learning for Sequential Portfolio Optimization
von: Yuan, Haochen, et al.
Veröffentlicht: (2025)
von: Yuan, Haochen, et al.
Veröffentlicht: (2025)
PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement Learning
von: Gong, Chen, et al.
Veröffentlicht: (2025)
von: Gong, Chen, et al.
Veröffentlicht: (2025)
TrajDeleter: Enabling Trajectory Forgetting in Offline Reinforcement Learning Agents
von: Gong, Chen, et al.
Veröffentlicht: (2024)
von: Gong, Chen, et al.
Veröffentlicht: (2024)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2024)
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2024)
Beyond Accuracy: On the Effects of Fine-tuning Towards Vision-Language Model's Prediction Rationality
von: Wang, Qitong, et al.
Veröffentlicht: (2024)
von: Wang, Qitong, et al.
Veröffentlicht: (2024)
Trustworthy Classification through Rank-Based Conformal Prediction Sets
von: Luo, Rui, et al.
Veröffentlicht: (2024)
von: Luo, Rui, et al.
Veröffentlicht: (2024)
HARP: Hesitation-Aware Reframing in Transformer Inference Pass
von: Storaï, Romain, et al.
Veröffentlicht: (2024)
von: Storaï, Romain, et al.
Veröffentlicht: (2024)
Strong Transitivity Relations and Graph Neural Networks
von: Mohamadi, Yassin, et al.
Veröffentlicht: (2024)
von: Mohamadi, Yassin, et al.
Veröffentlicht: (2024)
Beyond Benchmarks: Dynamic, Automatic And Systematic Red-Teaming Agents For Trustworthy Medical Language Models
von: Pan, Jiazhen, et al.
Veröffentlicht: (2025)
von: Pan, Jiazhen, et al.
Veröffentlicht: (2025)
CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting
von: Kang, James Jin, et al.
Veröffentlicht: (2025)
von: Kang, James Jin, et al.
Veröffentlicht: (2025)
Evaluating Reinforcement Learning Safety and Trustworthiness in Cyber-Physical Systems
von: Dearstyne, Katherine, et al.
Veröffentlicht: (2025)
von: Dearstyne, Katherine, et al.
Veröffentlicht: (2025)
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
A Survey on Medical Large Language Models: Technology, Application, Trustworthiness, and Future Directions
von: Liu, Lei, et al.
Veröffentlicht: (2024)
von: Liu, Lei, et al.
Veröffentlicht: (2024)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
von: Xing, Shuo, et al.
Veröffentlicht: (2024)
von: Xing, Shuo, et al.
Veröffentlicht: (2024)
Open-World Reinforcement Learning over Long Short-Term Imagination
von: Li, Jiajian, et al.
Veröffentlicht: (2024)
von: Li, Jiajian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity
von: Xie, Shuo, et al.
Veröffentlicht: (2024) -
Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition
von: Mohamadi, Mohamad Amin, et al.
Veröffentlicht: (2024) -
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
von: Ren, Richard, et al.
Veröffentlicht: (2025) -
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024) -
Training LLMs for Honesty via Confessions
von: Joglekar, Manas, et al.
Veröffentlicht: (2025)