TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Zhepei, Yang, Xiao, Sun, Kai, Wang, Jiaqi, Shao, Rulin, Chen, Sean, Kachuee, Mohammad, Gollapudi, Teja, Liao, Tony, Scheffer, Nicolas, Wanga, Rakesh, Kumar, Anuj, Meng, Yu, Yih, Wen-tau, Dong, Xin Luna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Incentivized Truthful Communication for Federated Bandits
by: Wei, Zhepei, et al.
Published: (2024)
by: Wei, Zhepei, et al.
Published: (2024)
PrismRAG: Boosting RAG Factuality with Distractor Resilience and Strategized Reasoning
by: Kachuee, Mohammad, et al.
Published: (2025)
by: Kachuee, Mohammad, et al.
Published: (2025)
SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
by: Liu, Shicheng, et al.
Published: (2025)
by: Liu, Shicheng, et al.
Published: (2025)
Truth
by: Cubitt, Sean
Published: (2024)
by: Cubitt, Sean
Published: (2024)
Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs?
by: Sun, Kai, et al.
Published: (2025)
by: Sun, Kai, et al.
Published: (2025)
Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning
by: Sim, Rachael Hwee Ling, et al.
Published: (2026)
by: Sim, Rachael Hwee Ling, et al.
Published: (2026)
Incentivizing Truthful Collaboration in Heterogeneous Federated Learning
by: Chakarov, Dimitar, et al.
Published: (2024)
by: Chakarov, Dimitar, et al.
Published: (2024)
The Unreasonable Effectiveness of Eccentric Automatic Prompts
by: Battle, Rick, et al.
Published: (2024)
by: Battle, Rick, et al.
Published: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
by: Fu, Yao, et al.
Published: (2025)
by: Fu, Yao, et al.
Published: (2025)
Improving Factuality with Explicit Working Memory
by: Chen, Mingda, et al.
Published: (2024)
by: Chen, Mingda, et al.
Published: (2024)
Incentivizing Truthful Language Models via Peer Elicitation Games
by: Chen, Baiting, et al.
Published: (2025)
by: Chen, Baiting, et al.
Published: (2025)
Incentivizing Truthful Data Contributions in a Marketplace for Mean Estimation
by: Chen, Keran, et al.
Published: (2025)
by: Chen, Keran, et al.
Published: (2025)
Elegance, Facts, and Scientific Truths
by: Gisin, Nicolas
Published: (2024)
by: Gisin, Nicolas
Published: (2024)
Learning to Reason for Factuality
by: Chen, Xilun, et al.
Published: (2025)
by: Chen, Xilun, et al.
Published: (2025)
How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
by: Adarsh, Shivam, et al.
Published: (2026)
by: Adarsh, Shivam, et al.
Published: (2026)
A Cramér-von Mises Approach to Incentivizing Truthful Data Sharing
by: Clinton, Alex, et al.
Published: (2025)
by: Clinton, Alex, et al.
Published: (2025)
On the Universal Truthfulness Hyperplane Inside LLMs
by: Liu, Junteng, et al.
Published: (2024)
by: Liu, Junteng, et al.
Published: (2024)
Testing the Limits of Truth Directions in LLMs
by: Poulis, Angelos, et al.
Published: (2026)
by: Poulis, Angelos, et al.
Published: (2026)
Love the Truth, the Whole Truth, and the Truth about Everything. An Interview with Josef Seifert
by: Rodrigo Guerra López
Published: (2014)
by: Rodrigo Guerra López
Published: (2014)
Latency Adjustable Transformer Encoder for Language Understanding
by: Kachuee, Sajjad, et al.
Published: (2022)
by: Kachuee, Sajjad, et al.
Published: (2022)
Efficient Large Language Models with Zero-Shot Adjustable Acceleration
by: Kachuee, Sajjad, et al.
Published: (2025)
by: Kachuee, Sajjad, et al.
Published: (2025)
Geometry-Preserving Aggregation for Mixture-of-Experts Embedding Models
by: Kachuee, Sajjad, et al.
Published: (2026)
by: Kachuee, Sajjad, et al.
Published: (2026)
The Truth, the Whole Truth, and Nothing but the Truth: Automatic Visualization Evaluation from Reconstruction Quality
by: Bujack, Roxana, et al.
Published: (2026)
by: Bujack, Roxana, et al.
Published: (2026)
Few-Shot Data Synthesis for Open Domain Multi-Hop Question Answering
by: Chen, Mingda, et al.
Published: (2023)
by: Chen, Mingda, et al.
Published: (2023)
Truthful Aggregation of LLMs with an Application to Online Advertising
by: Soumalias, Ermis, et al.
Published: (2024)
by: Soumalias, Ermis, et al.
Published: (2024)
Truth is Universal: Robust Detection of Lies in LLMs
by: Bürger, Lennart, et al.
Published: (2024)
by: Bürger, Lennart, et al.
Published: (2024)
Truth Knows No Language: Evaluating Truthfulness Beyond English
by: Figueras, Blanca Calvo, et al.
Published: (2025)
by: Figueras, Blanca Calvo, et al.
Published: (2025)
TruthStance: An Annotated Dataset of Conversations on Truth Social
by: Ameen, Fathima, et al.
Published: (2026)
by: Ameen, Fathima, et al.
Published: (2026)
ConfRAG: Confidence-Guided Retrieval-Augmenting Generation
by: Huang, Yin, et al.
Published: (2025)
by: Huang, Yin, et al.
Published: (2025)
Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks
by: Bao, Yuntai, et al.
Published: (2025)
by: Bao, Yuntai, et al.
Published: (2025)
In Search of Truth: In memory of Balraj Singh
by: Orce, José Nicolás, et al.
Published: (2024)
by: Orce, José Nicolás, et al.
Published: (2024)
The Shape of Truth
by: FERNANDEZ, DAHLIA D.
Published: (2025)
by: FERNANDEZ, DAHLIA D.
Published: (2025)
Similar Items
-
Incentivized Truthful Communication for Federated Bandits
by: Wei, Zhepei, et al.
Published: (2024) -
PrismRAG: Boosting RAG Factuality with Distractor Resilience and Strategized Reasoning
by: Kachuee, Mohammad, et al.
Published: (2025) -
SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
by: Liu, Shicheng, et al.
Published: (2025) -
Truth
by: Cubitt, Sean
Published: (2024) -
Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs?
by: Sun, Kai, et al.
Published: (2025)