Debating with More Persuasive LLMs Leads to More Truthful Answers
Fuente:
arXiv
Guardado en:
| Autores principales: | Khan, Akbir, Hughes, John, Valentine, Dan, Ruis, Laura, Sachan, Kshitij, Radhakrishnan, Ansh, Grefenstette, Edward, Bowman, Samuel R., Rocktäschel, Tim, Perez, Ethan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
por: Cook, Jonathan, et al.
Publicado: (2025)
por: Cook, Jonathan, et al.
Publicado: (2025)
Scaling Opponent Shaping to High Dimensional Games
por: Khan, Akbir, et al.
Publicado: (2023)
por: Khan, Akbir, et al.
Publicado: (2023)
minimax: Efficient Baselines for Autocurricula in JAX
por: Jiang, Minqi, et al.
Publicado: (2023)
por: Jiang, Minqi, et al.
Publicado: (2023)
Investigating Non-Transitivity in LLM-as-a-Judge
por: Xu, Yi, et al.
Publicado: (2025)
por: Xu, Yi, et al.
Publicado: (2025)
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
por: Ruis, Laura, et al.
Publicado: (2024)
por: Ruis, Laura, et al.
Publicado: (2024)
Infusion: Shaping Model Behavior by Editing Training Data via Influence Functions
por: Rosser, J, et al.
Publicado: (2026)
por: Rosser, J, et al.
Publicado: (2026)
AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
por: Carro, María Victoria, et al.
Publicado: (2025)
por: Carro, María Victoria, et al.
Publicado: (2025)
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
por: Wen, Jiaxin, et al.
Publicado: (2024)
por: Wen, Jiaxin, et al.
Publicado: (2024)
Interaction Dynamics as a Reward Signal for LLMs
por: Gooding, Sian, et al.
Publicado: (2025)
por: Gooding, Sian, et al.
Publicado: (2025)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
Do Papers with Titles Ending in a Question Mark Usually Have the Answer "No"?
por: Stern, Daniel, et al.
Publicado: (2026)
por: Stern, Daniel, et al.
Publicado: (2026)
What's More Important: The Questions or the Answers?
por: Morgan, Eric Lease
Publicado: (1999)
por: Morgan, Eric Lease
Publicado: (1999)
Providing More than Just an Answer.
por: Low, Kathleen
Publicado: (1990)
por: Low, Kathleen
Publicado: (1990)
Language Models Learn to Mislead Humans via RLHF
por: Wen, Jiaxin, et al.
Publicado: (2024)
por: Wen, Jiaxin, et al.
Publicado: (2024)
Learning to Lead Themselves: Agentic AI in MAS using MARL
por: Kamthan, Ansh
Publicado: (2025)
por: Kamthan, Ansh
Publicado: (2025)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
por: Agarwal, Mahak, et al.
Publicado: (2025)
por: Agarwal, Mahak, et al.
Publicado: (2025)
More Than Saying “It's AI”: How Role Disclosure Transparency in AI‐Generated Ads Influences Persuasion
por: Khanh Bao Quang Le, et al.
Publicado: (2026)
por: Khanh Bao Quang Le, et al.
Publicado: (2026)
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
por: Jain, Samyak, et al.
Publicado: (2023)
por: Jain, Samyak, et al.
Publicado: (2023)
Make Your VLA More Robust Without More Data By Interleaving Motion Planning
por: Choe, Dan BW, et al.
Publicado: (2026)
por: Choe, Dan BW, et al.
Publicado: (2026)
Technology--More than an Answer in Search of a Question.
Publicado: (1997)
Publicado: (1997)
Do More Suspicious Transaction Reports Lead to More Convictions for Money Laundering?
por: Jensen, Rasmus Ingemann Tuffveson, et al.
Publicado: (2025)
por: Jensen, Rasmus Ingemann Tuffveson, et al.
Publicado: (2025)
Factorio Learning Environment
por: Hopkins, Jack, et al.
Publicado: (2025)
por: Hopkins, Jack, et al.
Publicado: (2025)
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
por: Paglieri, Davide, et al.
Publicado: (2024)
por: Paglieri, Davide, et al.
Publicado: (2024)
Less is More: Towards Sustainability-Aware Persuasive Explanations in Recommender Systems
por: Tran, Thi Ngoc Trang, et al.
Publicado: (2024)
por: Tran, Thi Ngoc Trang, et al.
Publicado: (2024)
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
por: Paglieri, Davide, et al.
Publicado: (2025)
por: Paglieri, Davide, et al.
Publicado: (2025)
PINNs in More General Geometry
por: Hirst, Edward
Publicado: (2026)
por: Hirst, Edward
Publicado: (2026)
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
por: Huang, Jing, et al.
Publicado: (2026)
por: Huang, Jing, et al.
Publicado: (2026)
When Large Language Models are More PersuasiveThan Incentivized Humans, and Why
por: Schoenegger, Philipp, et al.
Publicado: (2025)
por: Schoenegger, Philipp, et al.
Publicado: (2025)
LLM-First Search: Self-Guided Exploration of the Solution Space
por: Herr, Nathan, et al.
Publicado: (2025)
por: Herr, Nathan, et al.
Publicado: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
por: Pres, Itamar, et al.
Publicado: (2024)
por: Pres, Itamar, et al.
Publicado: (2024)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
por: Paglieri, Davide, et al.
Publicado: (2024)
por: Paglieri, Davide, et al.
Publicado: (2024)
More Than Machines?
por: Voss, Laura
Publicado: (2023)
por: Voss, Laura
Publicado: (2023)
LIMO: Less is More for Reasoning
por: Ye, Yixin, et al.
Publicado: (2025)
por: Ye, Yixin, et al.
Publicado: (2025)
PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
por: Çelebi, Yusuf, et al.
Publicado: (2025)
por: Çelebi, Yusuf, et al.
Publicado: (2025)
Echo Chambers and Information Brokers on Truth Social: A Study of Network Dynamics and Political Discourse
por: Hughes, Emelia May, et al.
Publicado: (2025)
por: Hughes, Emelia May, et al.
Publicado: (2025)
SH2: Self-Highlighted Hesitation Helps You Decode More Truthfully
por: Kai, Jushi, et al.
Publicado: (2024)
por: Kai, Jushi, et al.
Publicado: (2024)
AI Control: Improving Safety Despite Intentional Subversion
por: Greenblatt, Ryan, et al.
Publicado: (2023)
por: Greenblatt, Ryan, et al.
Publicado: (2023)
Right this way: Can VLMs Guide Us to See More to Answer Questions?
por: Liu, Li, et al.
Publicado: (2024)
por: Liu, Li, et al.
Publicado: (2024)
Open-Endedness is Essential for Artificial Superhuman Intelligence
por: Hughes, Edward, et al.
Publicado: (2024)
por: Hughes, Edward, et al.
Publicado: (2024)
More on Property($M$) in separable Banach spaces: Revised
por: Dalby, Tim
Publicado: (2024)
por: Dalby, Tim
Publicado: (2024)
Ejemplares similares
-
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
por: Cook, Jonathan, et al.
Publicado: (2025) -
Scaling Opponent Shaping to High Dimensional Games
por: Khan, Akbir, et al.
Publicado: (2023) -
minimax: Efficient Baselines for Autocurricula in JAX
por: Jiang, Minqi, et al.
Publicado: (2023) -
Investigating Non-Transitivity in LLM-as-a-Judge
por: Xu, Yi, et al.
Publicado: (2025) -
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
por: Ruis, Laura, et al.
Publicado: (2024)