It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics
Fuente:
arXiv
Saved in:
| Main Authors: | Kowal, Matthew, Timm, Jasper, Godbout, Jean-Francois, Costello, Thomas, Arechar, Antonio A., Pennycook, Gordon, Rand, David, Gleave, Adam, Pelrine, Kellin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large language models can effectively convince people to believe conspiracies
by: Costello, Thomas H., et al.
Published: (2026)
by: Costello, Thomas H., et al.
Published: (2026)
Emergent Persuasion: Will LLMs Persuade Without Being Prompted?
by: Chang, Vincent, et al.
Published: (2025)
by: Chang, Vincent, et al.
Published: (2025)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026)
by: Struppek, Lukas, et al.
Published: (2026)
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
by: Rivera, Mauricio, et al.
Published: (2024)
by: Rivera, Mauricio, et al.
Published: (2024)
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
by: Vergho, Tyler, et al.
Published: (2024)
by: Vergho, Tyler, et al.
Published: (2024)
Scaling Trends for Data Poisoning in LLMs
by: Bowen, Dillon, et al.
Published: (2024)
by: Bowen, Dillon, et al.
Published: (2024)
Can Go AIs be adversarially robust?
by: Tseng, Tom, et al.
Published: (2024)
by: Tseng, Tom, et al.
Published: (2024)
Exploiting Novel GPT-4 APIs
by: Pelrine, Kellin, et al.
Published: (2023)
by: Pelrine, Kellin, et al.
Published: (2023)
Uncertainty Resolution in Misinformation Detection
by: Orlovskiy, Yury, et al.
Published: (2024)
by: Orlovskiy, Yury, et al.
Published: (2024)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
by: Kowal, Matthew, et al.
Published: (2026)
by: Kowal, Matthew, et al.
Published: (2026)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
Online Influence Campaigns: Strategies and Vulnerabilities
by: Musulan, Andreea, et al.
Published: (2024)
by: Musulan, Andreea, et al.
Published: (2024)
A Guide to Misinformation Detection Data and Evaluation
by: Thibault, Camille, et al.
Published: (2024)
by: Thibault, Camille, et al.
Published: (2024)
Regional and Temporal Patterns of Partisan Polarization during the COVID-19 Pandemic in the United States and Canada
by: Yang, Zachary, et al.
Published: (2024)
by: Yang, Zachary, et al.
Published: (2024)
CrediBench: Building Web-Scale Network Datasets for Information Integrity
by: Kondrup, Emma, et al.
Published: (2025)
by: Kondrup, Emma, et al.
Published: (2025)
Web Retrieval Agents for Evidence-Based Misinformation Detection
by: Tian, Jacob-Junqi, et al.
Published: (2024)
by: Tian, Jacob-Junqi, et al.
Published: (2024)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Epistemic Integrity in Large Language Models
by: Ghafouri, Bijean, et al.
Published: (2024)
by: Ghafouri, Bijean, et al.
Published: (2024)
Nuevas tecnologías para entender el comportamiento económico
by: Antonio Alonso Aréchar
Published: (2017)
by: Antonio Alonso Aréchar
Published: (2017)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
by: Gibbs, Tom, et al.
Published: (2024)
by: Gibbs, Tom, et al.
Published: (2024)
Veracity: An Open-Source AI Fact-Checking System
by: Curtis, Taylor Lynn, et al.
Published: (2025)
by: Curtis, Taylor Lynn, et al.
Published: (2025)
From Intuition to Understanding: Using AI Peers to Overcome Physics Misconceptions
by: Weijers, Ruben, et al.
Published: (2025)
by: Weijers, Ruben, et al.
Published: (2025)
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
by: Hossain, Saad, et al.
Published: (2026)
by: Hossain, Saad, et al.
Published: (2026)
Persuadability and LLMs as Legal Decision Tools
by: Suttle, Oisin, et al.
Published: (2026)
by: Suttle, Oisin, et al.
Published: (2026)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
by: Cundy, Chris, et al.
Published: (2025)
by: Cundy, Chris, et al.
Published: (2025)
WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
by: Maurya, Sneha, et al.
Published: (2026)
by: Maurya, Sneha, et al.
Published: (2026)
La política del texto
by: Alastair Pennycook
Published: (2011)
by: Alastair Pennycook
Published: (2011)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
by: Potter, Yujin, et al.
Published: (2024)
by: Potter, Yujin, et al.
Published: (2024)
When Agents Persuade: Rhetoric Generation and Mitigation in LLMs
by: Jose, Julia, et al.
Published: (2026)
by: Jose, Julia, et al.
Published: (2026)
Paying and Persuading
by: Luo, Daniel
Published: (2025)
by: Luo, Daniel
Published: (2025)
Persuaded Search
by: Mekonnen, Teddy, et al.
Published: (2023)
by: Mekonnen, Teddy, et al.
Published: (2023)
Topics of Thought
by: Berto, Francesco
Published: (2022)
by: Berto, Francesco
Published: (2022)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
by: Sun, Xiangkun, et al.
Published: (2026)
by: Sun, Xiangkun, et al.
Published: (2026)
A Simulation System Towards Solving Societal-Scale Manipulation
by: Touzel, Maximilian Puelma, et al.
Published: (2024)
by: Touzel, Maximilian Puelma, et al.
Published: (2024)
Understanding Chain-of-Thought in LLMs through Information Theory
by: Ton, Jean-Francois, et al.
Published: (2024)
by: Ton, Jean-Francois, et al.
Published: (2024)
Is It More Common to Persuade Others to Break Up Online? The Influence of Perceived Anonymity on Online Breakup Persuasion Attempts in Others' Romantic Conflict
by: Hongyi Lin, et al.
Published: (2025)
by: Hongyi Lin, et al.
Published: (2025)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
by: Bowen, Dillon, et al.
Published: (2025)
by: Bowen, Dillon, et al.
Published: (2025)
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
by: Dombrowski, Ann-Kathrin, et al.
Published: (2025)
by: Dombrowski, Ann-Kathrin, et al.
Published: (2025)
Persuading Stable Matching
by: Shaki, Jonathan, et al.
Published: (2025)
by: Shaki, Jonathan, et al.
Published: (2025)
Similar Items
-
Large language models can effectively convince people to believe conspiracies
by: Costello, Thomas H., et al.
Published: (2026) -
Emergent Persuasion: Will LLMs Persuade Without Being Prompted?
by: Chang, Vincent, et al.
Published: (2025) -
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026) -
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
by: Rivera, Mauricio, et al.
Published: (2024) -
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
by: Vergho, Tyler, et al.
Published: (2024)