LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Minqian, Xu, Zhiyang, Zhang, Xinyi, An, Heajun, Qadir, Sarvech, Zhang, Qi, Wisniewski, Pamela J., Cho, Jin-Hee, Lee, Sang Won, Jia, Ruoxi, Huang, Lifu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Integrated Solutions: A Systematic Interdisciplinary Review of Cybergrooming Research
by: An, Heajun, et al.
Published: (2025)
by: An, Heajun, et al.
Published: (2025)
StagePilot: A Deep Reinforcement Learning Agent for Stage-Controlled Cybergrooming Simulation
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
From Vulnerable to Resilient: Examining Parent and Teen Perceptions on How to Respond to Unwanted Cybergrooming Advances
by: Zhang, Xinyi, et al.
Published: (2026)
by: Zhang, Xinyi, et al.
Published: (2026)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
Generating A Crowdsourced Conversation Dataset to Combat Cybergrooming
by: Zhang, Xinyi, et al.
Published: (2024)
by: Zhang, Xinyi, et al.
Published: (2024)
Building a Village: A Multi-stakeholder Approach to Open Innovation and Shared Governance to Promote Youth Online Safety
by: Caddle, Xavier V., et al.
Published: (2025)
by: Caddle, Xavier V., et al.
Published: (2025)
MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
by: Qi, Jingyuan, et al.
Published: (2023)
by: Qi, Jingyuan, et al.
Published: (2023)
Holistic Evaluation for Interleaved Text-and-Image Generation
by: Liu, Minqian, et al.
Published: (2024)
by: Liu, Minqian, et al.
Published: (2024)
Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
by: Xu, Zhiyang, et al.
Published: (2024)
by: Xu, Zhiyang, et al.
Published: (2024)
AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments
by: Saenger, Till Raphael, et al.
Published: (2024)
by: Saenger, Till Raphael, et al.
Published: (2024)
Markov Persuasion Processes: Learning to Persuade from Scratch
by: Bacchiocchi, Francesco, et al.
Published: (2024)
by: Bacchiocchi, Francesco, et al.
Published: (2024)
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
by: Liu, Minqian, et al.
Published: (2023)
by: Liu, Minqian, et al.
Published: (2023)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
by: Jaipersaud, Brandon, et al.
Published: (2025)
by: Jaipersaud, Brandon, et al.
Published: (2025)
Enhancing Dialogue Generation in Werewolf Game Through Situation Analysis and Persuasion Strategies
by: Qi, Zhiyang, et al.
Published: (2024)
by: Qi, Zhiyang, et al.
Published: (2024)
Navigating Ideation Space: Decomposed Conceptual Representations for Positioning Scientific Ideas
by: Shen, Yuexi, et al.
Published: (2026)
by: Shen, Yuexi, et al.
Published: (2026)
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
VEXA: Evidence-Grounded and Persona-Adaptive Explanations for Scam Risk Sensemaking
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
From Parental Control to Joint Family Oversight: Can Parents and Teens Manage Mobile Online Safety and Privacy as Equals?
by: Akter, Mamtaj, et al.
Published: (2022)
by: Akter, Mamtaj, et al.
Published: (2022)
AMELI: Enhancing Multimodal Entity Linking with Fine-Grained Attributes
by: Yao, Barry Menglong, et al.
Published: (2023)
by: Yao, Barry Menglong, et al.
Published: (2023)
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
by: Qi, Jingyuan, et al.
Published: (2025)
by: Qi, Jingyuan, et al.
Published: (2025)
When AI Gets Persuaded, Humans Follow: Inducing the Conformity Effect in Persuasive Dialogue
by: Sasaki, Rikuo, et al.
Published: (2025)
by: Sasaki, Rikuo, et al.
Published: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Moving Beyond Parental Control toward Community-based Approaches to Adolescent Online Safety
by: Akter, Mamtaj, et al.
Published: (2025)
by: Akter, Mamtaj, et al.
Published: (2025)
Persuading a Credible Agent
by: Gan, Jiarui, et al.
Published: (2024)
by: Gan, Jiarui, et al.
Published: (2024)
Understanding the Perceptions of Trigger Warning and Content Warning on Social Media Platforms in the U.S
by: Zhang, Xinyi, et al.
Published: (2025)
by: Zhang, Xinyi, et al.
Published: (2025)
An Event-triggered System for Social Persuasion and Danger Alert in Elder Home Monitoring
by: Liu, Jun-Yi, et al.
Published: (2025)
by: Liu, Jun-Yi, et al.
Published: (2025)
UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers
by: Min, Dehai, et al.
Published: (2024)
by: Min, Dehai, et al.
Published: (2024)
Persuading while Learning
by: Arieli, Itai, et al.
Published: (2024)
by: Arieli, Itai, et al.
Published: (2024)
Towards Resilience and Autonomy-based Approaches for Adolescents Online Safety
by: Park, Jinkyung, et al.
Published: (2025)
by: Park, Jinkyung, et al.
Published: (2025)
Towards Collaborative Family-Centered Design for Online Safety, Privacy and Security
by: Akter, Mamtaj, et al.
Published: (2024)
by: Akter, Mamtaj, et al.
Published: (2024)
LLMs Can Plan Only If We Tell Them
by: Sel, Bilgehan, et al.
Published: (2025)
by: Sel, Bilgehan, et al.
Published: (2025)
How Do Large Language Models Learn Concepts During Continual Pre-Training?
by: Yao, Barry Menglong, et al.
Published: (2026)
by: Yao, Barry Menglong, et al.
Published: (2026)
SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
LLM Braces: Straightening Out LLM Predictions with Relevant Sub-Updates
by: Shen, Ying, et al.
Published: (2025)
by: Shen, Ying, et al.
Published: (2025)
Make an Offer They Can't Refuse: Grounding Bayesian Persuasion in Real-World Dialogues without Pre-Commitment
by: He, Buwei, et al.
Published: (2025)
by: He, Buwei, et al.
Published: (2025)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
by: Al-Tawaha, Ahmad, et al.
Published: (2026)
by: Al-Tawaha, Ahmad, et al.
Published: (2026)
Persuasion and Safety in the Era of Generative AI
by: Kong, Haein
Published: (2025)
by: Kong, Haein
Published: (2025)
Persuading Stable Matching
by: Shaki, Jonathan, et al.
Published: (2025)
by: Shaki, Jonathan, et al.
Published: (2025)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
by: Potter, Yujin, et al.
Published: (2024)
by: Potter, Yujin, et al.
Published: (2024)
Similar Items
-
Toward Integrated Solutions: A Systematic Interdisciplinary Review of Cybergrooming Research
by: An, Heajun, et al.
Published: (2025) -
StagePilot: A Deep Reinforcement Learning Agent for Stage-Controlled Cybergrooming Simulation
by: An, Heajun, et al.
Published: (2026) -
From Vulnerable to Resilient: Examining Parent and Teen Perceptions on How to Respond to Unwanted Cybergrooming Advances
by: Zhang, Xinyi, et al.
Published: (2026) -
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026) -
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
by: Zeng, Yi, et al.
Published: (2024)