Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
Fuente:
arXiv
Saved in:
| Main Authors: | Creo, Aldan, Fernandez, Raul Castro, Cebrian, Manuel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Complete Evasion, Zero Modification: PDF Attacks on AI Text Detection
by: Creo, Aldan
Published: (2025)
by: Creo, Aldan
Published: (2025)
SilverSpeak: Evading AI-Generated Text Detectors using Homoglyphs
by: Creo, Aldan, et al.
Published: (2024)
by: Creo, Aldan, et al.
Published: (2024)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
by: Hawkins, John, et al.
Published: (2025)
by: Hawkins, John, et al.
Published: (2025)
Emergent evaluation hubs in a decentralizing large language model ecosystem
by: Cebrian, Manuel, et al.
Published: (2025)
by: Cebrian, Manuel, et al.
Published: (2025)
Ask a Local: Detecting Hallucinations With Specialized Model Divergence
by: Creo, Aldan, et al.
Published: (2025)
by: Creo, Aldan, et al.
Published: (2025)
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
Can adversarial attacks by large language models be attributed?
by: Cebrian, Manuel, et al.
Published: (2024)
by: Cebrian, Manuel, et al.
Published: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection
by: Kazemi, Arefeh, et al.
Published: (2025)
by: Kazemi, Arefeh, et al.
Published: (2025)
Evaluation Framework for AI Systems in "the Wild"
by: Jabbour, Sarah, et al.
Published: (2025)
by: Jabbour, Sarah, et al.
Published: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Conversational Complexity for Assessing Risk in Large Language Models
by: Burden, John, et al.
Published: (2024)
by: Burden, John, et al.
Published: (2024)
All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks
by: Takemoto, Kazuhiro
Published: (2024)
by: Takemoto, Kazuhiro
Published: (2024)
"What's Up, Doc?": Analyzing How Users Seek Health Information in Large-Scale Conversational AI Datasets
by: Paruchuri, Akshay, et al.
Published: (2025)
by: Paruchuri, Akshay, et al.
Published: (2025)
Founder effects shape the evolutionary dynamics of multimodality in open LLM families
by: Cebrian, Manuel
Published: (2026)
by: Cebrian, Manuel
Published: (2026)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
by: Liu, Yunting, et al.
Published: (2024)
by: Liu, Yunting, et al.
Published: (2024)
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
by: Bai, Yuqi, et al.
Published: (2025)
by: Bai, Yuqi, et al.
Published: (2025)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024)
by: Dong, Zhichen, et al.
Published: (2024)
Commercial Persuasion in AI-Mediated Conversations
by: Salvi, Francesco, et al.
Published: (2026)
by: Salvi, Francesco, et al.
Published: (2026)
Mechanical Enforcement for LLM Governance:Evidence of Governance-Task Decoupling in Financial Decision Systems
by: Rodríguez, José Manuel de la Chica, et al.
Published: (2026)
by: Rodríguez, José Manuel de la Chica, et al.
Published: (2026)
Large Language Models as Misleading Assistants in Conversation
by: Hou, Betty Li, et al.
Published: (2024)
by: Hou, Betty Li, et al.
Published: (2024)
Hanging in the Balance: Pivotal Moments in Crisis Counseling Conversations
by: Nguyen, Vivian, et al.
Published: (2025)
by: Nguyen, Vivian, et al.
Published: (2025)
The Carbon Cost of Conversation, Sustainability in the Age of Language Models
by: Amiri, Sayed Mahbub Hasan, et al.
Published: (2025)
by: Amiri, Sayed Mahbub Hasan, et al.
Published: (2025)
How Did We Get Here? Summarizing Conversation Dynamics
by: Hua, Yilun, et al.
Published: (2024)
by: Hua, Yilun, et al.
Published: (2024)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
by: Ding, Junchen, et al.
Published: (2025)
by: Ding, Junchen, et al.
Published: (2025)
Incorporating LLMs for Large-Scale Urban Complex Mobility Simulation
by: Song, Yu-Lun, et al.
Published: (2025)
by: Song, Yu-Lun, et al.
Published: (2025)
A Linguistic Comparison between Human and ChatGPT-Generated Conversations
by: Sandler, Morgan, et al.
Published: (2024)
by: Sandler, Morgan, et al.
Published: (2024)
The Unequal Opportunities of Large Language Models: Revealing Demographic Bias through Job Recommendations
by: Salinas, Abel, et al.
Published: (2023)
by: Salinas, Abel, et al.
Published: (2023)
Knowledge Acquisition on Mass-shooting Events via LLMs for AI-Driven Justice
by: Ihugba, Benign John, et al.
Published: (2025)
by: Ihugba, Benign John, et al.
Published: (2025)
Wait! There's a Way Out: A Decision Mechanism for Forecasting Conversational Derailment
by: Kim, Laerdon, et al.
Published: (2026)
by: Kim, Laerdon, et al.
Published: (2026)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024)
by: Ji, Haoxuan, et al.
Published: (2024)
Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild
by: Hicke, Rebecca M. M., et al.
Published: (2026)
by: Hicke, Rebecca M. M., et al.
Published: (2026)
Teaching Language Models How to Code Like Learners: Conversational Serialization for Student Simulation
by: Koutcheme, Charles, et al.
Published: (2026)
by: Koutcheme, Charles, et al.
Published: (2026)
Taking a turn for the better: Conversation redirection throughout the course of mental-health therapy
by: Nguyen, Vivian, et al.
Published: (2024)
by: Nguyen, Vivian, et al.
Published: (2024)
Conversational Agents for Building Energy Efficiency -- Advising Housing Cooperatives in Stockholm on Reducing Energy Consumption
by: Ghani, Shadaab, et al.
Published: (2025)
by: Ghani, Shadaab, et al.
Published: (2025)
NLP Meets the World: Toward Improving Conversations With the Public About Natural Language Processing Research
by: Wilson, Shomir
Published: (2025)
by: Wilson, Shomir
Published: (2025)
Cultural Compass: A Framework for Organizing Societal Norms to Detect Violations in Human-AI Conversations
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
LLM Nepotism in Organizational Governance
by: Mao, Shunqi, et al.
Published: (2026)
by: Mao, Shunqi, et al.
Published: (2026)
Similar Items
-
Complete Evasion, Zero Modification: PDF Attacks on AI Text Detection
by: Creo, Aldan
Published: (2025) -
SilverSpeak: Evading AI-Generated Text Detectors using Homoglyphs
by: Creo, Aldan, et al.
Published: (2024) -
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
by: Hawkins, John, et al.
Published: (2025) -
Emergent evaluation hubs in a decentralizing large language model ecosystem
by: Cebrian, Manuel, et al.
Published: (2025) -
Ask a Local: Detecting Hallucinations With Specialized Model Divergence
by: Creo, Aldan, et al.
Published: (2025)