h4rm3l: A language for Composable Jailbreak Attack Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Doumbouya, Moussa Koulako Bala, Nandi, Ananjan, Poesia, Gabriel, Ghilardi, Davide, Goldie, Anna, Bianchi, Federico, Jurafsky, Dan, Manning, Christopher D. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tversky Neural Networks: Psychologically Plausible Deep Learning with Differentiable Tversky Similarity
by: Doumbouya, Moussa Koulako Bala, et al.
Published: (2025)
by: Doumbouya, Moussa Koulako Bala, et al.
Published: (2025)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
by: Kuric, Eduard, et al.
Published: (2026)
by: Kuric, Eduard, et al.
Published: (2026)
Selecting for Less Discriminatory Algorithms: A Relational Search Framework for Navigating Fairness-Accuracy Trade-offs in Practice
by: Samad, Hana, et al.
Published: (2025)
by: Samad, Hana, et al.
Published: (2025)
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
by: Qi, Jinhu, et al.
Published: (2026)
by: Qi, Jinhu, et al.
Published: (2026)
Membership Inference Attacks against Large Audio Language Models
by: Dong, Jia-Kai, et al.
Published: (2026)
by: Dong, Jia-Kai, et al.
Published: (2026)
Robust Uncertainty Quantification for Factual Generation of Large Language Models
by: Zhang, Yuhao, et al.
Published: (2026)
by: Zhang, Yuhao, et al.
Published: (2026)
Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI
by: Pasupuleti, Vinil, et al.
Published: (2026)
by: Pasupuleti, Vinil, et al.
Published: (2026)
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
by: Doda, Shravan
Published: (2026)
by: Doda, Shravan
Published: (2026)
T-Norm Operators for EU AI Act Compliance Classification: An Empirical Comparison of Lukasiewicz, Product, and Gödel Semantics in a Neuro-Symbolic Reasoning System
by: Laabs, Adam
Published: (2026)
by: Laabs, Adam
Published: (2026)
Sensitivity Uncertainty Alignment in Large Language Models
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
by: Hiremath, Prakul Sunil, et al.
Published: (2026)
Playing telephone with generative models: "verification disability," "compelled reliance," and accessibility in data visualization
by: Elavsky, Frank, et al.
Published: (2025)
by: Elavsky, Frank, et al.
Published: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
by: Azov, Guy, et al.
Published: (2026)
by: Azov, Guy, et al.
Published: (2026)
Explainable Artificial Intelligence (XAI) 2.0: A Manifesto of Open Challenges and Interdisciplinary Research Directions
by: Longo, Luca, et al.
Published: (2023)
by: Longo, Luca, et al.
Published: (2023)
Making AI Compliance Evidence Machine-Readable
by: Ugarte, Rodrigo Cilla, et al.
Published: (2026)
by: Ugarte, Rodrigo Cilla, et al.
Published: (2026)
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
by: Wu, Shuai, et al.
Published: (2026)
by: Wu, Shuai, et al.
Published: (2026)
Systematic Classification of Studies Investigating Social Media Conversations about Long COVID Using a Novel Zero-Shot Transformer Framework
by: Thakur, Nirmalya, et al.
Published: (2025)
by: Thakur, Nirmalya, et al.
Published: (2025)
IMUVIE: Pickup Timeline Action Localization via Motion Movies
by: Clapham, John, et al.
Published: (2024)
by: Clapham, John, et al.
Published: (2024)
Closing the SNAP Gap: Identifying Under-Enrollment in High-Poverty ZIP Codes
by: Ray, Auyona
Published: (2025)
by: Ray, Auyona
Published: (2025)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
by: Semenov, Andrei, et al.
Published: (2024)
by: Semenov, Andrei, et al.
Published: (2024)
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
by: Alberts, Lize, et al.
Published: (2024)
by: Alberts, Lize, et al.
Published: (2024)
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
by: Yun, Bhada, et al.
Published: (2026)
by: Yun, Bhada, et al.
Published: (2026)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
by: Zmanovskii, Nikita
Published: (2025)
by: Zmanovskii, Nikita
Published: (2025)
Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models
by: Ntais, Pavlos
Published: (2025)
by: Ntais, Pavlos
Published: (2025)
Classification of descriptions and summary using multiple passes of statistical and natural language toolkits
by: Banthia, Saumya, et al.
Published: (2020)
by: Banthia, Saumya, et al.
Published: (2020)
Adversarially Probing Cross-Family Sound Symbolism in 27 Languages
by: Sharma, Anika, et al.
Published: (2025)
by: Sharma, Anika, et al.
Published: (2025)
Dynamics of COVID-19 Misinformation: An Analysis of Conspiracy Theories, Fake Remedies, and False Reports
by: Thakur, Nirmalya, et al.
Published: (2025)
by: Thakur, Nirmalya, et al.
Published: (2025)
Beyond Incompatibility: Trade-offs between Mutually Exclusive Fairness Criteria in Machine Learning and Law
by: Zehlike, Meike, et al.
Published: (2022)
by: Zehlike, Meike, et al.
Published: (2022)
Can Graph-Based Microservice Performance Detection Be Used for Microservice Intrusion Detection?
by: Ma, Yunjian
Published: (2026)
by: Ma, Yunjian
Published: (2026)
A Human-Machine Collaboration Framework for the Development of Schemas
by: Isaak, Nicos
Published: (2024)
by: Isaak, Nicos
Published: (2024)
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
by: Lee, Christine, et al.
Published: (2025)
by: Lee, Christine, et al.
Published: (2025)
SmileyNet -- Towards the Prediction of the Lottery by Reading Tea Leaves with AI
by: Birk, Andreas
Published: (2024)
by: Birk, Andreas
Published: (2024)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
by: Mansour, Jad, et al.
Published: (2025)
by: Mansour, Jad, et al.
Published: (2025)
Large Language Models Are Not Strong Abstract Reasoners
by: Gendron, Gaël, et al.
Published: (2023)
by: Gendron, Gaël, et al.
Published: (2023)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
by: Mansour, Jad, et al.
Published: (2024)
by: Mansour, Jad, et al.
Published: (2024)
Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
by: Dar, Daniyal Kabir, et al.
Published: (2025)
by: Dar, Daniyal Kabir, et al.
Published: (2025)
Improving Omics-Based Classification: The Role of Feature Selection and Synthetic Data Generation
by: Perazzolo, Diego, et al.
Published: (2025)
by: Perazzolo, Diego, et al.
Published: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
by: Portelance, Eva, et al.
Published: (2023)
by: Portelance, Eva, et al.
Published: (2023)
"I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated Dialogue
by: Taheri, Atieh, et al.
Published: (2026)
by: Taheri, Atieh, et al.
Published: (2026)
CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates
by: Shaw, Ankit Kumar, et al.
Published: (2025)
by: Shaw, Ankit Kumar, et al.
Published: (2025)
Making High-Level AI Design Decisions Explicit Using a Binary Stream System-Designation Approach
by: Mossbridge, Julia
Published: (2024)
by: Mossbridge, Julia
Published: (2024)
Similar Items
-
Tversky Neural Networks: Psychologically Plausible Deep Learning with Differentiable Tversky Similarity
by: Doumbouya, Moussa Koulako Bala, et al.
Published: (2025) -
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
by: Kuric, Eduard, et al.
Published: (2026) -
Selecting for Less Discriminatory Algorithms: A Relational Search Framework for Navigating Fairness-Accuracy Trade-offs in Practice
by: Samad, Hana, et al.
Published: (2025) -
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
by: Qi, Jinhu, et al.
Published: (2026) -
Membership Inference Attacks against Large Audio Language Models
by: Dong, Jia-Kai, et al.
Published: (2026)