WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Liwei, Rao, Kavel, Han, Seungju, Ettinger, Allyson, Brahman, Faeze, Kumar, Sachin, Mireshghallah, Niloofar, Lu, Ximing, Sap, Maarten, Choi, Yejin, Dziri, Nouha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
di: Han, Seungju, et al.
Pubblicazione: (2024)
di: Han, Seungju, et al.
Pubblicazione: (2024)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
di: Rao, Kavel, et al.
Pubblicazione: (2023)
di: Rao, Kavel, et al.
Pubblicazione: (2023)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
di: Lu, Ximing, et al.
Pubblicazione: (2024)
di: Lu, Ximing, et al.
Pubblicazione: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
di: Baheti, Ashutosh, et al.
Pubblicazione: (2024)
di: Baheti, Ashutosh, et al.
Pubblicazione: (2024)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
di: Sorensen, Taylor, et al.
Pubblicazione: (2023)
di: Sorensen, Taylor, et al.
Pubblicazione: (2023)
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2024)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
di: Baheti, Ashutosh, et al.
Pubblicazione: (2023)
di: Baheti, Ashutosh, et al.
Pubblicazione: (2023)
A Roadmap to Pluralistic Alignment
di: Sorensen, Taylor, et al.
Pubblicazione: (2024)
di: Sorensen, Taylor, et al.
Pubblicazione: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
Information-Theoretic Distillation for Reference-less Summarization
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
To Err is AI : A Case Study Informing LLM Flaw Reporting Practices
di: McGregor, Sean, et al.
Pubblicazione: (2024)
di: McGregor, Sean, et al.
Pubblicazione: (2024)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
di: Jiang, Liwei, et al.
Pubblicazione: (2025)
di: Jiang, Liwei, et al.
Pubblicazione: (2025)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
di: Jung, Jaehun, et al.
Pubblicazione: (2023)
di: Jung, Jaehun, et al.
Pubblicazione: (2023)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2024)
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2024)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
di: Lee, Jaeyoung, et al.
Pubblicazione: (2024)
di: Lee, Jaeyoung, et al.
Pubblicazione: (2024)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
di: Su, Zhe, et al.
Pubblicazione: (2024)
di: Su, Zhe, et al.
Pubblicazione: (2024)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
di: Qiu, Linlu, et al.
Pubblicazione: (2023)
di: Qiu, Linlu, et al.
Pubblicazione: (2023)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
di: Chen, Tong, et al.
Pubblicazione: (2025)
di: Chen, Tong, et al.
Pubblicazione: (2025)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
RESTOR: Knowledge Recovery in Machine Unlearning
di: Rezaei, Keivan, et al.
Pubblicazione: (2024)
di: Rezaei, Keivan, et al.
Pubblicazione: (2024)
WildChat: 1M ChatGPT Interaction Logs in the Wild
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
di: Sorensen, Taylor, et al.
Pubblicazione: (2025)
di: Sorensen, Taylor, et al.
Pubblicazione: (2025)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
di: Lu, Yujie, et al.
Pubblicazione: (2024)
di: Lu, Yujie, et al.
Pubblicazione: (2024)
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
di: Ravichander, Abhilasha, et al.
Pubblicazione: (2025)
di: Ravichander, Abhilasha, et al.
Pubblicazione: (2025)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
di: Mendelsohn, Julia, et al.
Pubblicazione: (2023)
di: Mendelsohn, Julia, et al.
Pubblicazione: (2023)
Position: Privacy Is Not Just Memorization!
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2025)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
di: Zheng, Mingqian, et al.
Pubblicazione: (2025)
di: Zheng, Mingqian, et al.
Pubblicazione: (2025)
In Search of the Long-Tail: Systematic Generation of Long-Tail Inferential Knowledge via Logical Rule Guided Search
di: Li, Huihan, et al.
Pubblicazione: (2023)
di: Li, Huihan, et al.
Pubblicazione: (2023)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
di: Rahman, Salman, et al.
Pubblicazione: (2025)
di: Rahman, Salman, et al.
Pubblicazione: (2025)
Reasoning Up the Instruction Ladder for Controllable Language Models
di: Zheng, Zishuo, et al.
Pubblicazione: (2025)
di: Zheng, Zishuo, et al.
Pubblicazione: (2025)
PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
di: Hallinan, Skyler, et al.
Pubblicazione: (2025)
di: Hallinan, Skyler, et al.
Pubblicazione: (2025)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
di: Li, Huihan, et al.
Pubblicazione: (2024)
di: Li, Huihan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
di: Han, Seungju, et al.
Pubblicazione: (2024) -
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
di: Rao, Kavel, et al.
Pubblicazione: (2023) -
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
di: Lu, Ximing, et al.
Pubblicazione: (2024) -
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024) -
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
di: Baheti, Ashutosh, et al.
Pubblicazione: (2024)