Best-of-N Jailbreaking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hughes, John, Price, Sara, Lynch, Aengus, Schaeffer, Rylan, Barez, Fazl, Koyejo, Sanmi, Sleight, Henry, Jones, Erik, Perez, Ethan, Sharma, Mrinank |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
von: Wang, Tony T., et al.
Veröffentlicht: (2024)
von: Wang, Tony T., et al.
Veröffentlicht: (2024)
Chain-of-Thought Hijacking
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
How Do Large Language Monkeys Get Their Power (Laws)?
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
von: Peng, Alwin, et al.
Veröffentlicht: (2024)
von: Peng, Alwin, et al.
Veröffentlicht: (2024)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
von: Obbad, Elyas, et al.
Veröffentlicht: (2024)
von: Obbad, Elyas, et al.
Veröffentlicht: (2024)
Investigating Data Contamination for Pre-training Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
Position: Model Collapse Does Not Mean What You Think
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2024)
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2024)
In-Context Learning of Energy Functions
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
von: Kaunismaa, Jackson, et al.
Veröffentlicht: (2026)
von: Kaunismaa, Jackson, et al.
Veröffentlicht: (2026)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
What Causes Polysemanticity? An Alternative Origin Story of Mixed Selectivity from Incidental Causes
von: Lecomte, Victor, et al.
Veröffentlicht: (2023)
von: Lecomte, Victor, et al.
Veröffentlicht: (2023)
Pretraining Scaling Laws for Generative Evaluations of Language Models
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
The Persistent Vulnerability of Aligned AI Systems
von: Lynch, Aengus
Veröffentlicht: (2026)
von: Lynch, Aengus
Veröffentlicht: (2026)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2026)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2026)
Looking Inward: Language Models Can Learn About Themselves by Introspection
von: Binder, Felix J, et al.
Veröffentlicht: (2024)
von: Binder, Felix J, et al.
Veröffentlicht: (2024)
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Efficient Prediction of Pass@k Scaling in Large Language Models
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
von: Kazdan, Joshua, et al.
Veröffentlicht: (2024)
von: Kazdan, Joshua, et al.
Veröffentlicht: (2024)
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
von: Denisov-Blanch, Yegor, et al.
Veröffentlicht: (2026)
von: Denisov-Blanch, Yegor, et al.
Veröffentlicht: (2026)
Query Circuits: Explaining How Language Models Answer User Prompts
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
Why Do Safety Guardrails Degrade Across Languages?
von: Zhang, Max, et al.
Veröffentlicht: (2026)
von: Zhang, Max, et al.
Veröffentlicht: (2026)
Logits are All We Need to Adapt Closed Models
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
Understanding Addition and Subtraction in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
Extracting books from production language models
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2026)
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
von: Wang, Tony T., et al.
Veröffentlicht: (2024) -
Chain-of-Thought Hijacking
von: Zhao, Jianli, et al.
Veröffentlicht: (2025) -
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024) -
How Do Large Language Monkeys Get Their Power (Laws)?
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025) -
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
von: Peng, Alwin, et al.
Veröffentlicht: (2024)