Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Tony T., Hughes, John, Sleight, Henry, Schaeffer, Rylan, Agrawal, Rajashree, Barez, Fazl, Sharma, Mrinank, Mu, Jesse, Shavit, Nir, Perez, Ethan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Best-of-N Jailbreaking
por: Hughes, John, et al.
Publicado: (2024)
por: Hughes, John, et al.
Publicado: (2024)
Chain-of-Thought Hijacking
por: Zhao, Jianli, et al.
Publicado: (2025)
por: Zhao, Jianli, et al.
Publicado: (2025)
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
por: Schaeffer, Rylan, et al.
Publicado: (2024)
por: Schaeffer, Rylan, et al.
Publicado: (2024)
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
por: Peng, Alwin, et al.
Publicado: (2024)
por: Peng, Alwin, et al.
Publicado: (2024)
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
por: Youstra, Jack, et al.
Publicado: (2025)
por: Youstra, Jack, et al.
Publicado: (2025)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
por: Fu, Tingchen, et al.
Publicado: (2024)
por: Fu, Tingchen, et al.
Publicado: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
por: Chaudhary, Maheep, et al.
Publicado: (2025)
por: Chaudhary, Maheep, et al.
Publicado: (2025)
Understanding Addition in Transformers
por: Quirke, Philip, et al.
Publicado: (2023)
por: Quirke, Philip, et al.
Publicado: (2023)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
por: Fu, Tingchen, et al.
Publicado: (2025)
por: Fu, Tingchen, et al.
Publicado: (2025)
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
por: Cunningham, Hoagy, et al.
Publicado: (2026)
por: Cunningham, Hoagy, et al.
Publicado: (2026)
Query Circuits: Explaining How Language Models Answer User Prompts
por: Wu, Tung-Yu, et al.
Publicado: (2025)
por: Wu, Tung-Yu, et al.
Publicado: (2025)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
por: Gerstgrasser, Matthias, et al.
Publicado: (2024)
por: Gerstgrasser, Matthias, et al.
Publicado: (2024)
Rethinking AI Cultural Alignment
por: Bravansky, Michal, et al.
Publicado: (2025)
por: Bravansky, Michal, et al.
Publicado: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
por: Lan, Michael, et al.
Publicado: (2023)
por: Lan, Michael, et al.
Publicado: (2023)
Understanding Addition and Subtraction in Transformers
por: Quirke, Philip, et al.
Publicado: (2024)
por: Quirke, Philip, et al.
Publicado: (2024)
Token Taxes: mitigating AGI's economic risks
por: Irwin, Lucas, et al.
Publicado: (2026)
por: Irwin, Lucas, et al.
Publicado: (2026)
Large Language Models Relearn Removed Concepts
por: Lo, Michelle, et al.
Publicado: (2024)
por: Lo, Michelle, et al.
Publicado: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
por: Gupta, Aman, et al.
Publicado: (2025)
por: Gupta, Aman, et al.
Publicado: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
por: Neo, Clement, et al.
Publicado: (2024)
por: Neo, Clement, et al.
Publicado: (2024)
Learning to Interpret Weight Differences in Language Models
por: Goel, Avichal, et al.
Publicado: (2025)
por: Goel, Avichal, et al.
Publicado: (2025)
Forbidden Facts: An Investigation of Competing Objectives in Llama-2
por: Wang, Tony T., et al.
Publicado: (2023)
por: Wang, Tony T., et al.
Publicado: (2023)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
por: Marks, Luke, et al.
Publicado: (2024)
por: Marks, Luke, et al.
Publicado: (2024)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
por: Heindrich, Lovis, et al.
Publicado: (2025)
por: Heindrich, Lovis, et al.
Publicado: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
por: Schrodi, Simon, et al.
Publicado: (2025)
por: Schrodi, Simon, et al.
Publicado: (2025)
Visualizing Neural Network Imagination
por: Wichers, Nevan, et al.
Publicado: (2024)
por: Wichers, Nevan, et al.
Publicado: (2024)
On the Complexity of Neural Computation in Superposition
por: Adler, Micah, et al.
Publicado: (2024)
por: Adler, Micah, et al.
Publicado: (2024)
Toy Combinatorial Interpretability Models Reveal Lottery Tickets in Early Feature Space
por: Bebchuk, Alon, et al.
Publicado: (2026)
por: Bebchuk, Alon, et al.
Publicado: (2026)
In-Context Learning of Energy Functions
por: Schaeffer, Rylan, et al.
Publicado: (2024)
por: Schaeffer, Rylan, et al.
Publicado: (2024)
NeuroADDA: Active Discriminative Domain Adaptation in Connectomic
por: Sawmya, Shashata, et al.
Publicado: (2025)
por: Sawmya, Shashata, et al.
Publicado: (2025)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
por: Kim, Taeyoun, et al.
Publicado: (2024)
por: Kim, Taeyoun, et al.
Publicado: (2024)
Scalable Energy-Based Models via Adversarial Training: Unifying Discrimination and Generation
por: Yin, Xuwang, et al.
Publicado: (2025)
por: Yin, Xuwang, et al.
Publicado: (2025)
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
por: Wen, Jiaxin, et al.
Publicado: (2024)
por: Wen, Jiaxin, et al.
Publicado: (2024)
Embodied AI: Emerging Risks and Opportunities for Policy Action
por: Perlo, Jared, et al.
Publicado: (2025)
por: Perlo, Jared, et al.
Publicado: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
por: Oozeer, Narmeen, et al.
Publicado: (2025)
por: Oozeer, Narmeen, et al.
Publicado: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
por: Simhi, Adi, et al.
Publicado: (2025)
por: Simhi, Adi, et al.
Publicado: (2025)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
por: Oldfield, James, et al.
Publicado: (2025)
por: Oldfield, James, et al.
Publicado: (2025)
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
por: Li, Changyi, et al.
Publicado: (2026)
por: Li, Changyi, et al.
Publicado: (2026)
Scaling sparse feature circuit finding for in-context learning
por: Kharlapenko, Dmitrii, et al.
Publicado: (2025)
por: Kharlapenko, Dmitrii, et al.
Publicado: (2025)
Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models
por: Schaeffer, Rylan, et al.
Publicado: (2025)
por: Schaeffer, Rylan, et al.
Publicado: (2025)
Forecasting Rare Language Model Behaviors
por: Jones, Erik, et al.
Publicado: (2025)
por: Jones, Erik, et al.
Publicado: (2025)
Ejemplares similares
-
Best-of-N Jailbreaking
por: Hughes, John, et al.
Publicado: (2024) -
Chain-of-Thought Hijacking
por: Zhao, Jianli, et al.
Publicado: (2025) -
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
por: Schaeffer, Rylan, et al.
Publicado: (2024) -
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
por: Peng, Alwin, et al.
Publicado: (2024) -
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
por: Youstra, Jack, et al.
Publicado: (2025)