Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Kumar, Priyanshu, Lau, Elaine, Vijayakumar, Saranya, Trinh, Tu, Team, Scale Red, Chang, Elaine, Robinson, Vaughn, Hendryx, Sean, Zhou, Shuyan, Fredrikson, Matt, Yue, Summer, Wang, Zifan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Jailbreaking to Jailbreak
por: Kritz, Jeremy, et al.
Publicado: (2025)
por: Kritz, Jeremy, et al.
Publicado: (2025)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
por: Nath, Vaskar, et al.
Publicado: (2025)
por: Nath, Vaskar, et al.
Publicado: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025)
por: Gunjal, Anisha, et al.
Publicado: (2025)
A Recipe for Improved Certifiable Robustness
por: Hu, Kai, et al.
Publicado: (2023)
por: Hu, Kai, et al.
Publicado: (2023)
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
por: Knight, Christina Q., et al.
Publicado: (2025)
por: Knight, Christina Q., et al.
Publicado: (2025)
Reliable Weak-to-Strong Monitoring of LLM Agents
por: Kale, Neil, et al.
Publicado: (2025)
por: Kale, Neil, et al.
Publicado: (2025)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
por: Whitehead, Spencer, et al.
Publicado: (2024)
por: Whitehead, Spencer, et al.
Publicado: (2024)
VeriSplit: Secure and Practical Offloading of Machine Learning Inferences across IoT Devices
por: Zhang, Han, et al.
Publicado: (2024)
por: Zhang, Han, et al.
Publicado: (2024)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
por: Yuan, Youliang, et al.
Publicado: (2024)
por: Yuan, Youliang, et al.
Publicado: (2024)
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
por: Chen, Luoyu, et al.
Publicado: (2026)
por: Chen, Luoyu, et al.
Publicado: (2026)
Jailbroken Frontier Models Retain Their Capabilities
por: Zhu, Daniel, et al.
Publicado: (2026)
por: Zhu, Daniel, et al.
Publicado: (2026)
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
por: Lin, Weiran, et al.
Publicado: (2024)
por: Lin, Weiran, et al.
Publicado: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
por: Palumbo, Nils, et al.
Publicado: (2024)
por: Palumbo, Nils, et al.
Publicado: (2024)
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)
por: Priyanshu, Aman, et al.
Publicado: (2024)
por: Priyanshu, Aman, et al.
Publicado: (2024)
LipNeXt: Scaling up Lipschitz-based Certified Robustness to Billion-parameter Models
por: Hu, Kai, et al.
Publicado: (2026)
por: Hu, Kai, et al.
Publicado: (2026)
Who Trains the Trainer? Library Staff are OPAC Users, Too.
por: Coppola, Elaine
Publicado: (1983)
por: Coppola, Elaine
Publicado: (1983)
Revisiting the Superficial Alignment Hypothesis
por: Raghavendra, Mohit, et al.
Publicado: (2024)
por: Raghavendra, Mohit, et al.
Publicado: (2024)
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
por: Zhang, Hugh, et al.
Publicado: (2024)
por: Zhang, Hugh, et al.
Publicado: (2024)
ADAPTACIÓN AL ESPAÑOL DE LA ESCALA DE AMBIENTE INVALIDANTE INFANTIL
por: Scale Martín M. Puddington
Publicado: (2017)
por: Scale Martín M. Puddington
Publicado: (2017)
Does Refusal Training in LLMs Generalize to the Past Tense?
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Progress over Points: Reframing LM Benchmarks Around Scientific Objectives
por: Jin, Alwin, et al.
Publicado: (2025)
por: Jin, Alwin, et al.
Publicado: (2025)
Going 3D with Technology: An Overarching Approach for Language Teachers
por: Jason D. Hendryx
Publicado: (2016)
por: Jason D. Hendryx
Publicado: (2016)
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
por: Da, Jeff, et al.
Publicado: (2025)
por: Da, Jeff, et al.
Publicado: (2025)
Herz P1 Smart Scale More Than Just Weight Tracking
por: Herz P1 Smart Scale
Publicado: (2026)
por: Herz P1 Smart Scale
Publicado: (2026)
ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark
por: Nath, Vaskar, et al.
Publicado: (2025)
por: Nath, Vaskar, et al.
Publicado: (2025)
A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift
por: LeVine, Will, et al.
Publicado: (2023)
por: LeVine, Will, et al.
Publicado: (2023)
STARS (Secondary Training for Alaskan Rural Students): Communications. Draft Copy.
por: Griffin, Elaine, et al.
Publicado: (1977)
por: Griffin, Elaine, et al.
Publicado: (1977)
Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation
por: Heiding, Fred, et al.
Publicado: (2025)
por: Heiding, Fred, et al.
Publicado: (2025)
Sex Differences in the Impact of Obesity on Immunity
por: Saranya Vijayakumar, et al.
Publicado: (2026)
por: Saranya Vijayakumar, et al.
Publicado: (2026)
Refusal in LLMs is an Affine Function
por: Marshall, Thomas, et al.
Publicado: (2024)
por: Marshall, Thomas, et al.
Publicado: (2024)
Planning In Natural Language Improves LLM Search For Code Generation
por: Wang, Evan, et al.
Publicado: (2024)
por: Wang, Evan, et al.
Publicado: (2024)
The discrete empirical interpolation method in class identification and data summarization
por: Emily P. Hendryx Lyons
Publicado: (2024)
por: Emily P. Hendryx Lyons
Publicado: (2024)
The Functional Imperative: Structural Integrity and the Fallacy of Ornamental Soul
por: Red, Pax
Publicado: (2026)
por: Red, Pax
Publicado: (2026)
Revista de Prensa
por: Red Ires
Publicado: (2009)
por: Red Ires
Publicado: (2009)
Declaración europea por una nueva cultura del agua
por: Red Euwater
Publicado: (2005)
por: Red Euwater
Publicado: (2005)
When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models
por: Liu, Xiaoze, et al.
Publicado: (2025)
por: Liu, Xiaoze, et al.
Publicado: (2025)
A Mixture of Linear Corrections Generates Secure Code
por: Yu, Weichen, et al.
Publicado: (2025)
por: Yu, Weichen, et al.
Publicado: (2025)
LLMs Encode Harmfulness and Refusal Separately
por: Zhao, Jiachen, et al.
Publicado: (2025)
por: Zhao, Jiachen, et al.
Publicado: (2025)
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
por: Plaut, Benjamin, et al.
Publicado: (2024)
por: Plaut, Benjamin, et al.
Publicado: (2024)
Hindu Pluralism
por: Fisher, Elaine
Publicado: (2020)
por: Fisher, Elaine
Publicado: (2020)
Ejemplares similares
-
Jailbreaking to Jailbreak
por: Kritz, Jeremy, et al.
Publicado: (2025) -
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
por: Nath, Vaskar, et al.
Publicado: (2025) -
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025) -
A Recipe for Improved Certifiable Robustness
por: Hu, Kai, et al.
Publicado: (2023) -
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
por: Knight, Christina Q., et al.
Publicado: (2025)