A StrongREJECT for Empty Jailbreaks
Fuente:
arXiv
Saved in:
| Main Authors: | Souly, Alexandra, Lu, Qingyuan, Bowen, Dillon, Trinh, Tu, Hsieh, Elvis, Pandey, Sana, Abbeel, Pieter, Svegliato, Justin, Emmons, Scott, Watkins, Olivia, Toyer, Sam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring and Addressing Reward Confusion in Offline Preference Learning
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Learning to Model the World with Language
by: Lin, Jessy, et al.
Published: (2023)
by: Lin, Jessy, et al.
Published: (2023)
Active teacher selection for reward learning
by: Freedman, Rachel, et al.
Published: (2023)
by: Freedman, Rachel, et al.
Published: (2023)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
by: CH-Wang, Sky, et al.
Published: (2025)
by: CH-Wang, Sky, et al.
Published: (2025)
Lightning Grasp: High Performance Procedural Grasp Synthesis with Contact Fields
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
Offline Imitation Learning Through Graph Search and Retrieval
by: Yin, Zhao-Heng, et al.
Published: (2024)
by: Yin, Zhao-Heng, et al.
Published: (2024)
A Stable Whitening Optimizer for Efficient Neural Network Training
by: Frans, Kevin, et al.
Published: (2025)
by: Frans, Kevin, et al.
Published: (2025)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
by: Lee, Vint, et al.
Published: (2023)
by: Lee, Vint, et al.
Published: (2023)
Reward-Conditioned Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2026)
by: Nauman, Michal, et al.
Published: (2026)
What Really Matters in Matrix-Whitening Optimizers?
by: Frans, Kevin, et al.
Published: (2025)
by: Frans, Kevin, et al.
Published: (2025)
SEMDICE: Off-policy State Entropy Maximization via Stationary Distribution Correction Estimation
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Welcome First--Books Later; The Service Center Branch, Richmond Public Library, December 1967 - June 1971.
by: Emmons, Karen
Published: (1971)
by: Emmons, Karen
Published: (1971)
She's Practiced What She Teaches
by: Emmons, Julia
Published: (1976)
by: Emmons, Julia
Published: (1976)
Empty Fields, Empty Promises
by: Ashwood, Loka, et al.
Published: (2023)
by: Ashwood, Loka, et al.
Published: (2023)
SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
by: Chen, Qianzhong, et al.
Published: (2025)
by: Chen, Qianzhong, et al.
Published: (2025)
EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations
by: Yu, Justin, et al.
Published: (2025)
by: Yu, Justin, et al.
Published: (2025)
Object-centric 3D Motion Field for Robot Learning from Human Videos
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
Cliqueformer: Model-Based Optimization with Structured Transformers
by: Kuba, Jakub Grudzien, et al.
Published: (2024)
by: Kuba, Jakub Grudzien, et al.
Published: (2024)
GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation
by: Elmaaroufi, Karim, et al.
Published: (2025)
by: Elmaaroufi, Karim, et al.
Published: (2025)
The Relationship Between Discomfort Intolerance And the Fear Of Self‐Injection And Testing In Patients With Diabetes Using Insulin: A Cross‐Sectional Study
by: Nilhan Töyer Şahin, et al.
Published: (2024)
by: Nilhan Töyer Şahin, et al.
Published: (2024)
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting
by: Liu, Fangchen, et al.
Published: (2024)
by: Liu, Fangchen, et al.
Published: (2024)
Learning a Diffusion Model Policy from Rewards via Q-Score Matching
by: Psenka, Michael, et al.
Published: (2023)
by: Psenka, Michael, et al.
Published: (2023)
World Model on Million-Length Video And Language With Blockwise RingAttention
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
by: Frans, Kevin, et al.
Published: (2024)
by: Frans, Kevin, et al.
Published: (2024)
Interactive Task Planning with Language Models
by: Li, Boyi, et al.
Published: (2023)
by: Li, Boyi, et al.
Published: (2023)
Diffusion Guidance Is a Controllable Policy Improvement Operator
by: Frans, Kevin, et al.
Published: (2025)
by: Frans, Kevin, et al.
Published: (2025)
From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control
by: Shentu, Yide, et al.
Published: (2024)
by: Shentu, Yide, et al.
Published: (2024)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
by: Kim, Dongyoung, et al.
Published: (2023)
by: Kim, Dongyoung, et al.
Published: (2023)
Closing the Visual Sim-to-Real Gap with Object-Composable NeRFs
by: Mishra, Nikhil, et al.
Published: (2024)
by: Mishra, Nikhil, et al.
Published: (2024)
One Step Diffusion via Shortcut Models
by: Frans, Kevin, et al.
Published: (2024)
by: Frans, Kevin, et al.
Published: (2024)
ALMANACS: A Simulatability Benchmark for Language Model Explainability
by: Mills, Edmund, et al.
Published: (2023)
by: Mills, Edmund, et al.
Published: (2023)
Observation Interference in Partially Observable Assistance Games
by: Emmons, Scott, et al.
Published: (2024)
by: Emmons, Scott, et al.
Published: (2024)
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
by: Bailey, Luke, et al.
Published: (2023)
by: Bailey, Luke, et al.
Published: (2023)
Reasoning and Tools for Human-Level Forecasting
by: Hsieh, Elvis, et al.
Published: (2024)
by: Hsieh, Elvis, et al.
Published: (2024)
Empty Innovation
by: Hallonsten, Olof
Published: (2023)
by: Hallonsten, Olof
Published: (2023)
GaussGym: An open-source real-to-sim framework for learning locomotion from pixels
by: Escontrela, Alejandro, et al.
Published: (2025)
by: Escontrela, Alejandro, et al.
Published: (2025)
Engaging Conversation: Evaluating the Contribution of Library Instruction to the Quality of Student Research.
by: Emmons, Mark, et al.
Published: (2002)
by: Emmons, Mark, et al.
Published: (2002)
Similar Items
-
Exploring and Addressing Reward Confusion in Offline Preference Learning
by: Chen, Xin, et al.
Published: (2024) -
Learning to Model the World with Language
by: Lin, Jessy, et al.
Published: (2023) -
Active teacher selection for reward learning
by: Freedman, Rachel, et al.
Published: (2023) -
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024) -
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)