A StrongREJECT for Empty Jailbreaks
Fuente:
arXiv
Guardado en:
| Autores principales: | Souly, Alexandra, Lu, Qingyuan, Bowen, Dillon, Trinh, Tu, Hsieh, Elvis, Pandey, Sana, Abbeel, Pieter, Svegliato, Justin, Emmons, Scott, Watkins, Olivia, Toyer, Sam |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring and Addressing Reward Confusion in Offline Preference Learning
por: Chen, Xin, et al.
Publicado: (2024)
por: Chen, Xin, et al.
Publicado: (2024)
Learning to Model the World with Language
por: Lin, Jessy, et al.
Publicado: (2023)
por: Lin, Jessy, et al.
Publicado: (2023)
Active teacher selection for reward learning
por: Freedman, Rachel, et al.
Publicado: (2023)
por: Freedman, Rachel, et al.
Publicado: (2023)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
por: Seo, Younggyo, et al.
Publicado: (2024)
por: Seo, Younggyo, et al.
Publicado: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
por: Murphy, Brendan, et al.
Publicado: (2025)
por: Murphy, Brendan, et al.
Publicado: (2025)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
por: Subramani, Nishant, et al.
Publicado: (2025)
por: Subramani, Nishant, et al.
Publicado: (2025)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
por: CH-Wang, Sky, et al.
Publicado: (2025)
por: CH-Wang, Sky, et al.
Publicado: (2025)
Lightning Grasp: High Performance Procedural Grasp Synthesis with Contact Fields
por: Yin, Zhao-Heng, et al.
Publicado: (2025)
por: Yin, Zhao-Heng, et al.
Publicado: (2025)
Offline Imitation Learning Through Graph Search and Retrieval
por: Yin, Zhao-Heng, et al.
Publicado: (2024)
por: Yin, Zhao-Heng, et al.
Publicado: (2024)
A Stable Whitening Optimizer for Efficient Neural Network Training
por: Frans, Kevin, et al.
Publicado: (2025)
por: Frans, Kevin, et al.
Publicado: (2025)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
por: Lee, Vint, et al.
Publicado: (2023)
por: Lee, Vint, et al.
Publicado: (2023)
Reward-Conditioned Reinforcement Learning
por: Nauman, Michal, et al.
Publicado: (2026)
por: Nauman, Michal, et al.
Publicado: (2026)
What Really Matters in Matrix-Whitening Optimizers?
por: Frans, Kevin, et al.
Publicado: (2025)
por: Frans, Kevin, et al.
Publicado: (2025)
SEMDICE: Off-policy State Entropy Maximization via Stationary Distribution Correction Estimation
por: Lee, Jongmin, et al.
Publicado: (2025)
por: Lee, Jongmin, et al.
Publicado: (2025)
Welcome First--Books Later; The Service Center Branch, Richmond Public Library, December 1967 - June 1971.
por: Emmons, Karen
Publicado: (1971)
por: Emmons, Karen
Publicado: (1971)
She's Practiced What She Teaches
por: Emmons, Julia
Publicado: (1976)
por: Emmons, Julia
Publicado: (1976)
Empty Fields, Empty Promises
por: Ashwood, Loka, et al.
Publicado: (2023)
por: Ashwood, Loka, et al.
Publicado: (2023)
SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
por: Chen, Qianzhong, et al.
Publicado: (2025)
por: Chen, Qianzhong, et al.
Publicado: (2025)
EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations
por: Yu, Justin, et al.
Publicado: (2025)
por: Yu, Justin, et al.
Publicado: (2025)
Object-centric 3D Motion Field for Robot Learning from Human Videos
por: Yin, Zhao-Heng, et al.
Publicado: (2025)
por: Yin, Zhao-Heng, et al.
Publicado: (2025)
Cliqueformer: Model-Based Optimization with Structured Transformers
por: Kuba, Jakub Grudzien, et al.
Publicado: (2024)
por: Kuba, Jakub Grudzien, et al.
Publicado: (2024)
GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation
por: Elmaaroufi, Karim, et al.
Publicado: (2025)
por: Elmaaroufi, Karim, et al.
Publicado: (2025)
The Relationship Between Discomfort Intolerance And the Fear Of Self‐Injection And Testing In Patients With Diabetes Using Insulin: A Cross‐Sectional Study
por: Nilhan Töyer Şahin, et al.
Publicado: (2024)
por: Nilhan Töyer Şahin, et al.
Publicado: (2024)
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting
por: Liu, Fangchen, et al.
Publicado: (2024)
por: Liu, Fangchen, et al.
Publicado: (2024)
Learning a Diffusion Model Policy from Rewards via Q-Score Matching
por: Psenka, Michael, et al.
Publicado: (2023)
por: Psenka, Michael, et al.
Publicado: (2023)
World Model on Million-Length Video And Language With Blockwise RingAttention
por: Liu, Hao, et al.
Publicado: (2024)
por: Liu, Hao, et al.
Publicado: (2024)
Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
por: Frans, Kevin, et al.
Publicado: (2024)
por: Frans, Kevin, et al.
Publicado: (2024)
Interactive Task Planning with Language Models
por: Li, Boyi, et al.
Publicado: (2023)
por: Li, Boyi, et al.
Publicado: (2023)
Diffusion Guidance Is a Controllable Policy Improvement Operator
por: Frans, Kevin, et al.
Publicado: (2025)
por: Frans, Kevin, et al.
Publicado: (2025)
From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control
por: Shentu, Yide, et al.
Publicado: (2024)
por: Shentu, Yide, et al.
Publicado: (2024)
Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
por: Kim, Dongyoung, et al.
Publicado: (2023)
por: Kim, Dongyoung, et al.
Publicado: (2023)
Closing the Visual Sim-to-Real Gap with Object-Composable NeRFs
por: Mishra, Nikhil, et al.
Publicado: (2024)
por: Mishra, Nikhil, et al.
Publicado: (2024)
One Step Diffusion via Shortcut Models
por: Frans, Kevin, et al.
Publicado: (2024)
por: Frans, Kevin, et al.
Publicado: (2024)
ALMANACS: A Simulatability Benchmark for Language Model Explainability
por: Mills, Edmund, et al.
Publicado: (2023)
por: Mills, Edmund, et al.
Publicado: (2023)
Observation Interference in Partially Observable Assistance Games
por: Emmons, Scott, et al.
Publicado: (2024)
por: Emmons, Scott, et al.
Publicado: (2024)
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
por: Bailey, Luke, et al.
Publicado: (2023)
por: Bailey, Luke, et al.
Publicado: (2023)
Reasoning and Tools for Human-Level Forecasting
por: Hsieh, Elvis, et al.
Publicado: (2024)
por: Hsieh, Elvis, et al.
Publicado: (2024)
Empty Innovation
por: Hallonsten, Olof
Publicado: (2023)
por: Hallonsten, Olof
Publicado: (2023)
GaussGym: An open-source real-to-sim framework for learning locomotion from pixels
por: Escontrela, Alejandro, et al.
Publicado: (2025)
por: Escontrela, Alejandro, et al.
Publicado: (2025)
Engaging Conversation: Evaluating the Contribution of Library Instruction to the Quality of Student Research.
por: Emmons, Mark, et al.
Publicado: (2002)
por: Emmons, Mark, et al.
Publicado: (2002)
Ejemplares similares
-
Exploring and Addressing Reward Confusion in Offline Preference Learning
por: Chen, Xin, et al.
Publicado: (2024) -
Learning to Model the World with Language
por: Lin, Jessy, et al.
Publicado: (2023) -
Active teacher selection for reward learning
por: Freedman, Rachel, et al.
Publicado: (2023) -
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
por: Seo, Younggyo, et al.
Publicado: (2024) -
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
por: Murphy, Brendan, et al.
Publicado: (2025)