Learning diverse attacks on large language models for robust red-teaming and safety tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Seanie, Kim, Minsu, Cherif, Lynn, Dobre, David, Lee, Juho, Hwang, Sung Ju, Kawaguchi, Kenji, Gidel, Gauthier, Bengio, Yoshua, Malkin, Nikolay, Jain, Moksh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Machine learning and information theory concepts towards an AI Mathematician
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
Amortizing intractable inference in large language models
by: Hu, Edward J., et al.
Published: (2023)
by: Hu, Edward J., et al.
Published: (2023)
Expected flow networks in stochastic environments and two-player zero-sum games
by: Jiralerspong, Marco, et al.
Published: (2023)
by: Jiralerspong, Marco, et al.
Published: (2023)
Set-based Meta-Interpolation for Few-Task Meta-Learning
by: Lee, Seanie, et al.
Published: (2022)
by: Lee, Seanie, et al.
Published: (2022)
PhyloGFN: Phylogenetic inference with generative flow networks
by: Zhou, Mingyang, et al.
Published: (2023)
by: Zhou, Mingyang, et al.
Published: (2023)
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
by: Bartoldson, Brian, et al.
Published: (2025)
by: Bartoldson, Brian, et al.
Published: (2025)
Amortizing intractable inference in diffusion models for vision, language, and control
by: Venkatraman, Siddarth, et al.
Published: (2024)
by: Venkatraman, Siddarth, et al.
Published: (2024)
Self-Supervised Dataset Distillation for Transfer Learning
by: Lee, Dong Bok, et al.
Published: (2023)
by: Lee, Dong Bok, et al.
Published: (2023)
Action abstractions for amortized sampling
by: Boussif, Oussama, et al.
Published: (2024)
by: Boussif, Oussama, et al.
Published: (2024)
Drug Discovery with Dynamic Goal-aware Fragments
by: Lee, Seul, et al.
Published: (2023)
by: Lee, Seul, et al.
Published: (2023)
On Generalization for Generative Flow Networks
by: Krichel, Anas, et al.
Published: (2024)
by: Krichel, Anas, et al.
Published: (2024)
In-Context Parametric Inference: Point or Distribution Estimators?
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Discrete Probabilistic Inference as Control in Multi-path Environments
by: Deleu, Tristan, et al.
Published: (2024)
by: Deleu, Tristan, et al.
Published: (2024)
Active Attacks: Red-teaming LLMs via Adaptive Environments
by: Yun, Taeyoung, et al.
Published: (2025)
by: Yun, Taeyoung, et al.
Published: (2025)
Outsourced diffusion sampling: Efficient posterior inference in latent spaces of generative models
by: Venkatraman, Siddarth, et al.
Published: (2025)
by: Venkatraman, Siddarth, et al.
Published: (2025)
Multi-Fidelity Active Learning with GFlowNets
by: Hernandez-Garcia, Alex, et al.
Published: (2023)
by: Hernandez-Garcia, Alex, et al.
Published: (2023)
Latent Veracity Inference for Identifying Errors in Stepwise Reasoning
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Adaptive teachers for amortized samplers
by: Kim, Minsu, et al.
Published: (2024)
by: Kim, Minsu, et al.
Published: (2024)
Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
by: Venkatraman, Siddarth, et al.
Published: (2025)
by: Venkatraman, Siddarth, et al.
Published: (2025)
Proof Flow: Preliminary Study on Generative Flow Network Language Model Tuning for Formal Reasoning
by: Ho, Matthew, et al.
Published: (2024)
by: Ho, Matthew, et al.
Published: (2024)
Discrete, compositional, and symbolic representations through attractor dynamics
by: Nam, Andrew, et al.
Published: (2023)
by: Nam, Andrew, et al.
Published: (2023)
Improved off-policy training of diffusion samplers
by: Sendera, Marcin, et al.
Published: (2024)
by: Sendera, Marcin, et al.
Published: (2024)
Solving Bayesian inverse problems with diffusion priors and off-policy RL
by: Scimeca, Luca, et al.
Published: (2025)
by: Scimeca, Luca, et al.
Published: (2025)
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
by: Shah, Vedant, et al.
Published: (2025)
by: Shah, Vedant, et al.
Published: (2025)
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
by: Schwinn, Leo, et al.
Published: (2024)
by: Schwinn, Leo, et al.
Published: (2024)
In-Context Learning Can Re-learn Forbidden Tasks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
Iterated Denoising Energy Matching for Sampling from Boltzmann Densities
by: Akhound-Sadegh, Tara, et al.
Published: (2024)
by: Akhound-Sadegh, Tara, et al.
Published: (2024)
Reliable Decision Making via Calibration Oriented Retrieval Augmented Generation
by: Jang, Chaeyun, et al.
Published: (2024)
by: Jang, Chaeyun, et al.
Published: (2024)
Learning Decision Trees as Amortized Structure Inference
by: Mahfoud, Mohammed, et al.
Published: (2025)
by: Mahfoud, Mohammed, et al.
Published: (2025)
Local Search GFlowNets
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
Can a Bayesian Oracle Prevent Harm from an Agent?
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
Simulation-free Schrödinger bridges via score and flow matching
by: Tong, Alexander, et al.
Published: (2023)
by: Tong, Alexander, et al.
Published: (2023)
Visual symbolic mechanisms: Emergent symbol processing in vision language models
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
Adaptive Inference-Time Scaling via Cyclic Diffusion Search
by: Lee, Gyubin, et al.
Published: (2025)
by: Lee, Gyubin, et al.
Published: (2025)
Delta-AI: Local objectives for amortized inference in sparse graphical models
by: Falet, Jean-Pierre, et al.
Published: (2023)
by: Falet, Jean-Pierre, et al.
Published: (2023)
Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning
by: Jiralerspong, Marco, et al.
Published: (2025)
by: Jiralerspong, Marco, et al.
Published: (2025)
Improving and generalizing flow-based generative models with minibatch optimal transport
by: Tong, Alexander, et al.
Published: (2023)
by: Tong, Alexander, et al.
Published: (2023)
Similar Items
-
Machine learning and information theory concepts towards an AI Mathematician
by: Bengio, Yoshua, et al.
Published: (2024) -
Amortizing intractable inference in large language models
by: Hu, Edward J., et al.
Published: (2023) -
Expected flow networks in stochastic environments and two-player zero-sum games
by: Jiralerspong, Marco, et al.
Published: (2023) -
Set-based Meta-Interpolation for Few-Task Meta-Learning
by: Lee, Seanie, et al.
Published: (2022) -
PhyloGFN: Phylogenetic inference with generative flow networks
by: Zhou, Mingyang, et al.
Published: (2023)