Break it, Imitate it, Fix it: Robustness by Generating Human-Like Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Sinha, Aradhana, Balashankar, Ananth, Beirami, Ahmad, Avrahami, Thi, Chen, Jilin, Beutel, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inducing Group Fairness in Prompt-Based Language Model Decisions
by: Atwood, James, et al.
Published: (2024)
by: Atwood, James, et al.
Published: (2024)
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
by: Wang, Xiangwen, et al.
Published: (2026)
by: Wang, Xiangwen, et al.
Published: (2026)
Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings
by: Akbarian, Fatemeh, et al.
Published: (2025)
by: Akbarian, Fatemeh, et al.
Published: (2025)
InfAlign: Inference-aware language model alignment
by: Balashankar, Ananth, et al.
Published: (2024)
by: Balashankar, Ananth, et al.
Published: (2024)
FieldFormer: Locality-Aware Transformers for Spatio-Temporal Modeling on Sparse Sensor Networks
by: Bhardwaj, Ankit, et al.
Published: (2025)
by: Bhardwaj, Ankit, et al.
Published: (2025)
BEFT: Bias-Efficient Fine-Tuning of Language Models in Low-Data Regimes
by: Huang, Baichuan, et al.
Published: (2025)
by: Huang, Baichuan, et al.
Published: (2025)
Multi-Group Fairness Evaluation via Conditional Value-at-Risk Testing
by: Paes, Lucas Monteiro, et al.
Published: (2023)
by: Paes, Lucas Monteiro, et al.
Published: (2023)
Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges
by: Feng, Chen, et al.
Published: (2026)
by: Feng, Chen, et al.
Published: (2026)
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment
by: Wu, Zhaofeng, et al.
Published: (2024)
by: Wu, Zhaofeng, et al.
Published: (2024)
Automated Adversarial Discovery for Safety Classifiers
by: Lal, Yash Kumar, et al.
Published: (2024)
by: Lal, Yash Kumar, et al.
Published: (2024)
Generalized People Diversity: Learning a Human Perception-Aligned Diversity Representation for People Images
by: Srinivasan, Hansa, et al.
Published: (2024)
by: Srinivasan, Hansa, et al.
Published: (2024)
Controlled Decoding from Language Models
by: Mudgal, Sidharth, et al.
Published: (2023)
by: Mudgal, Sidharth, et al.
Published: (2023)
Generalization and Robustness of the Tilted Empirical Risk
by: Aminian, Gholamali, et al.
Published: (2024)
by: Aminian, Gholamali, et al.
Published: (2024)
Improving Robustness via Tilted Exponential Layer: A Communication-Theoretic Perspective
by: Puranik, Bhagyashree, et al.
Published: (2023)
by: Puranik, Bhagyashree, et al.
Published: (2023)
Comprehensive Monitoring of Air Pollution Hotspots Using Sparse Sensor Networks
by: Bhardwaj, Ankit, et al.
Published: (2024)
by: Bhardwaj, Ankit, et al.
Published: (2024)
Conformal Signal Temporal Logic for Robust Reinforcement Learning Control: A Case Study
by: Beirami, Hani, et al.
Published: (2026)
by: Beirami, Hani, et al.
Published: (2026)
Imitative Membership Inference Attack
by: Du, Yuntao, et al.
Published: (2025)
by: Du, Yuntao, et al.
Published: (2025)
Adversarial Reinforcement Learning for Large Language Model Agent Safety
by: Wang, Zizhao, et al.
Published: (2025)
by: Wang, Zizhao, et al.
Published: (2025)
ILRR: Inference-Time Steering Method for Masked Diffusion Language Models
by: Avrahami, Eden, et al.
Published: (2026)
by: Avrahami, Eden, et al.
Published: (2026)
Click2Mask: Local Editing with Dynamic Mask Generation
by: Regev, Omer, et al.
Published: (2024)
by: Regev, Omer, et al.
Published: (2024)
Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields
by: Gordon, Ori, et al.
Published: (2023)
by: Gordon, Ori, et al.
Published: (2023)
On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression
by: Ge, Zichang, et al.
Published: (2025)
by: Ge, Zichang, et al.
Published: (2025)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
Generation from Noisy Examples
by: Raman, Ananth, et al.
Published: (2025)
by: Raman, Ananth, et al.
Published: (2025)
Robust Online Classification: From Estimation to Denoising
by: Wu, Changlong, et al.
Published: (2023)
by: Wu, Changlong, et al.
Published: (2023)
CoDe: Blockwise Control for Denoising Diffusion Models
by: Singh, Anuj, et al.
Published: (2025)
by: Singh, Anuj, et al.
Published: (2025)
On the Benefits of Inducing Local Lipschitzness for Robust Generative Adversarial Imitation Learning
by: Memarian, Farzan, et al.
Published: (2021)
by: Memarian, Farzan, et al.
Published: (2021)
Bayesian Robust Optimization for Imitation Learning
by: Brown, Daniel S., et al.
Published: (2020)
by: Brown, Daniel S., et al.
Published: (2020)
Robotic Imitation of Human Actions
by: Spisak, Josua, et al.
Published: (2024)
by: Spisak, Josua, et al.
Published: (2024)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023)
by: Cao, Bochuan, et al.
Published: (2023)
TimeMar: Multi-Scale Autoregressive Modeling for Unconditional Time Series Generation
by: Xu, Xiangyu, et al.
Published: (2026)
by: Xu, Xiangyu, et al.
Published: (2026)
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
by: Beutel, Alex, et al.
Published: (2024)
by: Beutel, Alex, et al.
Published: (2024)
Flow-Enabled Generalization to Human Demonstrations in Few-Shot Imitation Learning
by: Tang, Runze, et al.
Published: (2026)
by: Tang, Runze, et al.
Published: (2026)
Robust Imitation Learning for Automated Game Testing
by: Amadori, Pierluigi Vito, et al.
Published: (2024)
by: Amadori, Pierluigi Vito, et al.
Published: (2024)
FRAPPE: A Group Fairness Framework for Post-Processing Everything
by: Tifrea, Alexandru, et al.
Published: (2023)
by: Tifrea, Alexandru, et al.
Published: (2023)
Asymptotics of Language Model Alignment
by: Yang, Joy Qiping, et al.
Published: (2024)
by: Yang, Joy Qiping, et al.
Published: (2024)
Revisiting Synthetic Human Trajectories: Imitative Generation and Benchmarks Beyond Datasaurus
by: Deng, Bangchao, et al.
Published: (2024)
by: Deng, Bangchao, et al.
Published: (2024)
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
by: Liu, Xuefeng, et al.
Published: (2023)
by: Liu, Xuefeng, et al.
Published: (2023)
Story2Board: A Training-Free Approach for Expressive Storyboard Generation
by: Dinkevich, David, et al.
Published: (2025)
by: Dinkevich, David, et al.
Published: (2025)
Similar Items
-
Inducing Group Fairness in Prompt-Based Language Model Decisions
by: Atwood, James, et al.
Published: (2024) -
Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models
by: Wang, Xiangwen, et al.
Published: (2026) -
Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings
by: Akbarian, Fatemeh, et al.
Published: (2025) -
InfAlign: Inference-aware language model alignment
by: Balashankar, Ananth, et al.
Published: (2024) -
FieldFormer: Locality-Aware Transformers for Spatio-Temporal Modeling on Sparse Sensor Networks
by: Bhardwaj, Ankit, et al.
Published: (2025)