Automatic LLM Red Teaming
Fuente:
arXiv
Guardado en:
| Autores principales: | Belaire, Roman, Sinha, Arunesh, Varakantham, Pradeep |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On Minimizing Adversarial Counterfactual Error in Adversarial RL
por: Belaire, Roman, et al.
Publicado: (2024)
por: Belaire, Roman, et al.
Publicado: (2024)
Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning
por: Lu, Yuxiao, et al.
Publicado: (2023)
por: Lu, Yuxiao, et al.
Publicado: (2023)
Regret-Based Defense in Adversarial Reinforcement Learning
por: Belaire, Roman, et al.
Publicado: (2023)
por: Belaire, Roman, et al.
Publicado: (2023)
On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression
por: Ge, Zichang, et al.
Publicado: (2025)
por: Ge, Zichang, et al.
Publicado: (2025)
Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
por: Lu, Yuxiao, et al.
Publicado: (2024)
por: Lu, Yuxiao, et al.
Publicado: (2024)
Enhancing the Hierarchical Environment Design via Generative Trajectory Modeling
por: Li, Dexun, et al.
Publicado: (2023)
por: Li, Dexun, et al.
Publicado: (2023)
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
por: Hoang, Huy, et al.
Publicado: (2024)
por: Hoang, Huy, et al.
Publicado: (2024)
Offline Safe Reinforcement Learning Using Trajectory Classification
por: Gong, Ze, et al.
Publicado: (2024)
por: Gong, Ze, et al.
Publicado: (2024)
Imitate the Good and Avoid the Bad: An Incremental Approach to Safe Reinforcement Learning
por: Hoang, Huy, et al.
Publicado: (2023)
por: Hoang, Huy, et al.
Publicado: (2023)
SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning
por: Hoang, Huy, et al.
Publicado: (2024)
por: Hoang, Huy, et al.
Publicado: (2024)
Imitating Cost-Constrained Behaviors in Reinforcement Learning
por: Shao, Qian, et al.
Publicado: (2024)
por: Shao, Qian, et al.
Publicado: (2024)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
por: Hoang, Huy, et al.
Publicado: (2025)
por: Hoang, Huy, et al.
Publicado: (2025)
Induced Numerical Instability: Hidden Costs in Multimodal Large Language Models
por: Wong, Wai Tuck, et al.
Publicado: (2026)
por: Wong, Wai Tuck, et al.
Publicado: (2026)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
por: Sinha, Anusha, et al.
Publicado: (2025)
por: Sinha, Anusha, et al.
Publicado: (2025)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
por: Jain, Gauri, et al.
Publicado: (2024)
por: Jain, Gauri, et al.
Publicado: (2024)
Towards Neural Network based Cognitive Models of Dynamic Decision-Making by Humans
por: Chen, Changyu, et al.
Publicado: (2024)
por: Chen, Changyu, et al.
Publicado: (2024)
Geometric Red-Teaming for Robotic Manipulation
por: Goel, Divyam, et al.
Publicado: (2025)
por: Goel, Divyam, et al.
Publicado: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
por: Zymet, Jesse, et al.
Publicado: (2026)
por: Zymet, Jesse, et al.
Publicado: (2026)
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
por: Pavlova, Maya, et al.
Publicado: (2024)
por: Pavlova, Maya, et al.
Publicado: (2024)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
por: Panfilov, Alexander, et al.
Publicado: (2025)
por: Panfilov, Alexander, et al.
Publicado: (2025)
Red-Team Multi-Agent Reinforcement Learning for Emergency Braking Scenario
por: Chen, Yinsong, et al.
Publicado: (2025)
por: Chen, Yinsong, et al.
Publicado: (2025)
Embodied Red Teaming for Auditing Robotic Foundation Models
por: Karnik, Sathwik, et al.
Publicado: (2024)
por: Karnik, Sathwik, et al.
Publicado: (2024)
Bootstrapping Language Models with DPO Implicit Rewards
por: Chen, Changyu, et al.
Publicado: (2024)
por: Chen, Changyu, et al.
Publicado: (2024)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
por: Zhang, Jiawei, et al.
Publicado: (2025)
por: Zhang, Jiawei, et al.
Publicado: (2025)
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
por: Chen, Zhaorun, et al.
Publicado: (2025)
por: Chen, Zhaorun, et al.
Publicado: (2025)
Red-Teaming Segment Anything Model
por: Jankowski, Krzysztof, et al.
Publicado: (2024)
por: Jankowski, Krzysztof, et al.
Publicado: (2024)
Automatic Configuration of LLM Post-Training Pipelines
por: Chwa, Channe, et al.
Publicado: (2026)
por: Chwa, Channe, et al.
Publicado: (2026)
Text2Insight: Transform natural language text into insights seamlessly using multi-model architecture
por: Sain, Pradeep
Publicado: (2024)
por: Sain, Pradeep
Publicado: (2024)
Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
por: Guo, Ruohao, et al.
Publicado: (2025)
por: Guo, Ruohao, et al.
Publicado: (2025)
Leveraging Reinforcement Learning in Red Teaming for Advanced Ransomware Attack Simulations
por: Wang, Cheng, et al.
Publicado: (2024)
por: Wang, Cheng, et al.
Publicado: (2024)
Automatic Causal Fairness Analysis with LLM-Generated Reporting
por: Berarducci, Alessia, et al.
Publicado: (2026)
por: Berarducci, Alessia, et al.
Publicado: (2026)
Predictive Red Teaming: Breaking Policies Without Breaking Robots
por: Majumdar, Anirudha, et al.
Publicado: (2025)
por: Majumdar, Anirudha, et al.
Publicado: (2025)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
por: Ding, Jiale, et al.
Publicado: (2025)
por: Ding, Jiale, et al.
Publicado: (2025)
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
por: Liang, Jiacheng, et al.
Publicado: (2026)
por: Liang, Jiacheng, et al.
Publicado: (2026)
Team, Then Trim: An Assembly-Line LLM Framework for High-Quality Tabular Data Generation
por: Zhang, Congjing, et al.
Publicado: (2026)
por: Zhang, Congjing, et al.
Publicado: (2026)
Automatic Demonstration Selection for LLM-based Tabular Data Classification
por: Han, Shuchu, et al.
Publicado: (2025)
por: Han, Shuchu, et al.
Publicado: (2025)
Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing
por: Bhatt, Manish, et al.
Publicado: (2025)
por: Bhatt, Manish, et al.
Publicado: (2025)
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
por: Cheng, Gang, et al.
Publicado: (2025)
por: Cheng, Gang, et al.
Publicado: (2025)
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
por: Beutel, Alex, et al.
Publicado: (2024)
por: Beutel, Alex, et al.
Publicado: (2024)
Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI
por: Rawat, Ambrish, et al.
Publicado: (2024)
por: Rawat, Ambrish, et al.
Publicado: (2024)
Ejemplares similares
-
On Minimizing Adversarial Counterfactual Error in Adversarial RL
por: Belaire, Roman, et al.
Publicado: (2024) -
Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning
por: Lu, Yuxiao, et al.
Publicado: (2023) -
Regret-Based Defense in Adversarial Reinforcement Learning
por: Belaire, Roman, et al.
Publicado: (2023) -
On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression
por: Ge, Zichang, et al.
Publicado: (2025) -
Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
por: Lu, Yuxiao, et al.
Publicado: (2024)