Gradient-Based Language Model Red Teaming
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wichers, Nevan, Denison, Carson, Beirami, Ahmad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
von: Shu, Lei, et al.
Veröffentlicht: (2023)
von: Shu, Lei, et al.
Veröffentlicht: (2023)
Red-Teaming for Inducing Societal Bias in Large Language Models
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024)
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024)
Red Teaming Visual Language Models
von: Li, Mukai, et al.
Veröffentlicht: (2024)
von: Li, Mukai, et al.
Veröffentlicht: (2024)
Red Teaming Large Language Models for Healthcare
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2025)
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2025)
RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models
von: Dang, Quy-Anh, et al.
Veröffentlicht: (2026)
von: Dang, Quy-Anh, et al.
Veröffentlicht: (2026)
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2024)
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2024)
Red Teaming Language Models for Processing Contradictory Dialogues
von: Wen, Xiaofei, et al.
Veröffentlicht: (2024)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2024)
RRTL: Red Teaming Reasoning Large Language Models in Tool Learning
von: Liu, Yifei, et al.
Veröffentlicht: (2025)
von: Liu, Yifei, et al.
Veröffentlicht: (2025)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
von: Wei, Zhang, et al.
Veröffentlicht: (2025)
von: Wei, Zhang, et al.
Veröffentlicht: (2025)
Resource Consumption Red-Teaming for Large Vision-Language Models
von: Gao, Haoran, et al.
Veröffentlicht: (2025)
von: Gao, Haoran, et al.
Veröffentlicht: (2025)
Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models
von: Van Doren, Madison, et al.
Veröffentlicht: (2025)
von: Van Doren, Madison, et al.
Veröffentlicht: (2025)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
STAR: SocioTechnical Approach to Red Teaming Language Models
von: Weidinger, Laura, et al.
Veröffentlicht: (2024)
von: Weidinger, Laura, et al.
Veröffentlicht: (2024)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
von: Xu, Huiyu, et al.
Veröffentlicht: (2024)
von: Xu, Huiyu, et al.
Veröffentlicht: (2024)
Anecdoctoring: Automated Red-Teaming Across Language and Place
von: Cuevas, Alejandro, et al.
Veröffentlicht: (2025)
von: Cuevas, Alejandro, et al.
Veröffentlicht: (2025)
ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts
von: Hardy, Amelia F., et al.
Veröffentlicht: (2024)
von: Hardy, Amelia F., et al.
Veröffentlicht: (2024)
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models
von: Purpura, Alberto, et al.
Veröffentlicht: (2025)
von: Purpura, Alberto, et al.
Veröffentlicht: (2025)
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis
von: Wu, Xiaorui, et al.
Veröffentlicht: (2025)
von: Wu, Xiaorui, et al.
Veröffentlicht: (2025)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
von: Hu, Yujia, et al.
Veröffentlicht: (2025)
von: Hu, Yujia, et al.
Veröffentlicht: (2025)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2024)
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2024)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models
von: Shamsi, Zafir, et al.
Veröffentlicht: (2026)
von: Shamsi, Zafir, et al.
Veröffentlicht: (2026)
Automated Progressive Red Teaming
von: Jiang, Bojian, et al.
Veröffentlicht: (2024)
von: Jiang, Bojian, et al.
Veröffentlicht: (2024)
Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
von: Pala, Tej Deep, et al.
Veröffentlicht: (2024)
von: Pala, Tej Deep, et al.
Veröffentlicht: (2024)
Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
von: Lin, Lizhi, et al.
Veröffentlicht: (2024)
von: Lin, Lizhi, et al.
Veröffentlicht: (2024)
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
von: Verma, Apurv, et al.
Veröffentlicht: (2024)
von: Verma, Apurv, et al.
Veröffentlicht: (2024)
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
von: Asad, Ali, et al.
Veröffentlicht: (2025)
von: Asad, Ali, et al.
Veröffentlicht: (2025)
Exploring Straightforward Conversational Red-Teaming
von: Kour, George, et al.
Veröffentlicht: (2024)
von: Kour, George, et al.
Veröffentlicht: (2024)
Training a General Purpose Automated Red Teaming Model
von: Padmakumar, Aishwarya, et al.
Veröffentlicht: (2026)
von: Padmakumar, Aishwarya, et al.
Veröffentlicht: (2026)
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
von: Jung, MinJae, et al.
Veröffentlicht: (2026)
von: Jung, MinJae, et al.
Veröffentlicht: (2026)
For Those Who May Find Themselves on the Red Team
von: Shoemaker, Tyler
Veröffentlicht: (2025)
von: Shoemaker, Tyler
Veröffentlicht: (2025)
Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks
von: Buszydlik, Aleksander, et al.
Veröffentlicht: (2023)
von: Buszydlik, Aleksander, et al.
Veröffentlicht: (2023)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
von: Tedeschi, Simone, et al.
Veröffentlicht: (2024)
von: Tedeschi, Simone, et al.
Veröffentlicht: (2024)
Visualizing Neural Network Imagination
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
von: Diao, Muxi, et al.
Veröffentlicht: (2025)
von: Diao, Muxi, et al.
Veröffentlicht: (2025)
Multi-lingual Multi-turn Automated Red Teaming for LLMs
von: Singhania, Abhishek, et al.
Veröffentlicht: (2025)
von: Singhania, Abhishek, et al.
Veröffentlicht: (2025)
TroubleLLM: Align to Red Team Expert
von: Xu, Zhuoer, et al.
Veröffentlicht: (2024)
von: Xu, Zhuoer, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
von: Cao, Meng, et al.
Veröffentlicht: (2024) -
Fusion-Eval: Integrating Assistant Evaluators with LLMs
von: Shu, Lei, et al.
Veröffentlicht: (2023) -
Red-Teaming for Inducing Societal Bias in Large Language Models
von: Luo, Chu Fei, et al.
Veröffentlicht: (2024) -
Red Teaming Visual Language Models
von: Li, Mukai, et al.
Veröffentlicht: (2024) -
Red Teaming Large Language Models for Healthcare
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2025)