STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
Fuente:
arXiv
Saved in:
| Main Authors: | Jung, MinJae, Lim, YongTaek, Kim, Chaeyun, Kim, Junghwan, Kim, Kihyun, Kim, Minwoo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
by: Lee, Jiyoung, et al.
Published: (2024)
by: Lee, Jiyoung, et al.
Published: (2024)
STAR: SocioTechnical Approach to Red Teaming Language Models
by: Weidinger, Laura, et al.
Published: (2024)
by: Weidinger, Laura, et al.
Published: (2024)
M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs
by: Ha, Junwoo, et al.
Published: (2025)
by: Ha, Junwoo, et al.
Published: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025)
by: Freenor, Michael, et al.
Published: (2025)
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024)
by: Jiang, Bojian, et al.
Published: (2024)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
by: Han, Vernon Toh Yan, et al.
Published: (2024)
by: Han, Vernon Toh Yan, et al.
Published: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
Towards Red Teaming in Multimodal and Multilingual Translation
by: Ropers, Christophe, et al.
Published: (2024)
by: Ropers, Christophe, et al.
Published: (2024)
Multi-lingual Multi-turn Automated Red Teaming for LLMs
by: Singhania, Abhishek, et al.
Published: (2025)
by: Singhania, Abhishek, et al.
Published: (2025)
Anecdoctoring: Automated Red-Teaming Across Language and Place
by: Cuevas, Alejandro, et al.
Published: (2025)
by: Cuevas, Alejandro, et al.
Published: (2025)
X-Teaming Evolutionary M2S: Automated Discovery of Multi-turn to Single-turn Jailbreak Templates
by: Kim, Hyunjun, et al.
Published: (2025)
by: Kim, Hyunjun, et al.
Published: (2025)
TroubleLLM: Align to Red Team Expert
by: Xu, Zhuoer, et al.
Published: (2024)
by: Xu, Zhuoer, et al.
Published: (2024)
The Effect of Gender Diversity on Scientific Team Impact: A Team Roles Perspective
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
by: Asad, Ali, et al.
Published: (2025)
by: Asad, Ali, et al.
Published: (2025)
Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains
by: Kim, Junghwan, et al.
Published: (2025)
by: Kim, Junghwan, et al.
Published: (2025)
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning
by: Jung, Min Jae, et al.
Published: (2024)
by: Jung, Min Jae, et al.
Published: (2024)
Model Fusion through Bayesian Optimization in Language Model Fine-Tuning
by: Jang, Chaeyun, et al.
Published: (2024)
by: Jang, Chaeyun, et al.
Published: (2024)
Training a General Purpose Automated Red Teaming Model
by: Padmakumar, Aishwarya, et al.
Published: (2026)
by: Padmakumar, Aishwarya, et al.
Published: (2026)
Fine-grained Gender Control in Machine Translation with Large Language Models
by: Lee, Minwoo, et al.
Published: (2024)
by: Lee, Minwoo, et al.
Published: (2024)
MAS-LitEval : Multi-Agent System for Literary Translation Quality Assessment
by: Kim, Junghwan, et al.
Published: (2025)
by: Kim, Junghwan, et al.
Published: (2025)
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
by: Kim, Mingyeong, et al.
Published: (2026)
by: Kim, Mingyeong, et al.
Published: (2026)
Exploring Coding Spot: Understanding Parametric Contributions to LLM Coding Performance
by: Kim, Dongjun, et al.
Published: (2024)
by: Kim, Dongjun, et al.
Published: (2024)
Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
by: Pala, Tej Deep, et al.
Published: (2024)
by: Pala, Tej Deep, et al.
Published: (2024)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
by: Horal, Artur, et al.
Published: (2025)
by: Horal, Artur, et al.
Published: (2025)
A Two-Step Approach for Data-Efficient French Pronunciation Learning
by: Lee, Hoyeon, et al.
Published: (2024)
by: Lee, Hoyeon, et al.
Published: (2024)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
by: Song, Jiwon, et al.
Published: (2025)
by: Song, Jiwon, et al.
Published: (2025)
Gradient-Based Language Model Red Teaming
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations
by: Raheja, Tarun, et al.
Published: (2024)
by: Raheja, Tarun, et al.
Published: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
TeamCMU at Touché: Adversarial Co-Evolution for Advertisement Integration and Detection in Conversational Search
by: Kim, To Eun, et al.
Published: (2025)
by: Kim, To Eun, et al.
Published: (2025)
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)
by: Kour, George, et al.
Published: (2024)
Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
by: Jang, Chaeyun, et al.
Published: (2025)
by: Jang, Chaeyun, et al.
Published: (2025)
PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
by: Deng, Wesley Hanwen, et al.
Published: (2025)
by: Deng, Wesley Hanwen, et al.
Published: (2025)
Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents
by: Mao, Yanxu, et al.
Published: (2026)
by: Mao, Yanxu, et al.
Published: (2026)
Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts
by: Ziakas, Christos, et al.
Published: (2025)
by: Ziakas, Christos, et al.
Published: (2025)
Similar Items
-
CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
by: Kim, Chaeyun, et al.
Published: (2026) -
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
by: Lee, Jiyoung, et al.
Published: (2024) -
STAR: SocioTechnical Approach to Red Teaming Language Models
by: Weidinger, Laura, et al.
Published: (2024) -
M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs
by: Ha, Junwoo, et al.
Published: (2025) -
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025)