RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Jiale, Zheng, Xiang, Wu, Yutao, Wang, Cong, Lee, Wei-Bin, Pan, Ling, Ma, Xingjun, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
by: Wang, Zilong, et al.
Published: (2025)
by: Wang, Zilong, et al.
Published: (2025)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming
by: Zheng, Xiang, et al.
Published: (2025)
by: Zheng, Xiang, et al.
Published: (2025)
Towards Red Teaming in Multimodal and Multilingual Translation
by: Ropers, Christophe, et al.
Published: (2024)
by: Ropers, Christophe, et al.
Published: (2024)
Anecdoctoring: Automated Red-Teaming Across Language and Place
by: Cuevas, Alejandro, et al.
Published: (2025)
by: Cuevas, Alejandro, et al.
Published: (2025)
STAR: SocioTechnical Approach to Red Teaming Language Models
by: Weidinger, Laura, et al.
Published: (2024)
by: Weidinger, Laura, et al.
Published: (2024)
Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories
by: Bilalpur, Maneesh, et al.
Published: (2025)
by: Bilalpur, Maneesh, et al.
Published: (2025)
TATA: Stance Detection via Topic-Agnostic and Topic-Aware Embeddings
by: Hanley, Hans W. A., et al.
Published: (2023)
by: Hanley, Hans W. A., et al.
Published: (2023)
ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
by: Tedeschi, Simone, et al.
Published: (2024)
by: Tedeschi, Simone, et al.
Published: (2024)
PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
by: Deng, Wesley Hanwen, et al.
Published: (2026)
by: Deng, Wesley Hanwen, et al.
Published: (2026)
Effective Automation to Support the Human Infrastructure in AI Red Teaming
by: Zhang, Alice Qian, et al.
Published: (2025)
by: Zhang, Alice Qian, et al.
Published: (2025)
How Proposal Novelty, Topical Diversity, and Theory-Practice Balance Shape Scholarly Outcomes in Funded Education Research
by: Gao, Yunfeng, et al.
Published: (2026)
by: Gao, Yunfeng, et al.
Published: (2026)
Promising Topics for U.S.-China Dialogues on AI Risks and Governance
by: Siddiqui, Saad, et al.
Published: (2025)
by: Siddiqui, Saad, et al.
Published: (2025)
Red-Teaming for Generative AI: Silver Bullet or Security Theater?
by: Feffer, Michael, et al.
Published: (2024)
by: Feffer, Michael, et al.
Published: (2024)
Assessing Risks of Large Language Models in Mental Health Support: A Framework for Automated Clinical AI Red Teaming
by: Steenstra, Ian, et al.
Published: (2026)
by: Steenstra, Ian, et al.
Published: (2026)
A Comparative Evaluation of Structural Topic Models and BERTopic for Short, Open-Ended Survey Responses
by: Jiang, Yan, et al.
Published: (2026)
by: Jiang, Yan, et al.
Published: (2026)
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
by: Dimino, Fabrizio, et al.
Published: (2026)
by: Dimino, Fabrizio, et al.
Published: (2026)
CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
An Image is Worth $K$ Topics: A Visual Structural Topic Model with Pretrained Image Embeddings
by: Piqueras, Matías, et al.
Published: (2025)
by: Piqueras, Matías, et al.
Published: (2025)
Google Topics as a way out of the cookie dilemma?
by: Köppel, Marius, et al.
Published: (2024)
by: Köppel, Marius, et al.
Published: (2024)
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
by: Rastogi, Charvi, et al.
Published: (2026)
by: Rastogi, Charvi, et al.
Published: (2026)
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
by: Movva, Rajiv, et al.
Published: (2023)
by: Movva, Rajiv, et al.
Published: (2023)
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation
by: Quaye, Jessica, et al.
Published: (2024)
by: Quaye, Jessica, et al.
Published: (2024)
The Human Factor in AI Red Teaming: Perspectives from Social and Collaborative Computing
by: Zhang, Alice Qian, et al.
Published: (2024)
by: Zhang, Alice Qian, et al.
Published: (2024)
Red Teaming AI Policy: A Taxonomy of Avoision and the EU AI Act
by: Yew, Rui-Jie, et al.
Published: (2025)
by: Yew, Rui-Jie, et al.
Published: (2025)
Fifteen Years of Learning Analytics Research: Topics, Trends, and Challenges
by: Švábenský, Valdemar, et al.
Published: (2026)
by: Švábenský, Valdemar, et al.
Published: (2026)
Hiden Topics in Robotic Process Automation -- an Approach based on AI
by: Prucha, Petr, et al.
Published: (2024)
by: Prucha, Petr, et al.
Published: (2024)
RedTeamLLM: an Agentic AI framework for offensive security
by: Challita, Brian, et al.
Published: (2025)
by: Challita, Brian, et al.
Published: (2025)
Cross-Platform Short-Video Diplomacy: Topic and Sentiment Analysis of China-US Relations on Douyin and TikTok
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
OpenAI's Approach to External Red Teaming for AI Models and Systems
by: Ahmad, Lama, et al.
Published: (2025)
by: Ahmad, Lama, et al.
Published: (2025)
LLMs Homogenize Values in Constructive Arguments on Value-Laden Topics
by: Shahid, Farhana, et al.
Published: (2025)
by: Shahid, Farhana, et al.
Published: (2025)
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
Understanding the Progression of Educational Topics via Semantic Matching
by: Alkhidir, Tamador, et al.
Published: (2024)
by: Alkhidir, Tamador, et al.
Published: (2024)
Ask What Your Country Can Do For You: Towards a Public Red Teaming Model
by: Kennedy, Wm. Matthew, et al.
Published: (2025)
by: Kennedy, Wm. Matthew, et al.
Published: (2025)
Landscape of Generative AI in Global News: Topics, Sentiments, and Spatiotemporal Analysis
by: Xian, Lu, et al.
Published: (2024)
by: Xian, Lu, et al.
Published: (2024)
Improving the TENOR of Labeling: Re-evaluating Topic Models for Content Analysis
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming
by: Tong, Baoshun, et al.
Published: (2026)
by: Tong, Baoshun, et al.
Published: (2026)
On the Affinity, Rationality, and Diversity of Hierarchical Topic Modeling
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
Automating Thematic Analysis: How LLMs Analyse Controversial Topics
by: Khan, Awais Hameed, et al.
Published: (2024)
by: Khan, Awais Hameed, et al.
Published: (2024)
Similar Items
-
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
by: Wang, Zilong, et al.
Published: (2025) -
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025) -
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming
by: Zheng, Xiang, et al.
Published: (2025) -
Towards Red Teaming in Multimodal and Multilingual Translation
by: Ropers, Christophe, et al.
Published: (2024) -
Anecdoctoring: Automated Red-Teaming Across Language and Place
by: Cuevas, Alejandro, et al.
Published: (2025)