TroubleLLM: Align to Red Team Expert
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Zhuoer, Zhang, Jianping, Cui, Shiwen, Meng, Changhua, Wang, Weiqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agent Safety Alignment via Reinforcement Learning
von: Sha, Zeyang, et al.
Veröffentlicht: (2025)
von: Sha, Zeyang, et al.
Veröffentlicht: (2025)
SEM: Reinforcement Learning for Search-Efficient Large Language Models
von: Sha, Zeyang, et al.
Veröffentlicht: (2025)
von: Sha, Zeyang, et al.
Veröffentlicht: (2025)
Mirror-Consistency: Harnessing Inconsistency in Majority Voting
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
Tiny Refinements Elicit Resilience: Toward Efficient Prefix-Model Against LLM Red-Teaming
von: Liu, Jiaxu, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxu, et al.
Veröffentlicht: (2024)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)
Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations
von: Raheja, Tarun, et al.
Veröffentlicht: (2024)
von: Raheja, Tarun, et al.
Veröffentlicht: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2025)
Exploring Straightforward Conversational Red-Teaming
von: Kour, George, et al.
Veröffentlicht: (2024)
von: Kour, George, et al.
Veröffentlicht: (2024)
Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models
von: Liu, Yule, et al.
Veröffentlicht: (2025)
von: Liu, Yule, et al.
Veröffentlicht: (2025)
Red Teaming Large Language Models for Healthcare
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2025)
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2025)
FERRET: Framework for Expansion Reliant Red Teaming
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2026)
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2026)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
von: Lee, Hyomin, et al.
Veröffentlicht: (2026)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
von: Xu, Huiyu, et al.
Veröffentlicht: (2024)
von: Xu, Huiyu, et al.
Veröffentlicht: (2024)
Red Teaming Language Models for Processing Contradictory Dialogues
von: Wen, Xiaofei, et al.
Veröffentlicht: (2024)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2024)
LEC-KG: An LLM-Embedding Collaborative Framework for Domain-Specific Knowledge Graph Construction -- A Case Study on SDGs
von: Zeng, Yikai, et al.
Veröffentlicht: (2026)
von: Zeng, Yikai, et al.
Veröffentlicht: (2026)
Adaptive Instruction Composition for Automated LLM Red-Teaming
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
Reason-Align-Respond: Aligning LLM Reasoning with Knowledge Graphs for KGQA
von: Shen, Xiangqing, et al.
Veröffentlicht: (2025)
von: Shen, Xiangqing, et al.
Veröffentlicht: (2025)
Truth, Trust, and Trouble: Medical AI on the Edge
von: Azeez, Mohammad Anas, et al.
Veröffentlicht: (2025)
von: Azeez, Mohammad Anas, et al.
Veröffentlicht: (2025)
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
von: Lin, Tianwei, et al.
Veröffentlicht: (2024)
von: Lin, Tianwei, et al.
Veröffentlicht: (2024)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
Anecdoctoring: Automated Red-Teaming Across Language and Place
von: Cuevas, Alejandro, et al.
Veröffentlicht: (2025)
von: Cuevas, Alejandro, et al.
Veröffentlicht: (2025)
TeamLLM: A Human-Like Team-Oriented Collaboration Framework for Multi-Step Contextualized Tasks
von: Wang, Xiangyu, et al.
Veröffentlicht: (2026)
von: Wang, Xiangyu, et al.
Veröffentlicht: (2026)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
Red Teaming Visual Language Models
von: Li, Mukai, et al.
Veröffentlicht: (2024)
von: Li, Mukai, et al.
Veröffentlicht: (2024)
M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs
von: Ha, Junwoo, et al.
Veröffentlicht: (2025)
von: Ha, Junwoo, et al.
Veröffentlicht: (2025)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
von: Wang, Shaojie, et al.
Veröffentlicht: (2026)
Accuracy is Not Agreement: Expert-Aligned Evaluation of Crash Narrative Classification Models
von: Bhagat, Sudesh Ramesh, et al.
Veröffentlicht: (2025)
von: Bhagat, Sudesh Ramesh, et al.
Veröffentlicht: (2025)
Aligning Multilingual Reasoning with Verifiable Semantics from a High-Resource Expert Model
von: Faisal, Fahim, et al.
Veröffentlicht: (2025)
von: Faisal, Fahim, et al.
Veröffentlicht: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
von: Cheng, Gang, et al.
Veröffentlicht: (2025)
von: Cheng, Gang, et al.
Veröffentlicht: (2025)
When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models
von: Shamsi, Zafir, et al.
Veröffentlicht: (2026)
von: Shamsi, Zafir, et al.
Veröffentlicht: (2026)
Split and Merge: Aligning Position Biases in LLM-based Evaluators
von: Li, Zongjie, et al.
Veröffentlicht: (2023)
von: Li, Zongjie, et al.
Veröffentlicht: (2023)
CoE-Ops: Collaboration of LLM-based Experts for AIOps Question-Answering
von: Zhao, Jinkun, et al.
Veröffentlicht: (2025)
von: Zhao, Jinkun, et al.
Veröffentlicht: (2025)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
von: Yan, Siyu, et al.
Veröffentlicht: (2025)
von: Yan, Siyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Agent Safety Alignment via Reinforcement Learning
von: Sha, Zeyang, et al.
Veröffentlicht: (2025) -
SEM: Reinforcement Learning for Search-Efficient Large Language Models
von: Sha, Zeyang, et al.
Veröffentlicht: (2025) -
Mirror-Consistency: Harnessing Inconsistency in Majority Voting
von: Huang, Siyuan, et al.
Veröffentlicht: (2024) -
Tiny Refinements Elicit Resilience: Toward Efficient Prefix-Model Against LLM Red-Teaming
von: Liu, Jiaxu, et al.
Veröffentlicht: (2024) -
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
von: Dong, Jianshuo, et al.
Veröffentlicht: (2025)