Kaleidoscopic Teaming in Multi Agent Simulations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mehrabi, Ninareh, Kumarage, Tharindu, Chang, Kai-Wei, Galstyan, Aram, Gupta, Rahul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
FLIRT: Feedback Loop In-context Red Teaming
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2023)
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2023)
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
von: Liang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Liang, Jiacheng, et al.
Veröffentlicht: (2026)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
K-Edit: Language Model Editing with Contextual Knowledge Awareness
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge Graphs
von: Markowitz, Elan, et al.
Veröffentlicht: (2024)
von: Markowitz, Elan, et al.
Veröffentlicht: (2024)
Strategize Globally, Adapt Locally: A Multi-Turn Red Teaming Agent with Dual-Level Learning
von: Chen, Si, et al.
Veröffentlicht: (2025)
von: Chen, Si, et al.
Veröffentlicht: (2025)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2023)
Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2026)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2026)
SWAN: Semantic Watermarking with Abstract Meaning Representation
von: Ye, Ziping, et al.
Veröffentlicht: (2026)
von: Ye, Ziping, et al.
Veröffentlicht: (2026)
FERRET: Framework for Expansion Reliant Red Teaming
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2026)
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2026)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
On the steerability of large language models toward data-driven personas
von: Li, Junyi, et al.
Veröffentlicht: (2023)
von: Li, Junyi, et al.
Veröffentlicht: (2023)
Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Time
von: Li, Huihan, et al.
Veröffentlicht: (2025)
von: Li, Huihan, et al.
Veröffentlicht: (2025)
Co-Evolving Agents: Learning from Failures as Hard Negatives
von: Jung, Yeonsung, et al.
Veröffentlicht: (2025)
von: Jung, Yeonsung, et al.
Veröffentlicht: (2025)
Attribute Controlled Fine-tuning for Large Language Models: A Case Study on Detoxification
von: Meng, Tao, et al.
Veröffentlicht: (2024)
von: Meng, Tao, et al.
Veröffentlicht: (2024)
Adaptive Video Understanding Agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning
von: Jeoung, Sullam, et al.
Veröffentlicht: (2024)
von: Jeoung, Sullam, et al.
Veröffentlicht: (2024)
RedditESS: A Mental Health Social Support Interaction Dataset -- Understanding Effective Social Support to Refine AI-Driven Support Tools
von: Alghamdi, Zeyad, et al.
Veröffentlicht: (2025)
von: Alghamdi, Zeyad, et al.
Veröffentlicht: (2025)
IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents
von: Choi, Daewon, et al.
Veröffentlicht: (2026)
von: Choi, Daewon, et al.
Veröffentlicht: (2026)
Asking Back: Interaction-Layer Antidistillation Watermarks
von: Yang, Guang, et al.
Veröffentlicht: (2026)
von: Yang, Guang, et al.
Veröffentlicht: (2026)
Generative Kaleidoscopic Networks
von: Shrivastava, Harsh
Veröffentlicht: (2024)
von: Shrivastava, Harsh
Veröffentlicht: (2024)
A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2024)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
MultiLS: A Multi-task Lexical Simplification Framework
von: North, Kai, et al.
Veröffentlicht: (2024)
von: North, Kai, et al.
Veröffentlicht: (2024)
Graph Based Deep Reinforcement Learning Aided by Transformers for Multi-Agent Cooperation
von: Elrod, Michael, et al.
Veröffentlicht: (2025)
von: Elrod, Michael, et al.
Veröffentlicht: (2025)
MDTeamGPT: A Self-Evolving LLM-based Multi-Agent Framework for Multi-Disciplinary Team Medical Consultation
von: Chen, Kai, et al.
Veröffentlicht: (2025)
von: Chen, Kai, et al.
Veröffentlicht: (2025)
Kaleidoscope: Learnable Masks for Heterogeneous Multi-agent Reinforcement Learning
von: Li, Xinran, et al.
Veröffentlicht: (2024)
von: Li, Xinran, et al.
Veröffentlicht: (2024)
SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
Ontology-Aware RAG for Improved Question-Answering in Cybersecurity Education
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2024)
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2024)
ExComm: Exploration-Stage Communication for Error-Resilient Agentic Test-Time Scaling
von: Song, Woomin, et al.
Veröffentlicht: (2026)
von: Song, Woomin, et al.
Veröffentlicht: (2026)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
Prompt Perturbation Consistency Learning for Robust Language Models
von: Qiang, Yao, et al.
Veröffentlicht: (2024)
von: Qiang, Yao, et al.
Veröffentlicht: (2024)
DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Multi-Agent Adversarial Reinforcement Learning
von: Fernando, Tharindu, et al.
Veröffentlicht: (2025)
von: Fernando, Tharindu, et al.
Veröffentlicht: (2025)
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification
von: North, Kai, et al.
Veröffentlicht: (2022)
von: North, Kai, et al.
Veröffentlicht: (2022)
Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art
von: Issak, Alayt, et al.
Veröffentlicht: (2025)
von: Issak, Alayt, et al.
Veröffentlicht: (2025)
Claw AI Lab: An Autonomous Multi-Agent Research Team
von: Wu, Fan, et al.
Veröffentlicht: (2026)
von: Wu, Fan, et al.
Veröffentlicht: (2026)
Multi-Agent Teams Hold Experts Back
von: Pappu, Aneesh, et al.
Veröffentlicht: (2026)
von: Pappu, Aneesh, et al.
Veröffentlicht: (2026)
CyberBOT: Towards Reliable Cybersecurity Education via Ontology-Grounded Retrieval Augmented Generation
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2025)
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2025)
Before Humans Join the Team: Diagnosing Coordination Failures in Healthcare Robot Team Simulation
von: Bai, Yuanchen, et al.
Veröffentlicht: (2025)
von: Bai, Yuanchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025) -
FLIRT: Feedback Loop In-context Red Teaming
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2023) -
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
von: Liang, Jiacheng, et al.
Veröffentlicht: (2026) -
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024) -
K-Edit: Language Model Editing with Contextual Knowledge Awareness
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)