Combating Adversarial Attacks with Multi-Agent Debate
Fuente:
arXiv
Saved in:
| Main Authors: | Chern, Steffi, Fan, Zhen, Liu, Andy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024)
by: Wang, Binjie, et al.
Published: (2024)
MultiAgent Collaboration Attack: Investigating Adversarial Attacks in Large Language Model Collaborations via Debate
by: Amayuelas, Alfonso, et al.
Published: (2024)
by: Amayuelas, Alfonso, et al.
Published: (2024)
BeHonest: Benchmarking Honesty in Large Language Models
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
GroupDebate: Enhancing the Efficiency of Multi-Agent Debate Using Group Discussion
by: Liu, Tongxuan, et al.
Published: (2024)
by: Liu, Tongxuan, et al.
Published: (2024)
Thinking with Generated Images
by: Chern, Ethan, et al.
Published: (2025)
by: Chern, Ethan, et al.
Published: (2025)
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
by: Pisano, Matthew, et al.
Published: (2023)
by: Pisano, Matthew, et al.
Published: (2023)
Dynamic Role Assignment for Multi-Agent Debate
by: Zhang, Miao, et al.
Published: (2026)
by: Zhang, Miao, et al.
Published: (2026)
Fast Adversarial Training against Textual Adversarial Attacks
by: Yang, Yichen, et al.
Published: (2024)
by: Yang, Yichen, et al.
Published: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Gradual Vigilance and Interval Communication: Enhancing Value Alignment in Multi-Agent Debates
by: Zou, Rui, et al.
Published: (2024)
by: Zou, Rui, et al.
Published: (2024)
MADIAVE: Multi-Agent Debate for Implicit Attribute Value Extraction
by: Huang, Wei-Chieh, et al.
Published: (2025)
by: Huang, Wei-Chieh, et al.
Published: (2025)
M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation
by: Feng, Zhaopeng, et al.
Published: (2024)
by: Feng, Zhaopeng, et al.
Published: (2024)
S$^2$-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency
by: Zeng, Yuting, et al.
Published: (2025)
by: Zeng, Yuting, et al.
Published: (2025)
iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM Inference
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
CascadeDebate: Multi-Agent Deliberation for Cost-Aware LLM Cascades
by: Chang, Raeyoung, et al.
Published: (2026)
by: Chang, Raeyoung, et al.
Published: (2026)
Learning to Break: Knowledge-Enhanced Reasoning in Multi-Agent Debate System
by: Wang, Haotian, et al.
Published: (2023)
by: Wang, Haotian, et al.
Published: (2023)
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
by: Mo, Lingbo, et al.
Published: (2024)
by: Mo, Lingbo, et al.
Published: (2024)
Generative AI Act II: Test Time Scaling Drives Cognition Engineering
by: Xia, Shijie, et al.
Published: (2025)
by: Xia, Shijie, et al.
Published: (2025)
AgentCourt: Simulating Court with Adversarial Evolvable Lawyer Agents
by: Chen, Guhong, et al.
Published: (2024)
by: Chen, Guhong, et al.
Published: (2024)
Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks
by: Liu, Haoyu, et al.
Published: (2026)
by: Liu, Haoyu, et al.
Published: (2026)
Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument Generation
by: Hu, Zhe, et al.
Published: (2024)
by: Hu, Zhe, et al.
Published: (2024)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
by: Zhu, Shenzhe
Published: (2025)
by: Zhu, Shenzhe
Published: (2025)
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
by: Smit, Andries, et al.
Published: (2023)
by: Smit, Andries, et al.
Published: (2023)
Voting or Consensus? Decision-Making in Multi-Agent Debate
by: Kaesberg, Lars Benedikt, et al.
Published: (2025)
by: Kaesberg, Lars Benedikt, et al.
Published: (2025)
SEP-Attack: A Simple and Effective Paradigm for Transfer-Based Textual Adversarial Attack
by: Liu, Han, et al.
Published: (2026)
by: Liu, Han, et al.
Published: (2026)
LIMO: Less is More for Reasoning
by: Ye, Yixin, et al.
Published: (2025)
by: Ye, Yixin, et al.
Published: (2025)
MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
by: Ning, Yucheng, et al.
Published: (2025)
by: Ning, Yucheng, et al.
Published: (2025)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
by: Ban, Minjeong, et al.
Published: (2026)
by: Ban, Minjeong, et al.
Published: (2026)
HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on Text
by: Liu, Han, et al.
Published: (2024)
by: Liu, Han, et al.
Published: (2024)
Negotiating with LLMS: Prompt Hacks, Skill Gaps, and Reasoning Deficits
by: Schneider, Johannes, et al.
Published: (2023)
by: Schneider, Johannes, et al.
Published: (2023)
Adversarial Attacks and Defense for Conversation Entailment Task
by: Yang, Zhenning, et al.
Published: (2024)
by: Yang, Zhenning, et al.
Published: (2024)
STACK: Adversarial Attacks on LLM Safeguard Pipelines
by: McKenzie, Ian R., et al.
Published: (2025)
by: McKenzie, Ian R., et al.
Published: (2025)
Multiple LLM Agents Debate for Equitable Cultural Alignment
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
by: Wang, Jianze, et al.
Published: (2026)
by: Wang, Jianze, et al.
Published: (2026)
AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation
by: He, Zhitao, et al.
Published: (2024)
by: He, Zhitao, et al.
Published: (2024)
Tree-of-Debate: Multi-Persona Debate Trees Elicit Critical Thinking for Scientific Comparative Analysis
by: Kargupta, Priyanka, et al.
Published: (2025)
by: Kargupta, Priyanka, et al.
Published: (2025)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
by: Zhou, Xiaofeng, et al.
Published: (2025)
by: Zhou, Xiaofeng, et al.
Published: (2025)
CORBA: Contagious Recursive Blocking Attacks on Multi-Agent Systems Based on Large Language Models
by: Zhou, Zhenhong, et al.
Published: (2025)
by: Zhou, Zhenhong, et al.
Published: (2025)
Similar Items
-
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024) -
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024) -
MultiAgent Collaboration Attack: Investigating Adversarial Attacks in Large Language Model Collaborations via Debate
by: Amayuelas, Alfonso, et al.
Published: (2024) -
BeHonest: Benchmarking Honesty in Large Language Models
by: Chern, Steffi, et al.
Published: (2024) -
GroupDebate: Enhancing the Efficiency of Multi-Agent Debate Using Group Discussion
by: Liu, Tongxuan, et al.
Published: (2024)