Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Harrasse, Abir, Bandi, Chaithanya, Bandi, Hari |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReDAct: Uncertainty-Aware Deferral for LLM Agents
von: Piatrashyn, Dzianis, et al.
Veröffentlicht: (2026)
von: Piatrashyn, Dzianis, et al.
Veröffentlicht: (2026)
Context, Reasoning, and Hierarchy: A Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP
von: Bogdanov, Igor, et al.
Veröffentlicht: (2026)
von: Bogdanov, Igor, et al.
Veröffentlicht: (2026)
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
von: Oh, Gyutaek, et al.
Veröffentlicht: (2025)
von: Oh, Gyutaek, et al.
Veröffentlicht: (2025)
MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety
von: Wen, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Wen, Xiaoyu, et al.
Veröffentlicht: (2026)
GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
RunAgent: Interpreting Natural-Language Plans with Constraint-Guided Execution
von: Srivastava, Arunabh, et al.
Veröffentlicht: (2026)
von: Srivastava, Arunabh, et al.
Veröffentlicht: (2026)
An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring
von: Ebrahimi, Sana, et al.
Veröffentlicht: (2025)
von: Ebrahimi, Sana, et al.
Veröffentlicht: (2025)
Towards Reliable ML Feature Engineering via Planning in Constrained-Topology of LLM Agents
von: Thakur, Himanshu, et al.
Veröffentlicht: (2026)
von: Thakur, Himanshu, et al.
Veröffentlicht: (2026)
PAACE: A Plan-Aware Automated Agent Context Engineering Framework
von: Yuksel, Kamer Ali
Veröffentlicht: (2025)
von: Yuksel, Kamer Ali
Veröffentlicht: (2025)
From Debate to Deliberation: Structured Collective Reasoning with Typed Epistemic Acts
von: Prakash, Sunil
Veröffentlicht: (2026)
von: Prakash, Sunil
Veröffentlicht: (2026)
MAATS: A Multi-Agent Automated Translation System Based on MQM Evaluation
von: Wang, George, et al.
Veröffentlicht: (2025)
von: Wang, George, et al.
Veröffentlicht: (2025)
In-Context Environments Induce Evaluation-Awareness in Language Models
von: Chaudhary, Maheep
Veröffentlicht: (2026)
von: Chaudhary, Maheep
Veröffentlicht: (2026)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML
von: Trirat, Patara, et al.
Veröffentlicht: (2024)
von: Trirat, Patara, et al.
Veröffentlicht: (2024)
EdgeAgentX: A Novel Framework for Agentic AI at the Edge in Military Communication Networks
von: Ray, Abir
Veröffentlicht: (2025)
von: Ray, Abir
Veröffentlicht: (2025)
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
von: Bozdag, Nimet Beyza, et al.
Veröffentlicht: (2025)
von: Bozdag, Nimet Beyza, et al.
Veröffentlicht: (2025)
Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline
von: Xu, Jiawei, et al.
Veröffentlicht: (2026)
von: Xu, Jiawei, et al.
Veröffentlicht: (2026)
Image, Word and Thought: A More Challenging Language Task for the Iterated Learning Model
von: Lee, Hyoyeon, et al.
Veröffentlicht: (2026)
von: Lee, Hyoyeon, et al.
Veröffentlicht: (2026)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
von: Liu, Yuhan, et al.
Veröffentlicht: (2024)
von: Liu, Yuhan, et al.
Veröffentlicht: (2024)
RDMA: Cost Effective Agent-Driven Rare Disease Mining from Electronic Health Records
von: Wu, John, et al.
Veröffentlicht: (2025)
von: Wu, John, et al.
Veröffentlicht: (2025)
Opponent Shaping in LLM Agents
von: Segura, Marta Emili Garcia, et al.
Veröffentlicht: (2025)
von: Segura, Marta Emili Garcia, et al.
Veröffentlicht: (2025)
Verification-Aware Planning for Multi-Agent Systems
von: Xu, Tianyang, et al.
Veröffentlicht: (2025)
von: Xu, Tianyang, et al.
Veröffentlicht: (2025)
LLM Agents Making Agent Tools
von: Wölflein, Georg, et al.
Veröffentlicht: (2025)
von: Wölflein, Georg, et al.
Veröffentlicht: (2025)
Stochastic Self-Organization in Multi-Agent Systems
von: Tastan, Nurbek, et al.
Veröffentlicht: (2025)
von: Tastan, Nurbek, et al.
Veröffentlicht: (2025)
Multi-agent Architecture Search via Agentic Supernet
von: Zhang, Guibin, et al.
Veröffentlicht: (2025)
von: Zhang, Guibin, et al.
Veröffentlicht: (2025)
Safe Multi-agent Reinforcement Learning with Natural Language Constraints
von: Wang, Ziyan, et al.
Veröffentlicht: (2024)
von: Wang, Ziyan, et al.
Veröffentlicht: (2024)
Exploring Modularity of Agentic Systems for Drug Discovery
von: van Weesep, Laura, et al.
Veröffentlicht: (2025)
von: van Weesep, Laura, et al.
Veröffentlicht: (2025)
Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
von: Tastan, Nurbek, et al.
Veröffentlicht: (2026)
von: Tastan, Nurbek, et al.
Veröffentlicht: (2026)
BRIDGE: Bootstrapping Text to Control Time-Series Generation via Multi-Agent Iterative Optimization and Diffusion Modeling
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Agents' Room: Narrative Generation through Multi-step Collaboration
von: Huot, Fantine, et al.
Veröffentlicht: (2024)
von: Huot, Fantine, et al.
Veröffentlicht: (2024)
Self-guided Knowledgeable Network of Thoughts: Amplifying Reasoning with Large Language Models
von: Chen, Chao-Chi, et al.
Veröffentlicht: (2024)
von: Chen, Chao-Chi, et al.
Veröffentlicht: (2024)
Under the Influence: Quantifying Persuasion and Vigilance in Large Language Models
von: Robinson, Sasha, et al.
Veröffentlicht: (2026)
von: Robinson, Sasha, et al.
Veröffentlicht: (2026)
ReaGAN: Node-as-Agent-Reasoning Graph Agentic Network
von: Guo, Minghao, et al.
Veröffentlicht: (2025)
von: Guo, Minghao, et al.
Veröffentlicht: (2025)
The emergence of numerical representations in communicating artificial agents
von: Mihai, Daniela, et al.
Veröffentlicht: (2026)
von: Mihai, Daniela, et al.
Veröffentlicht: (2026)
G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
von: Zhang, Guibin, et al.
Veröffentlicht: (2025)
von: Zhang, Guibin, et al.
Veröffentlicht: (2025)
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
von: Liu, Mickel, et al.
Veröffentlicht: (2025)
von: Liu, Mickel, et al.
Veröffentlicht: (2025)
LatentMem: Customizing Latent Memory for Multi-Agent Systems
von: Fu, Muxin, et al.
Veröffentlicht: (2026)
von: Fu, Muxin, et al.
Veröffentlicht: (2026)
Factorio Learning Environment
von: Hopkins, Jack, et al.
Veröffentlicht: (2025)
von: Hopkins, Jack, et al.
Veröffentlicht: (2025)
Learning Translations: Emergent Communication Pretraining for Cooperative Language Acquisition
von: Cope, Dylan, et al.
Veröffentlicht: (2024)
von: Cope, Dylan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReDAct: Uncertainty-Aware Deferral for LLM Agents
von: Piatrashyn, Dzianis, et al.
Veröffentlicht: (2026) -
Context, Reasoning, and Hierarchy: A Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP
von: Bogdanov, Igor, et al.
Veröffentlicht: (2026) -
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
von: Oh, Gyutaek, et al.
Veröffentlicht: (2025) -
MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety
von: Wen, Xiaoyu, et al.
Veröffentlicht: (2026) -
GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)