DeliberationBench: When Do More Voices Hurt? A Controlled Study of Multi-LLM Deliberation Protocols
Fuente:
arXiv
Salvato in:
| Autori principali: | Kaushal, Vaarunay, Singh, Taranveer |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DeliberationBench: A Normative Benchmark for the Influence of Large Language Models on Users' Views
di: Hewitt, Luke, et al.
Pubblicazione: (2026)
di: Hewitt, Luke, et al.
Pubblicazione: (2026)
CascadeDebate: Multi-Agent Deliberation for Cost-Aware LLM Cascades
di: Chang, Raeyoung, et al.
Pubblicazione: (2026)
di: Chang, Raeyoung, et al.
Pubblicazione: (2026)
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
di: Zhang, Zhiwei, et al.
Pubblicazione: (2025)
di: Zhang, Zhiwei, et al.
Pubblicazione: (2025)
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
di: Rakshit, Sushrita, et al.
Pubblicazione: (2026)
di: Rakshit, Sushrita, et al.
Pubblicazione: (2026)
Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2026)
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2026)
Achieving Unanimous Consensus Through Multi-Agent Deliberation
di: Pokharel, Apurba, et al.
Pubblicazione: (2025)
di: Pokharel, Apurba, et al.
Pubblicazione: (2025)
Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes
di: Nath, Abhijnan, et al.
Pubblicazione: (2025)
di: Nath, Abhijnan, et al.
Pubblicazione: (2025)
RDR: the Recap, Deliberate, and Respond Method for Enhanced Language Understanding
di: Zi, Yuxin, et al.
Pubblicazione: (2023)
di: Zi, Yuxin, et al.
Pubblicazione: (2023)
DeepCritic: Deliberate Critique with Large Language Models
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
Deliberation in Latent Space via Differentiable Cache Augmentation
di: Liu, Luyang, et al.
Pubblicazione: (2024)
di: Liu, Luyang, et al.
Pubblicazione: (2024)
How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
di: Fukui, Hiroki
Pubblicazione: (2026)
di: Fukui, Hiroki
Pubblicazione: (2026)
Collaborative Evaluation of Deepfake Text with Deliberation-Enhancing Dialogue Systems
di: Lee, Jooyoung, et al.
Pubblicazione: (2025)
di: Lee, Jooyoung, et al.
Pubblicazione: (2025)
One Panel Does Not Fit All: Case-Adaptive Multi-Agent Deliberation for Clinical Prediction
di: Lu, Yuxing, et al.
Pubblicazione: (2026)
di: Lu, Yuxing, et al.
Pubblicazione: (2026)
From Debate to Deliberation: Structured Collective Reasoning with Typed Epistemic Acts
di: Prakash, Sunil
Pubblicazione: (2026)
di: Prakash, Sunil
Pubblicazione: (2026)
A Survey of Test-Time Compute: From Intuitive Inference to Deliberate Reasoning
di: Ji, Yixin, et al.
Pubblicazione: (2025)
di: Ji, Yixin, et al.
Pubblicazione: (2025)
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
di: Kumarage, Tharindu, et al.
Pubblicazione: (2025)
di: Kumarage, Tharindu, et al.
Pubblicazione: (2025)
From Argumentation to Deliberation: Perspectivized Stance Vectors for Fine-grained (Dis)agreement Analysis
di: Plenz, Moritz, et al.
Pubblicazione: (2025)
di: Plenz, Moritz, et al.
Pubblicazione: (2025)
MedSimAI: Simulation and Formative Feedback Generation to Enhance Deliberate Practice in Medical Education
di: Hicke, Yann, et al.
Pubblicazione: (2025)
di: Hicke, Yann, et al.
Pubblicazione: (2025)
Agile Deliberation: Concept Deliberation for Subjective Visual Classification
di: Wang, Leijie, et al.
Pubblicazione: (2025)
di: Wang, Leijie, et al.
Pubblicazione: (2025)
When "Better" Prompts Hurt: Evaluation-Driven Iteration for LLM Applications
di: Commey, Daniel
Pubblicazione: (2026)
di: Commey, Daniel
Pubblicazione: (2026)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
AQuA -- Combining Experts' and Non-Experts' Views To Assess Deliberation Quality in Online Discussions Using LLMs
di: Behrendt, Maike, et al.
Pubblicazione: (2024)
di: Behrendt, Maike, et al.
Pubblicazione: (2024)
VoiceBench: Benchmarking LLM-Based Voice Assistants
di: Chen, Yiming, et al.
Pubblicazione: (2024)
di: Chen, Yiming, et al.
Pubblicazione: (2024)
Evaluating Social Bias in RAG Systems: When External Context Helps and Reasoning Hurts
di: Parihar, Shweta, et al.
Pubblicazione: (2026)
di: Parihar, Shweta, et al.
Pubblicazione: (2026)
The Wisdom of Deliberating AI Crowds: Does Deliberation Improve LLM-Based Forecasting?
di: Schneider, Paul, et al.
Pubblicazione: (2025)
di: Schneider, Paul, et al.
Pubblicazione: (2025)
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
di: Du, Yufeng, et al.
Pubblicazione: (2025)
di: Du, Yufeng, et al.
Pubblicazione: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
di: Tan, Haoran, et al.
Pubblicazione: (2025)
di: Tan, Haoran, et al.
Pubblicazione: (2025)
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
di: Liu, Ruikang, et al.
Pubblicazione: (2025)
di: Liu, Ruikang, et al.
Pubblicazione: (2025)
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
di: Zhao, Beidi, et al.
Pubblicazione: (2026)
di: Zhao, Beidi, et al.
Pubblicazione: (2026)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
di: Jain, Dhruv, et al.
Pubblicazione: (2025)
di: Jain, Dhruv, et al.
Pubblicazione: (2025)
When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
di: Faisal, Faizan
Pubblicazione: (2026)
di: Faisal, Faizan
Pubblicazione: (2026)
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
di: Xu, Ke, et al.
Pubblicazione: (2026)
di: Xu, Ke, et al.
Pubblicazione: (2026)
When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging
di: Li, Yayuan, et al.
Pubblicazione: (2026)
di: Li, Yayuan, et al.
Pubblicazione: (2026)
Drift No More? Context Equilibria in Multi-Turn LLM Interactions
di: Dongre, Vardhan, et al.
Pubblicazione: (2025)
di: Dongre, Vardhan, et al.
Pubblicazione: (2025)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
di: Zhu, Jie, et al.
Pubblicazione: (2026)
di: Zhu, Jie, et al.
Pubblicazione: (2026)
Algorithmic Approaches to Opinion Selection for Online Deliberation: A Comparative Study
di: Hafid, Salim, et al.
Pubblicazione: (2026)
di: Hafid, Salim, et al.
Pubblicazione: (2026)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
di: Du, Yuhao, et al.
Pubblicazione: (2025)
di: Du, Yuhao, et al.
Pubblicazione: (2025)
Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis
di: Doske, VD
Pubblicazione: (2026)
di: Doske, VD
Pubblicazione: (2026)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
di: Tonga, Junior Cedric, et al.
Pubblicazione: (2025)
di: Tonga, Junior Cedric, et al.
Pubblicazione: (2025)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
di: Schmidt, Jan-Philipp
Pubblicazione: (2026)
di: Schmidt, Jan-Philipp
Pubblicazione: (2026)
Documenti analoghi
-
DeliberationBench: A Normative Benchmark for the Influence of Large Language Models on Users' Views
di: Hewitt, Luke, et al.
Pubblicazione: (2026) -
CascadeDebate: Multi-Agent Deliberation for Cost-Aware LLM Cascades
di: Chang, Raeyoung, et al.
Pubblicazione: (2026) -
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
di: Zhang, Zhiwei, et al.
Pubblicazione: (2025) -
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
di: Rakshit, Sushrita, et al.
Pubblicazione: (2026) -
Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs
di: Anantaprayoon, Panatchakorn, et al.
Pubblicazione: (2026)