Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
Fuente:
arXiv
Guardado en:
| Autores principales: | Choi, Junhyuk, Kwon, Jeongyoun, Kim, Heeju, Cho, Haeun, Jung, Hayeong, Min, Sehee, Kim, Bugeun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
People will agree what I think: Investigating LLM's False Consensus Effect
por: Choi, Junhyuk, et al.
Publicado: (2024)
por: Choi, Junhyuk, et al.
Publicado: (2024)
A Stereotype Content Analysis on Color-related Social Bias in Large Vision Language Models
por: Choi, Junhyuk, et al.
Publicado: (2025)
por: Choi, Junhyuk, et al.
Publicado: (2025)
Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
por: Choi, Junhyuk, et al.
Publicado: (2025)
por: Choi, Junhyuk, et al.
Publicado: (2025)
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
por: Kim, Byungjun, et al.
Publicado: (2024)
por: Kim, Byungjun, et al.
Publicado: (2024)
Examining Identity Drift in Conversations of LLM Agents
por: Choi, Junhyuk, et al.
Publicado: (2024)
por: Choi, Junhyuk, et al.
Publicado: (2024)
DART: An AIGT Detector using AMR of Rephrased Text
por: Park, Hyeonchu, et al.
Publicado: (2024)
por: Park, Hyeonchu, et al.
Publicado: (2024)
Leveraging Large Language Models for Active Merchant Non-player Characters
por: Kim, Byungjun, et al.
Publicado: (2024)
por: Kim, Byungjun, et al.
Publicado: (2024)
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
por: Choi, Junhyuk, et al.
Publicado: (2026)
por: Choi, Junhyuk, et al.
Publicado: (2026)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
por: Mun, Jimin, et al.
Publicado: (2026)
por: Mun, Jimin, et al.
Publicado: (2026)
DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents
por: Kim, Jiho, et al.
Publicado: (2024)
por: Kim, Jiho, et al.
Publicado: (2024)
VorTEX: Various overlap ratio for Target speech EXtraction
por: Oh, Ro-hoon, et al.
Publicado: (2026)
por: Oh, Ro-hoon, et al.
Publicado: (2026)
Acoustic-based Gender Differentiation in Speech-aware Language Models
por: Choi, Junhyuk, et al.
Publicado: (2025)
por: Choi, Junhyuk, et al.
Publicado: (2025)
VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model
por: Choi, Junhyuk, et al.
Publicado: (2025)
por: Choi, Junhyuk, et al.
Publicado: (2025)
PSYCHE: A Multi-faceted Patient Simulation Framework for Evaluation of Psychiatric Assessment Conversational Agents
por: Lee, Jingoo, et al.
Publicado: (2025)
por: Lee, Jingoo, et al.
Publicado: (2025)
Beyond Single-User Dialogue: Assessing Multi-User Dialogue State Tracking Capabilities of Large Language Models
por: Song, Sangmin, et al.
Publicado: (2025)
por: Song, Sangmin, et al.
Publicado: (2025)
An Empirical Study of Group Conformity in Multi-Agent Systems
por: Choi, Min, et al.
Publicado: (2025)
por: Choi, Min, et al.
Publicado: (2025)
R2-KG: General-Purpose Dual-Agent Framework for Reliable Reasoning on Knowledge Graphs
por: Jo, Sumin, et al.
Publicado: (2025)
por: Jo, Sumin, et al.
Publicado: (2025)
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
por: Choi, Dongwook, et al.
Publicado: (2025)
por: Choi, Dongwook, et al.
Publicado: (2025)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
por: Lee, Jiyoung, et al.
Publicado: (2025)
por: Lee, Jiyoung, et al.
Publicado: (2025)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
por: Ban, Minjeong, et al.
Publicado: (2026)
por: Ban, Minjeong, et al.
Publicado: (2026)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
por: Lu, Dongxu, et al.
Publicado: (2025)
por: Lu, Dongxu, et al.
Publicado: (2025)
CUB: Benchmarking Context Utilisation Techniques for Language Models
por: Hagström, Lovisa, et al.
Publicado: (2025)
por: Hagström, Lovisa, et al.
Publicado: (2025)
Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification
por: Demir, M. Mikail, et al.
Publicado: (2026)
por: Demir, M. Mikail, et al.
Publicado: (2026)
MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
por: Lee, Woongkyu, et al.
Publicado: (2025)
por: Lee, Woongkyu, et al.
Publicado: (2025)
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
por: Choi, Juhwan, et al.
Publicado: (2024)
por: Choi, Juhwan, et al.
Publicado: (2024)
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
por: Jung, Jimin, et al.
Publicado: (2026)
por: Jung, Jimin, et al.
Publicado: (2026)
Personalized Author Obfuscation with Large Language Models
por: Shokri, Mohammad, et al.
Publicado: (2025)
por: Shokri, Mohammad, et al.
Publicado: (2025)
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks
por: Choi, Juhwan, et al.
Publicado: (2024)
por: Choi, Juhwan, et al.
Publicado: (2024)
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
por: Deng, Jiaqi, et al.
Publicado: (2025)
por: Deng, Jiaqi, et al.
Publicado: (2025)
Opacity as Authority: Arbitrariness and the Preclusion of Contestation
por: Kayembe, Naomi Omeonga wa
Publicado: (2025)
por: Kayembe, Naomi Omeonga wa
Publicado: (2025)
Anterior's Approach to Fairness Evaluation of Automated Prior Authorization System
por: Selvaraj, Sai P., et al.
Publicado: (2026)
por: Selvaraj, Sai P., et al.
Publicado: (2026)
Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings
por: Choi, Juhwan, et al.
Publicado: (2024)
por: Choi, Juhwan, et al.
Publicado: (2024)
PersonalHomeBench: Evaluating Agents in Personalized Smart Homes
por: Bharadwaj, Manasa, et al.
Publicado: (2026)
por: Bharadwaj, Manasa, et al.
Publicado: (2026)
Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records
por: Kwon, Yeonsu, et al.
Publicado: (2026)
por: Kwon, Yeonsu, et al.
Publicado: (2026)
MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference
por: Park, Jeonghyun, et al.
Publicado: (2025)
por: Park, Jeonghyun, et al.
Publicado: (2025)
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning
por: Jung, Min Jae, et al.
Publicado: (2024)
por: Jung, Min Jae, et al.
Publicado: (2024)
MATA: Multi-Agent Framework for Reliable and Flexible Table Question Answering
por: Hyeon, Sieun, et al.
Publicado: (2026)
por: Hyeon, Sieun, et al.
Publicado: (2026)
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation
por: Li, Yu, et al.
Publicado: (2024)
por: Li, Yu, et al.
Publicado: (2024)
SLM as Guardian: Pioneering AI Safety with Small Language Models
por: Kwon, Ohjoon, et al.
Publicado: (2024)
por: Kwon, Ohjoon, et al.
Publicado: (2024)
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework
por: Liu, Xuelin, et al.
Publicado: (2024)
por: Liu, Xuelin, et al.
Publicado: (2024)
Ejemplares similares
-
People will agree what I think: Investigating LLM's False Consensus Effect
por: Choi, Junhyuk, et al.
Publicado: (2024) -
A Stereotype Content Analysis on Color-related Social Bias in Large Vision Language Models
por: Choi, Junhyuk, et al.
Publicado: (2025) -
Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
por: Choi, Junhyuk, et al.
Publicado: (2025) -
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
por: Kim, Byungjun, et al.
Publicado: (2024) -
Examining Identity Drift in Conversations of LLM Agents
por: Choi, Junhyuk, et al.
Publicado: (2024)