Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Eddie, Wang, Dashun |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs
par: Allen, Bradley P., et autres
Publié: (2025)
par: Allen, Bradley P., et autres
Publié: (2025)
The Illusion of Stochasticity in LLMs
par: Gu, Xiangming, et autres
Publié: (2026)
par: Gu, Xiangming, et autres
Publié: (2026)
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
par: Andreev, Nikita, et autres
Publié: (2024)
par: Andreev, Nikita, et autres
Publié: (2024)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
par: Lu, Junyu, et autres
Publié: (2025)
par: Lu, Junyu, et autres
Publié: (2025)
When Disagreements Elicit Robustness: Investigating Self-Repair Capabilities under LLM Multi-Agent Disagreements
par: Ju, Tianjie, et autres
Publié: (2025)
par: Ju, Tianjie, et autres
Publié: (2025)
Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation
par: Wu, Weimin, et autres
Publié: (2025)
par: Wu, Weimin, et autres
Publié: (2025)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
par: Zhou, Yujun, et autres
Publié: (2024)
par: Zhou, Yujun, et autres
Publié: (2024)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
par: Rostamkhani, Mohammadmostafa, et autres
Publié: (2024)
par: Rostamkhani, Mohammadmostafa, et autres
Publié: (2024)
The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
par: Tomov, Tim, et autres
Publié: (2025)
par: Tomov, Tim, et autres
Publié: (2025)
SciDA: Scientific Dynamic Assessor of LLMs
par: Zhou, Junting, et autres
Publié: (2025)
par: Zhou, Junting, et autres
Publié: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
par: Chen, Haotian, et autres
Publié: (2025)
par: Chen, Haotian, et autres
Publié: (2025)
Quantifying the Benefit of Artificial Intelligence for Scientific Research
par: Gao, Jian, et autres
Publié: (2023)
par: Gao, Jian, et autres
Publié: (2023)
Leveraging Annotator Disagreement for Text Classification
par: Xu, Jin, et autres
Publié: (2024)
par: Xu, Jin, et autres
Publié: (2024)
Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis
par: Lu, Junyu, et autres
Publié: (2026)
par: Lu, Junyu, et autres
Publié: (2026)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
par: Janiak, Denis, et autres
Publié: (2025)
par: Janiak, Denis, et autres
Publié: (2025)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
par: Ullman, Tomer
Publié: (2024)
par: Ullman, Tomer
Publié: (2024)
Quantifying and Predicting Disagreement in Graded Human Ratings
par: Zhang, Leixin, et autres
Publié: (2026)
par: Zhang, Leixin, et autres
Publié: (2026)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
par: Xu, Wanghan, et autres
Publié: (2025)
par: Xu, Wanghan, et autres
Publié: (2025)
MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
par: Tian, Yang, et autres
Publié: (2025)
par: Tian, Yang, et autres
Publié: (2025)
MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors
par: Hikal, Baraa, et autres
Publié: (2025)
par: Hikal, Baraa, et autres
Publié: (2025)
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
par: Liu, Yujie, et autres
Publié: (2025)
par: Liu, Yujie, et autres
Publié: (2025)
Pun Unintended: LLMs and the Illusion of Humor Understanding
par: Zangari, Alessandro, et autres
Publié: (2025)
par: Zangari, Alessandro, et autres
Publié: (2025)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
par: Han, Pengrui, et autres
Publié: (2025)
par: Han, Pengrui, et autres
Publié: (2025)
NUTMEG: Separating Signal From Noise in Annotator Disagreement
par: Ivey, Jonathan, et autres
Publié: (2025)
par: Ivey, Jonathan, et autres
Publié: (2025)
Do Differences in Values Influence Disagreements in Online Discussions?
par: van der Meer, Michiel, et autres
Publié: (2023)
par: van der Meer, Michiel, et autres
Publié: (2023)
Bridging the Gap: In-Context Learning for Modeling Human Disagreement
par: Muscato, Benedetta, et autres
Publié: (2025)
par: Muscato, Benedetta, et autres
Publié: (2025)
From Disagreement to Understanding: The Case for Ambiguity Detection in NLI
par: Jayaweera, Chathuri, et autres
Publié: (2025)
par: Jayaweera, Chathuri, et autres
Publié: (2025)
LEGOBench: Scientific Leaderboard Generation Benchmark
par: Singh, Shruti, et autres
Publié: (2024)
par: Singh, Shruti, et autres
Publié: (2024)
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
par: Cai, Hengxing, et autres
Publié: (2024)
par: Cai, Hengxing, et autres
Publié: (2024)
Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts
par: Seshadri, Preethi, et autres
Publié: (2025)
par: Seshadri, Preethi, et autres
Publié: (2025)
Learning-to-Context Slope: Evaluating In-Context Learning Effectiveness Beyond Performance Illusions
par: Wang, Dingzriui, et autres
Publié: (2025)
par: Wang, Dingzriui, et autres
Publié: (2025)
Benchmarking LLMs via Uncertainty Quantification
par: Ye, Fanghua, et autres
Publié: (2024)
par: Ye, Fanghua, et autres
Publié: (2024)
IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
par: Shahgir, Haz Sameen, et autres
Publié: (2024)
par: Shahgir, Haz Sameen, et autres
Publié: (2024)
Extreme Miscalibration and the Illusion of Adversarial Robustness
par: Raina, Vyas, et autres
Publié: (2024)
par: Raina, Vyas, et autres
Publié: (2024)
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
par: Javaji, Shashidhar Reddy, et autres
Publié: (2025)
par: Javaji, Shashidhar Reddy, et autres
Publié: (2025)
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
par: Xu, Yinuo, et autres
Publié: (2026)
par: Xu, Yinuo, et autres
Publié: (2026)
Disagreement as Data: Reasoning Trace Analytics in Multi-Agent Systems
par: Tajik, Elham, et autres
Publié: (2026)
par: Tajik, Elham, et autres
Publié: (2026)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
par: Aghaebe, Favour Yahdii, et autres
Publié: (2026)
par: Aghaebe, Favour Yahdii, et autres
Publié: (2026)
The Illusion of State in State-Space Models
par: Merrill, William, et autres
Publié: (2024)
par: Merrill, William, et autres
Publié: (2024)
Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
par: Xu, Yinuo, et autres
Publié: (2025)
par: Xu, Yinuo, et autres
Publié: (2025)
Documents similaires
-
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs
par: Allen, Bradley P., et autres
Publié: (2025) -
The Illusion of Stochasticity in LLMs
par: Gu, Xiangming, et autres
Publié: (2026) -
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
par: Andreev, Nikita, et autres
Publié: (2024) -
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
par: Lu, Junyu, et autres
Publié: (2025) -
When Disagreements Elicit Robustness: Investigating Self-Repair Capabilities under LLM Multi-Agent Disagreements
par: Ju, Tianjie, et autres
Publié: (2025)