Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Eddie, Wang, Dashun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs
von: Allen, Bradley P., et al.
Veröffentlicht: (2025)
von: Allen, Bradley P., et al.
Veröffentlicht: (2025)
The Illusion of Stochasticity in LLMs
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
von: Gu, Xiangming, et al.
Veröffentlicht: (2026)
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
von: Andreev, Nikita, et al.
Veröffentlicht: (2024)
von: Andreev, Nikita, et al.
Veröffentlicht: (2024)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
When Disagreements Elicit Robustness: Investigating Self-Repair Capabilities under LLM Multi-Agent Disagreements
von: Ju, Tianjie, et al.
Veröffentlicht: (2025)
von: Ju, Tianjie, et al.
Veröffentlicht: (2025)
Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation
von: Wu, Weimin, et al.
Veröffentlicht: (2025)
von: Wu, Weimin, et al.
Veröffentlicht: (2025)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
von: Tomov, Tim, et al.
Veröffentlicht: (2025)
von: Tomov, Tim, et al.
Veröffentlicht: (2025)
SciDA: Scientific Dynamic Assessor of LLMs
von: Zhou, Junting, et al.
Veröffentlicht: (2025)
von: Zhou, Junting, et al.
Veröffentlicht: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
von: Chen, Haotian, et al.
Veröffentlicht: (2025)
von: Chen, Haotian, et al.
Veröffentlicht: (2025)
Quantifying the Benefit of Artificial Intelligence for Scientific Research
von: Gao, Jian, et al.
Veröffentlicht: (2023)
von: Gao, Jian, et al.
Veröffentlicht: (2023)
Leveraging Annotator Disagreement for Text Classification
von: Xu, Jin, et al.
Veröffentlicht: (2024)
von: Xu, Jin, et al.
Veröffentlicht: (2024)
Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis
von: Lu, Junyu, et al.
Veröffentlicht: (2026)
von: Lu, Junyu, et al.
Veröffentlicht: (2026)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
von: Ullman, Tomer
Veröffentlicht: (2024)
von: Ullman, Tomer
Veröffentlicht: (2024)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
von: Xu, Wanghan, et al.
Veröffentlicht: (2025)
von: Xu, Wanghan, et al.
Veröffentlicht: (2025)
Quantifying and Predicting Disagreement in Graded Human Ratings
von: Zhang, Leixin, et al.
Veröffentlicht: (2026)
von: Zhang, Leixin, et al.
Veröffentlicht: (2026)
MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
von: Tian, Yang, et al.
Veröffentlicht: (2025)
von: Tian, Yang, et al.
Veröffentlicht: (2025)
MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors
von: Hikal, Baraa, et al.
Veröffentlicht: (2025)
von: Hikal, Baraa, et al.
Veröffentlicht: (2025)
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
Pun Unintended: LLMs and the Illusion of Humor Understanding
von: Zangari, Alessandro, et al.
Veröffentlicht: (2025)
von: Zangari, Alessandro, et al.
Veröffentlicht: (2025)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
von: Han, Pengrui, et al.
Veröffentlicht: (2025)
von: Han, Pengrui, et al.
Veröffentlicht: (2025)
LEGOBench: Scientific Leaderboard Generation Benchmark
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
NUTMEG: Separating Signal From Noise in Annotator Disagreement
von: Ivey, Jonathan, et al.
Veröffentlicht: (2025)
von: Ivey, Jonathan, et al.
Veröffentlicht: (2025)
Do Differences in Values Influence Disagreements in Online Discussions?
von: van der Meer, Michiel, et al.
Veröffentlicht: (2023)
von: van der Meer, Michiel, et al.
Veröffentlicht: (2023)
Bridging the Gap: In-Context Learning for Modeling Human Disagreement
von: Muscato, Benedetta, et al.
Veröffentlicht: (2025)
von: Muscato, Benedetta, et al.
Veröffentlicht: (2025)
From Disagreement to Understanding: The Case for Ambiguity Detection in NLI
von: Jayaweera, Chathuri, et al.
Veröffentlicht: (2025)
von: Jayaweera, Chathuri, et al.
Veröffentlicht: (2025)
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
von: Cai, Hengxing, et al.
Veröffentlicht: (2024)
von: Cai, Hengxing, et al.
Veröffentlicht: (2024)
Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts
von: Seshadri, Preethi, et al.
Veröffentlicht: (2025)
von: Seshadri, Preethi, et al.
Veröffentlicht: (2025)
Benchmarking LLMs via Uncertainty Quantification
von: Ye, Fanghua, et al.
Veröffentlicht: (2024)
von: Ye, Fanghua, et al.
Veröffentlicht: (2024)
Learning-to-Context Slope: Evaluating In-Context Learning Effectiveness Beyond Performance Illusions
von: Wang, Dingzriui, et al.
Veröffentlicht: (2025)
von: Wang, Dingzriui, et al.
Veröffentlicht: (2025)
IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2024)
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2024)
Extreme Miscalibration and the Illusion of Adversarial Robustness
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2025)
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2025)
The Illusion of State in State-Space Models
von: Merrill, William, et al.
Veröffentlicht: (2024)
von: Merrill, William, et al.
Veröffentlicht: (2024)
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
Disagreement as Data: Reasoning Trace Analytics in Multi-Agent Systems
von: Tajik, Elham, et al.
Veröffentlicht: (2026)
von: Tajik, Elham, et al.
Veröffentlicht: (2026)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2026)
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2026)
Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
von: Xu, Yinuo, et al.
Veröffentlicht: (2025)
von: Xu, Yinuo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs
von: Allen, Bradley P., et al.
Veröffentlicht: (2025) -
The Illusion of Stochasticity in LLMs
von: Gu, Xiangming, et al.
Veröffentlicht: (2026) -
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
von: Andreev, Nikita, et al.
Veröffentlicht: (2024) -
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
von: Lu, Junyu, et al.
Veröffentlicht: (2025) -
When Disagreements Elicit Robustness: Investigating Self-Repair Capabilities under LLM Multi-Agent Disagreements
von: Ju, Tianjie, et al.
Veröffentlicht: (2025)