The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pinhanez, Claudio, Cavalin, Paulo, Sanctos, Cassia, Grave, Marcelo, Primerano, Yago |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
von: Gonçalves, Isabel, et al.
Veröffentlicht: (2025)
von: Gonçalves, Isabel, et al.
Veröffentlicht: (2025)
Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level Aggregation
von: Cavalin, Paulo, et al.
Veröffentlicht: (2024)
von: Cavalin, Paulo, et al.
Veröffentlicht: (2024)
Creating an African American-Sounding TTS: Guidelines, Technical Challenges,and Surprising Evaluations
von: Pinhanez, Claudio, et al.
Veröffentlicht: (2024)
von: Pinhanez, Claudio, et al.
Veröffentlicht: (2024)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
von: Zhang, Yushun, et al.
Veröffentlicht: (2026)
von: Zhang, Yushun, et al.
Veröffentlicht: (2026)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
Harnessing the Power of Artificial Intelligence to Vitalize Endangered Indigenous Languages: Technologies and Experiences
von: Pinhanez, Claudio, et al.
Veröffentlicht: (2024)
von: Pinhanez, Claudio, et al.
Veröffentlicht: (2024)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
von: Labat, Léo, et al.
Veröffentlicht: (2026)
von: Labat, Léo, et al.
Veröffentlicht: (2026)
Prompt Repetition Improves Non-Reasoning LLMs
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2025)
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
Prompt Sensitivity and Answer Consistency of Small Open-Source Language Models for Clinical Question Answering in Low-Resource Healthcare
von: Hariprasad, Shravani
Veröffentlicht: (2026)
von: Hariprasad, Shravani
Veröffentlicht: (2026)
A methodological analysis of prompt perturbations and their effect on attack success rates
von: Machado, Tiago, et al.
Veröffentlicht: (2025)
von: Machado, Tiago, et al.
Veröffentlicht: (2025)
Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation
von: Tang, Xuemei, et al.
Veröffentlicht: (2026)
von: Tang, Xuemei, et al.
Veröffentlicht: (2026)
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
von: Park, Yoonah, et al.
Veröffentlicht: (2025)
von: Park, Yoonah, et al.
Veröffentlicht: (2025)
LLMs as Agentic Cooperative Players in Multiplayer UNO
von: Matinez, Yago Romano, et al.
Veröffentlicht: (2025)
von: Matinez, Yago Romano, et al.
Veröffentlicht: (2025)
Benchmarking LLMs' Judgments with No Gold Standard
von: Xu, Shengwei, et al.
Veröffentlicht: (2024)
von: Xu, Shengwei, et al.
Veröffentlicht: (2024)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
UBench: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions
von: Wang, Xunzhi, et al.
Veröffentlicht: (2024)
von: Wang, Xunzhi, et al.
Veröffentlicht: (2024)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025)
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025)
Student Answer Forecasting: Transformer-Driven Answer Choice Prediction for Language Learning
von: Gado, Elena Grazia, et al.
Veröffentlicht: (2024)
von: Gado, Elena Grazia, et al.
Veröffentlicht: (2024)
Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options
von: Góral, Gracjan, et al.
Veröffentlicht: (2024)
von: Góral, Gracjan, et al.
Veröffentlicht: (2024)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
von: Schlicht, Ipek Baris, et al.
Veröffentlicht: (2025)
von: Schlicht, Ipek Baris, et al.
Veröffentlicht: (2025)
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
von: Vico, Gianluca, et al.
Veröffentlicht: (2026)
von: Vico, Gianluca, et al.
Veröffentlicht: (2026)
Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs
von: Sanz-Guerrero, Mario, et al.
Veröffentlicht: (2025)
von: Sanz-Guerrero, Mario, et al.
Veröffentlicht: (2025)
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
von: Yang, Running, et al.
Veröffentlicht: (2025)
von: Yang, Running, et al.
Veröffentlicht: (2025)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
von: Letzelter, Victor, et al.
Veröffentlicht: (2025)
von: Letzelter, Victor, et al.
Veröffentlicht: (2025)
Learning to Generate Answers with Citations via Factual Consistency Models
von: Aly, Rami, et al.
Veröffentlicht: (2024)
von: Aly, Rami, et al.
Veröffentlicht: (2024)
WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More Challenging
von: Elhady, Ahmed, et al.
Veröffentlicht: (2025)
von: Elhady, Ahmed, et al.
Veröffentlicht: (2025)
Representation Consistency for Accurate and Coherent LLM Answer Aggregation
von: Jiang, Junqi, et al.
Veröffentlicht: (2025)
von: Jiang, Junqi, et al.
Veröffentlicht: (2025)
Rethinking Repetition Problems of LLMs in Code Generation
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
Recursive Think-Answer Process for LLMs and VLMs
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2026)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2026)
Repetition Neurons: How Do Language Models Produce Repetitions?
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
von: Deng, Wenqing, et al.
Veröffentlicht: (2024)
von: Deng, Wenqing, et al.
Veröffentlicht: (2024)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2026)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025) -
CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025) -
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
von: Gonçalves, Isabel, et al.
Veröffentlicht: (2025) -
Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level Aggregation
von: Cavalin, Paulo, et al.
Veröffentlicht: (2024) -
Creating an African American-Sounding TTS: Guidelines, Technical Challenges,and Surprising Evaluations
von: Pinhanez, Claudio, et al.
Veröffentlicht: (2024)