The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | Pinhanez, Claudio, Cavalin, Paulo, Sanctos, Cassia, Grave, Marcelo, Primerano, Yago |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
por: Cavalin, Paulo, et al.
Publicado: (2025)
por: Cavalin, Paulo, et al.
Publicado: (2025)
CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
por: Cavalin, Paulo, et al.
Publicado: (2025)
por: Cavalin, Paulo, et al.
Publicado: (2025)
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
por: Gonçalves, Isabel, et al.
Publicado: (2025)
por: Gonçalves, Isabel, et al.
Publicado: (2025)
Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level Aggregation
por: Cavalin, Paulo, et al.
Publicado: (2024)
por: Cavalin, Paulo, et al.
Publicado: (2024)
Creating an African American-Sounding TTS: Guidelines, Technical Challenges,and Surprising Evaluations
por: Pinhanez, Claudio, et al.
Publicado: (2024)
por: Pinhanez, Claudio, et al.
Publicado: (2024)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
por: Balepur, Nishant, et al.
Publicado: (2024)
por: Balepur, Nishant, et al.
Publicado: (2024)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
por: Zhang, Yushun, et al.
Publicado: (2026)
por: Zhang, Yushun, et al.
Publicado: (2026)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
por: Wiegreffe, Sarah, et al.
Publicado: (2024)
por: Wiegreffe, Sarah, et al.
Publicado: (2024)
Harnessing the Power of Artificial Intelligence to Vitalize Endangered Indigenous Languages: Technologies and Experiences
por: Pinhanez, Claudio, et al.
Publicado: (2024)
por: Pinhanez, Claudio, et al.
Publicado: (2024)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
por: Chandak, Nikhil, et al.
Publicado: (2025)
por: Chandak, Nikhil, et al.
Publicado: (2025)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
por: Labat, Léo, et al.
Publicado: (2026)
por: Labat, Léo, et al.
Publicado: (2026)
Prompt Repetition Improves Non-Reasoning LLMs
por: Leviathan, Yaniv, et al.
Publicado: (2025)
por: Leviathan, Yaniv, et al.
Publicado: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
por: Li, Ruizhe, et al.
Publicado: (2024)
por: Li, Ruizhe, et al.
Publicado: (2024)
Prompt Sensitivity and Answer Consistency of Small Open-Source Language Models for Clinical Question Answering in Low-Resource Healthcare
por: Hariprasad, Shravani
Publicado: (2026)
por: Hariprasad, Shravani
Publicado: (2026)
A methodological analysis of prompt perturbations and their effect on attack success rates
por: Machado, Tiago, et al.
Publicado: (2025)
por: Machado, Tiago, et al.
Publicado: (2025)
Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation
por: Tang, Xuemei, et al.
Publicado: (2026)
por: Tang, Xuemei, et al.
Publicado: (2026)
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
por: Park, Yoonah, et al.
Publicado: (2025)
por: Park, Yoonah, et al.
Publicado: (2025)
LLMs as Agentic Cooperative Players in Multiplayer UNO
por: Matinez, Yago Romano, et al.
Publicado: (2025)
por: Matinez, Yago Romano, et al.
Publicado: (2025)
Benchmarking LLMs' Judgments with No Gold Standard
por: Xu, Shengwei, et al.
Publicado: (2024)
por: Xu, Shengwei, et al.
Publicado: (2024)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
por: Mustapha, Ahmad, et al.
Publicado: (2024)
por: Mustapha, Ahmad, et al.
Publicado: (2024)
UBench: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions
por: Wang, Xunzhi, et al.
Publicado: (2024)
por: Wang, Xunzhi, et al.
Publicado: (2024)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
por: Hidayat, Naila Shafirni, et al.
Publicado: (2025)
por: Hidayat, Naila Shafirni, et al.
Publicado: (2025)
Student Answer Forecasting: Transformer-Driven Answer Choice Prediction for Language Learning
por: Gado, Elena Grazia, et al.
Publicado: (2024)
por: Gado, Elena Grazia, et al.
Publicado: (2024)
Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options
por: Góral, Gracjan, et al.
Publicado: (2024)
por: Góral, Gracjan, et al.
Publicado: (2024)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
por: Palta, Shramay, et al.
Publicado: (2024)
por: Palta, Shramay, et al.
Publicado: (2024)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
por: Balepur, Nishant, et al.
Publicado: (2026)
por: Balepur, Nishant, et al.
Publicado: (2026)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
por: Schlicht, Ipek Baris, et al.
Publicado: (2025)
por: Schlicht, Ipek Baris, et al.
Publicado: (2025)
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
por: Vico, Gianluca, et al.
Publicado: (2026)
por: Vico, Gianluca, et al.
Publicado: (2026)
Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs
por: Sanz-Guerrero, Mario, et al.
Publicado: (2025)
por: Sanz-Guerrero, Mario, et al.
Publicado: (2025)
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
por: Yang, Running, et al.
Publicado: (2025)
por: Yang, Running, et al.
Publicado: (2025)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
por: Letzelter, Victor, et al.
Publicado: (2025)
por: Letzelter, Victor, et al.
Publicado: (2025)
Learning to Generate Answers with Citations via Factual Consistency Models
por: Aly, Rami, et al.
Publicado: (2024)
por: Aly, Rami, et al.
Publicado: (2024)
WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More Challenging
por: Elhady, Ahmed, et al.
Publicado: (2025)
por: Elhady, Ahmed, et al.
Publicado: (2025)
Representation Consistency for Accurate and Coherent LLM Answer Aggregation
por: Jiang, Junqi, et al.
Publicado: (2025)
por: Jiang, Junqi, et al.
Publicado: (2025)
Rethinking Repetition Problems of LLMs in Code Generation
por: Dong, Yihong, et al.
Publicado: (2025)
por: Dong, Yihong, et al.
Publicado: (2025)
Recursive Think-Answer Process for LLMs and VLMs
por: Lee, Byung-Kwan, et al.
Publicado: (2026)
por: Lee, Byung-Kwan, et al.
Publicado: (2026)
Repetition Neurons: How Do Language Models Produce Repetitions?
por: Hiraoka, Tatsuya, et al.
Publicado: (2024)
por: Hiraoka, Tatsuya, et al.
Publicado: (2024)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
por: Deng, Wenqing, et al.
Publicado: (2024)
por: Deng, Wenqing, et al.
Publicado: (2024)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2026)
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2026)
Ejemplares similares
-
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
por: Cavalin, Paulo, et al.
Publicado: (2025) -
CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
por: Cavalin, Paulo, et al.
Publicado: (2025) -
Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
por: Gonçalves, Isabel, et al.
Publicado: (2025) -
Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level Aggregation
por: Cavalin, Paulo, et al.
Publicado: (2024) -
Creating an African American-Sounding TTS: Guidelines, Technical Challenges,and Surprising Evaluations
por: Pinhanez, Claudio, et al.
Publicado: (2024)