Predictions from language models for multiple-choice tasks are not robust under variation of scoring methods
Fuente:
arXiv
Saved in:
| Main Authors: | Tsvilodub, Polina, Wang, Hening, Grosch, Sharon, Franke, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bayesian Statistical Modeling with Predictors from LLMs
by: Franke, Michael, et al.
Published: (2024)
by: Franke, Michael, et al.
Published: (2024)
Cognitive Modeling with Scaffolded LLMs: A Case Study of Referential Expression Generation
by: Tsvilodub, Polina, et al.
Published: (2024)
by: Tsvilodub, Polina, et al.
Published: (2024)
Integrating Neural and Symbolic Components in a Model of Pragmatic Question-Answering
by: Tsvilodub, Polina, et al.
Published: (2025)
by: Tsvilodub, Polina, et al.
Published: (2025)
Experimental Pragmatics with Machines: Testing LLM Predictions for the Inferences of Plain and Embedded Disjunctions
by: Tsvilodub, Polina, et al.
Published: (2024)
by: Tsvilodub, Polina, et al.
Published: (2024)
Act or Clarify? Modeling Sensitivity to Uncertainty and Cost in Communication
by: Tsvilodub, Polina, et al.
Published: (2026)
by: Tsvilodub, Polina, et al.
Published: (2026)
On Emergent Social World Models -- Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models
by: Tsvilodub, Polina, et al.
Published: (2026)
by: Tsvilodub, Polina, et al.
Published: (2026)
Non-literal Understanding of Number Words by Language Models
by: Tsvilodub, Polina, et al.
Published: (2025)
by: Tsvilodub, Polina, et al.
Published: (2025)
Robustness assessment of large audio language models in multiple-choice evaluation
by: López, Fernando, et al.
Published: (2025)
by: López, Fernando, et al.
Published: (2025)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
by: Ilia, Evgenia, et al.
Published: (2024)
by: Ilia, Evgenia, et al.
Published: (2024)
Evaluating language models as risk scores
by: Cruz, André F., et al.
Published: (2024)
by: Cruz, André F., et al.
Published: (2024)
Auxiliary task demands mask the capabilities of smaller language models
by: Hu, Jennifer, et al.
Published: (2024)
by: Hu, Jennifer, et al.
Published: (2024)
Do different prompting methods yield a common task representation in language models?
by: Davidson, Guy, et al.
Published: (2025)
by: Davidson, Guy, et al.
Published: (2025)
Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning
by: Netík, Jan, et al.
Published: (2026)
by: Netík, Jan, et al.
Published: (2026)
Just-in-time and distributed task representations in language models
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Testing AI on language comprehension tasks reveals insensitivity to underlying meaning
by: Dentella, Vittoria, et al.
Published: (2023)
by: Dentella, Vittoria, et al.
Published: (2023)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
Code-enabled language models can outperform reasoning models on diverse tasks
by: Zhang, Cedegao E., et al.
Published: (2025)
by: Zhang, Cedegao E., et al.
Published: (2025)
Markovian ODE-guided scoring can assess the quality of offline reasoning traces in language models
by: Nandi, Arghodeep, et al.
Published: (2026)
by: Nandi, Arghodeep, et al.
Published: (2026)
Evidence from counterfactual tasks supports emergent analogical reasoning in large language models
by: Webb, Taylor, et al.
Published: (2024)
by: Webb, Taylor, et al.
Published: (2024)
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
Multilingual BERT language model for medical tasks: Evaluation on domain-specific adaptation and cross-linguality
by: Luo, Yinghao, et al.
Published: (2025)
by: Luo, Yinghao, et al.
Published: (2025)
Superhuman performance of a large language model on the reasoning tasks of a physician
by: Brodeur, Peter G., et al.
Published: (2024)
by: Brodeur, Peter G., et al.
Published: (2024)
Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models
by: Lyu, Y., et al.
Published: (2025)
by: Lyu, Y., et al.
Published: (2025)
SemPool: Simple, robust, and interpretable KG pooling for enhancing language models
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
Unsupervised multiple choices question answering via universal corpus
by: Zhang, Qin, et al.
Published: (2024)
by: Zhang, Qin, et al.
Published: (2024)
Red and blue language: Word choices in the Trump & Harris 2024 presidential debate
by: Wicke, Philipp, et al.
Published: (2024)
by: Wicke, Philipp, et al.
Published: (2024)
Scaling behavior of large language models in emotional safety classification across sizes and tasks
by: Pinzuti, Edoardo, et al.
Published: (2025)
by: Pinzuti, Edoardo, et al.
Published: (2025)
Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks
by: Xian, R. Patrick, et al.
Published: (2024)
by: Xian, R. Patrick, et al.
Published: (2024)
A validity-guided workflow for robust large language model research in psychology
by: Lin, Zhicheng
Published: (2025)
by: Lin, Zhicheng
Published: (2025)
Can we teach language models to gloss endangered languages?
by: Ginn, Michael, et al.
Published: (2024)
by: Ginn, Michael, et al.
Published: (2024)
Can multiple-choice questions really be useful in detecting the abilities of LLMs?
by: Li, Wangyue, et al.
Published: (2024)
by: Li, Wangyue, et al.
Published: (2024)
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
by: Thelwall, Mike
Published: (2026)
by: Thelwall, Mike
Published: (2026)
TechGPT-2.0: A large language model project to solve the task of knowledge graph construction
by: Wang, Jiaqi, et al.
Published: (2024)
by: Wang, Jiaqi, et al.
Published: (2024)
Can large language models interpret unstructured chat data on dynamic group decision-making processes? Evidence on joint destination choice
by: Lim, Sung-Yoo, et al.
Published: (2026)
by: Lim, Sung-Yoo, et al.
Published: (2026)
Towards a cognitive architecture to enable natural language interaction in co-constructive task learning
by: Scheibl, Manuel, et al.
Published: (2025)
by: Scheibl, Manuel, et al.
Published: (2025)
Transformers need glasses! Information over-squashing in language tasks
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Retrieval augmentation of large language models for lay language generation
by: Guo, Yue, et al.
Published: (2022)
by: Guo, Yue, et al.
Published: (2022)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Aviary: training language agents on challenging scientific tasks
by: Narayanan, Siddharth, et al.
Published: (2024)
by: Narayanan, Siddharth, et al.
Published: (2024)
Do large language models resemble humans in language use?
by: Cai, Zhenguang G., et al.
Published: (2023)
by: Cai, Zhenguang G., et al.
Published: (2023)
Similar Items
-
Bayesian Statistical Modeling with Predictors from LLMs
by: Franke, Michael, et al.
Published: (2024) -
Cognitive Modeling with Scaffolded LLMs: A Case Study of Referential Expression Generation
by: Tsvilodub, Polina, et al.
Published: (2024) -
Integrating Neural and Symbolic Components in a Model of Pragmatic Question-Answering
by: Tsvilodub, Polina, et al.
Published: (2025) -
Experimental Pragmatics with Machines: Testing LLM Predictions for the Inferences of Plain and Embedded Disjunctions
by: Tsvilodub, Polina, et al.
Published: (2024) -
Act or Clarify? Modeling Sensitivity to Uncertainty and Cost in Communication
by: Tsvilodub, Polina, et al.
Published: (2026)