Can multiple-choice questions really be useful in detecting the abilities of LLMs?
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Wangyue, Li, Liangzhi, Xiang, Tong, Liu, Xiao, Deng, Wei, Garcia, Noa |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Cracking the Code: Enhancing Implicit Hate Speech Detection through Coding Classification
by: Wei, Lu, et al.
Published: (2025)
by: Wei, Lu, et al.
Published: (2025)
Unsupervised multiple choices question answering via universal corpus
by: Zhang, Qin, et al.
Published: (2024)
by: Zhang, Qin, et al.
Published: (2024)
"I know myself better, but not really greatly": How Well Can LLMs Detect and Explain LLM-Generated Texts?
by: Ji, Jiazhou, et al.
Published: (2025)
by: Ji, Jiazhou, et al.
Published: (2025)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning
by: Tong, Yongqi, et al.
Published: (2024)
by: Tong, Yongqi, et al.
Published: (2024)
Can LLMs Solve longer Math Word Problems Better?
by: Xu, Xin, et al.
Published: (2024)
by: Xu, Xin, et al.
Published: (2024)
Why are LLMs' abilities emergent?
by: Havlík, Vladimír
Published: (2025)
by: Havlík, Vladimír
Published: (2025)
Lifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-Experts
by: Chen, Qizhou, et al.
Published: (2024)
by: Chen, Qizhou, et al.
Published: (2024)
Can Multi-modal (reasoning) LLMs detect document manipulation?
by: Liang, Zisheng, et al.
Published: (2025)
by: Liang, Zisheng, et al.
Published: (2025)
Can the capability of Large Language Models be described by human ability? A Meta Study
by: Zan, Mingrui, et al.
Published: (2025)
by: Zan, Mingrui, et al.
Published: (2025)
BlockPruner: Fine-grained Pruning for Large Language Models
by: Zhong, Longguang, et al.
Published: (2024)
by: Zhong, Longguang, et al.
Published: (2024)
DCR: Divide-and-Conquer Reasoning for Multi-choice Question Answering with LLMs
by: Meng, Zijie, et al.
Published: (2024)
by: Meng, Zijie, et al.
Published: (2024)
Can AI mimic the human ability to define neologisms?
by: Georgiou, Georgios P.
Published: (2025)
by: Georgiou, Georgios P.
Published: (2025)
Can LLMs Extract Frame-Semantic Arguments?
by: Devasier, Jacob, et al.
Published: (2025)
by: Devasier, Jacob, et al.
Published: (2025)
Can Editing LLMs Inject Harm?
by: Chen, Canyu, et al.
Published: (2024)
by: Chen, Canyu, et al.
Published: (2024)
Can LLMs Outshine Conventional Recommenders? A Comparative Evaluation
by: Liu, Qijiong, et al.
Published: (2025)
by: Liu, Qijiong, et al.
Published: (2025)
Large Linguistic Models: Investigating LLMs' metalinguistic abilities
by: Beguš, Gašper, et al.
Published: (2023)
by: Beguš, Gašper, et al.
Published: (2023)
Do prompt positions really matter?
by: Mao, Junyu, et al.
Published: (2023)
by: Mao, Junyu, et al.
Published: (2023)
Transfer Learning Enhanced Single-choice Decision for Multi-choice Question Answering
by: Cui, Chenhao, et al.
Published: (2024)
by: Cui, Chenhao, et al.
Published: (2024)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
by: Xiong, Miao, et al.
Published: (2023)
by: Xiong, Miao, et al.
Published: (2023)
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
by: Zhang, Qinyan, et al.
Published: (2025)
by: Zhang, Qinyan, et al.
Published: (2025)
LLMs Know More About Numbers than They Can Say
by: Yuchi, Fengting, et al.
Published: (2026)
by: Yuchi, Fengting, et al.
Published: (2026)
Beyond the Link: Assessing LLMs' ability to Classify Political Content across Global Media
by: De La Fuente-Cuesta, Alejandro, et al.
Published: (2025)
by: De La Fuente-Cuesta, Alejandro, et al.
Published: (2025)
Can LLMs Improve Multimodal Fact-Checking by Asking Relevant Questions?
by: Beigi, Alimohammad, et al.
Published: (2024)
by: Beigi, Alimohammad, et al.
Published: (2024)
Beyond Decoder-only: Large Language Models Can be Good Encoders for Machine Translation
by: Luo, Yingfeng, et al.
Published: (2025)
by: Luo, Yingfeng, et al.
Published: (2025)
Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
by: Hegde, Pradyoth, et al.
Published: (2025)
by: Hegde, Pradyoth, et al.
Published: (2025)
LLMs Can Generate a Better Answer by Aggregating Their Own Responses
by: Li, Zichong, et al.
Published: (2025)
by: Li, Zichong, et al.
Published: (2025)
Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
by: Gao, Lirong, et al.
Published: (2026)
by: Gao, Lirong, et al.
Published: (2026)
PACIFIC: Can LLMs Discern the Traits Influencing Your Preferences? Evaluating Personality-Driven Preference Alignment in LLMs
by: Zhao, Tianyu, et al.
Published: (2026)
by: Zhao, Tianyu, et al.
Published: (2026)
LLMs Can Also Do Well! Breaking Barriers in Semantic Role Labeling via Large Language Models
by: Li, Xinxin, et al.
Published: (2025)
by: Li, Xinxin, et al.
Published: (2025)
Can LLMs Learn to Map the World from Local Descriptions?
by: Xia, Sirui, et al.
Published: (2025)
by: Xia, Sirui, et al.
Published: (2025)
On the effectiveness of LLMs for automatic grading of open-ended questions in Spanish
by: Capdehourat, Germán, et al.
Published: (2025)
by: Capdehourat, Germán, et al.
Published: (2025)
SepSeq: A Training-Free Framework for Long Numerical Sequence Processing in LLMs
by: Sun, Jie, et al.
Published: (2026)
by: Sun, Jie, et al.
Published: (2026)
Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs
by: Xia, Yu, et al.
Published: (2024)
by: Xia, Yu, et al.
Published: (2024)
Self-evolving Agents with reflective and memory-augmented abilities
by: Liang, Xuechen, et al.
Published: (2024)
by: Liang, Xuechen, et al.
Published: (2024)
Can LLMs Generate High-Quality Task-Specific Conversations?
by: Li, Shengqi, et al.
Published: (2025)
by: Li, Shengqi, et al.
Published: (2025)
Predictions from language models for multiple-choice tasks are not robust under variation of scoring methods
by: Tsvilodub, Polina, et al.
Published: (2024)
by: Tsvilodub, Polina, et al.
Published: (2024)
Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective
by: Mou, Yutao, et al.
Published: (2025)
by: Mou, Yutao, et al.
Published: (2025)
Similar Items
-
Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models
by: Liu, Xiao, et al.
Published: (2024) -
Cracking the Code: Enhancing Implicit Hate Speech Detection through Coding Classification
by: Wei, Lu, et al.
Published: (2025) -
Unsupervised multiple choices question answering via universal corpus
by: Zhang, Qin, et al.
Published: (2024) -
"I know myself better, but not really greatly": How Well Can LLMs Detect and Explain LLM-Generated Texts?
by: Ji, Jiazhou, et al.
Published: (2025) -
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
by: Li, Belinda Z., et al.
Published: (2025)