None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Tam, Zhi Rui, Wu, Cheng-Kuang, Lin, Chieh-Yen, Chen, Yun-Nung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
by: Wu, Cheng-Kuang, et al.
Published: (2024)
by: Wu, Cheng-Kuang, et al.
Published: (2024)
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
by: Wu, Cheng-Kuang, et al.
Published: (2025)
by: Wu, Cheng-Kuang, et al.
Published: (2025)
Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
by: Tam, Zhi Rui, et al.
Published: (2024)
by: Tam, Zhi Rui, et al.
Published: (2024)
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
by: Tam, Zhi Rui, et al.
Published: (2025)
by: Tam, Zhi Rui, et al.
Published: (2025)
I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL Generation
by: Wu, Cheng-Kuang, et al.
Published: (2024)
by: Wu, Cheng-Kuang, et al.
Published: (2024)
MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
by: Tam, Zhi Rui, et al.
Published: (2025)
by: Tam, Zhi Rui, et al.
Published: (2025)
None of the Above
by: Cohen, Mollie
Published: (2024)
by: Cohen, Mollie
Published: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
by: Haider, Batool, et al.
Published: (2025)
by: Haider, Batool, et al.
Published: (2025)
The Role of Exploration Modules in Small Language Models for Knowledge Graph Question Answering
by: Cheng, Yi-Jie, et al.
Published: (2025)
by: Cheng, Yi-Jie, et al.
Published: (2025)
VisTW: Benchmarking Vision-Language Models for Traditional Chinese in Taiwan
by: Tam, Zhi Rui, et al.
Published: (2025)
by: Tam, Zhi Rui, et al.
Published: (2025)
EduAgentQG: A Multi-Agent Workflow Framework for Personalized Question Generation
by: Jia, Rui, et al.
Published: (2025)
by: Jia, Rui, et al.
Published: (2025)
Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning
by: Wu, Chao-Chung, et al.
Published: (2025)
by: Wu, Chao-Chung, et al.
Published: (2025)
One Size Fits None: Rethinking Fairness in Medical AI
by: Roller, Roland, et al.
Published: (2025)
by: Roller, Roland, et al.
Published: (2025)
Towards Unsupervised Question Answering System with Multi-level Summarization for Legal Text
by: Prabhu, M Manvith, et al.
Published: (2024)
by: Prabhu, M Manvith, et al.
Published: (2024)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
by: Mohamadi, Alireza, et al.
Published: (2025)
by: Mohamadi, Alireza, et al.
Published: (2025)
One Size Fits None: A Personalized Framework for Urban Accessibility Using Exponential Decay
by: Ghuriki, Prabhanjana, et al.
Published: (2025)
by: Ghuriki, Prabhanjana, et al.
Published: (2025)
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models
by: Bae, Suyoung, et al.
Published: (2025)
by: Bae, Suyoung, et al.
Published: (2025)
Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice?
by: Tak, Ala N., et al.
Published: (2026)
by: Tak, Ala N., et al.
Published: (2026)
Social Bias in Popular Question-Answering Benchmarks
by: Kraft, Angelie, et al.
Published: (2025)
by: Kraft, Angelie, et al.
Published: (2025)
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
by: Bahaj, Adil, et al.
Published: (2025)
by: Bahaj, Adil, et al.
Published: (2025)
Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception
by: Lin, Luyang, et al.
Published: (2024)
by: Lin, Luyang, et al.
Published: (2024)
CAIRNS: Balancing Readability and Scientific Accuracy in Climate Adaptation Question Answering
by: Kong, Liangji, et al.
Published: (2025)
by: Kong, Liangji, et al.
Published: (2025)
Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models
by: Zhou, Ke, et al.
Published: (2025)
by: Zhou, Ke, et al.
Published: (2025)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
by: Singh, Shrutika, et al.
Published: (2025)
by: Singh, Shrutika, et al.
Published: (2025)
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language
by: Sammoudi, Mohammad, et al.
Published: (2024)
by: Sammoudi, Mohammad, et al.
Published: (2024)
SUKHSANDESH: An Avatar Therapeutic Question Answering Platform for Sexual Education in Rural India
by: Singh, Salam Michael, et al.
Published: (2024)
by: Singh, Salam Michael, et al.
Published: (2024)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
Human Rights for the Digital Age
by: Siddiqui, Shaleeza Yaqoob, et al.
Published: (2024)
by: Siddiqui, Shaleeza Yaqoob, et al.
Published: (2024)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
by: Hu, Zhanghao, et al.
Published: (2025)
by: Hu, Zhanghao, et al.
Published: (2025)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Technological Progress and Obsolescence: Analyzing the Environmental Economic Impacts of MacBook Pro I/O Devices
by: Cheng, Yun-Chieh, et al.
Published: (2024)
by: Cheng, Yun-Chieh, et al.
Published: (2024)
GG-BBQ: German Gender Bias Benchmark for Question Answering
by: Satheesh, Shalaka, et al.
Published: (2025)
by: Satheesh, Shalaka, et al.
Published: (2025)
LLMs Provide Unstable Answers to Legal Questions
by: Blair-Stanek, Andrew, et al.
Published: (2025)
by: Blair-Stanek, Andrew, et al.
Published: (2025)
More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
by: Vu, Duc Anh, et al.
Published: (2025)
by: Vu, Duc Anh, et al.
Published: (2025)
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
by: Javed, Rafiya, et al.
Published: (2025)
by: Javed, Rafiya, et al.
Published: (2025)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
by: Wu, Xuansheng, et al.
Published: (2024)
by: Wu, Xuansheng, et al.
Published: (2024)
Ontology-Aware RAG for Improved Question-Answering in Cybersecurity Education
by: Zhao, Chengshuai, et al.
Published: (2024)
by: Zhao, Chengshuai, et al.
Published: (2024)
CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation
by: Duan, Xiaojing, et al.
Published: (2026)
by: Duan, Xiaojing, et al.
Published: (2026)
Answering Students' Questions on Course Forums Using Multiple Chain-of-Thought Reasoning and Finetuning RAG-Enabled LLM
by: Wang, Neo, et al.
Published: (2025)
by: Wang, Neo, et al.
Published: (2025)
Similar Items
-
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026) -
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
by: Wu, Cheng-Kuang, et al.
Published: (2024) -
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
by: Wu, Cheng-Kuang, et al.
Published: (2025) -
Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
by: Tam, Zhi Rui, et al.
Published: (2024) -
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
by: Tam, Zhi Rui, et al.
Published: (2025)