Self-Correcting Large Language Models: Generation vs. Multiple Choice
Fuente:
arXiv
Saved in:
| Main Authors: | Rahmani, Hossein A., Krishna, Satyapriya, Wang, Xi, Naghiaei, Mohammadmehdi, Yilmaz, Emine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Personalized Framework for Consumer and Producer Group Fairness Optimization in Recommender Systems
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
Understanding the Role of User Profile in the Personalization of Large Language Models
by: Wu, Bin, et al.
Published: (2024)
by: Wu, Bin, et al.
Published: (2024)
Beyond Output Critique: Self-Correction via Task Distillation
by: Rahmani, Hossein A., et al.
Published: (2026)
by: Rahmani, Hossein A., et al.
Published: (2026)
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
by: Chen, Yinzhu, et al.
Published: (2026)
by: Chen, Yinzhu, et al.
Published: (2026)
Clarifying the Path to User Satisfaction: An Investigation into Clarification Usefulness
by: Rahmani, Hossein A., et al.
Published: (2024)
by: Rahmani, Hossein A., et al.
Published: (2024)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
by: Fu, Xiao, et al.
Published: (2025)
by: Fu, Xiao, et al.
Published: (2025)
Multiple-Choice Question Generation Using Large Language Models: Methodology and Educator Insights
by: Biancini, Giorgio, et al.
Published: (2025)
by: Biancini, Giorgio, et al.
Published: (2025)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
by: Li, Aaron J., et al.
Published: (2024)
by: Li, Aaron J., et al.
Published: (2024)
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback
by: Yao, Zonghai, et al.
Published: (2024)
by: Yao, Zonghai, et al.
Published: (2024)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
by: Fu, Tairan, et al.
Published: (2025)
by: Fu, Tairan, et al.
Published: (2025)
Balancing Rigor and Utility: Mitigating Cognitive Biases in Large Language Models for Multiple-Choice Questions
by: Zhong, Hanyang, et al.
Published: (2024)
by: Zhong, Hanyang, et al.
Published: (2024)
Cognitively Diverse Multiple-Choice Question Generation: A Hybrid Multi-Agent Framework with Large Language Models
by: Tian, Yu, et al.
Published: (2026)
by: Tian, Yu, et al.
Published: (2026)
Incorporating Error Level Noise Embedding for Improving LLM-Assisted Robustness in Persian Speech Recognition
by: Rahmani, Zahra, et al.
Published: (2025)
by: Rahmani, Zahra, et al.
Published: (2025)
Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
by: Schumann, Raphael, et al.
Published: (2025)
by: Schumann, Raphael, et al.
Published: (2025)
Overview of the TREC 2023 deep learning track
by: Craswell, Nick, et al.
Published: (2025)
by: Craswell, Nick, et al.
Published: (2025)
Exploring Iterative Enhancement for Improving Learnersourced Multiple-Choice Question Explanations with Large Language Models
by: Bao, Qiming, et al.
Published: (2023)
by: Bao, Qiming, et al.
Published: (2023)
RECALL-MM: A Multimodal Dataset of Consumer Product Recalls for Risk Analysis using Computational Methods and Large Language Models
by: Bolanos, Diana, et al.
Published: (2025)
by: Bolanos, Diana, et al.
Published: (2025)
Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
by: Liu, Yesheng, et al.
Published: (2025)
by: Liu, Yesheng, et al.
Published: (2025)
Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation
by: Li, Ruizhe, et al.
Published: (2025)
by: Li, Ruizhe, et al.
Published: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation
by: Doris, Anna C., et al.
Published: (2024)
by: Doris, Anna C., et al.
Published: (2024)
Large Language Models Cannot Self-Correct Reasoning Yet
by: Huang, Jie, et al.
Published: (2023)
by: Huang, Jie, et al.
Published: (2023)
Large Language Models have Intrinsic Self-Correction Ability
by: Liu, Dancheng, et al.
Published: (2024)
by: Liu, Dancheng, et al.
Published: (2024)
A Study on Large Language Models' Limitations in Multiple-Choice Question Answering
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
by: Deng, Wenqing, et al.
Published: (2024)
by: Deng, Wenqing, et al.
Published: (2024)
Generating Multiple-Choice Knowledge Questions with Interpretable Difficulty Estimation using Knowledge Graphs and Large Language Models
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models
by: Zhang, Che, et al.
Published: (2024)
by: Zhang, Che, et al.
Published: (2024)
MegaChat: A Synthetic Persian Q&A Dataset for High-Quality Sales Chatbot Evaluation
by: Rahmani, Mahdi, et al.
Published: (2025)
by: Rahmani, Mahdi, et al.
Published: (2025)
CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Models
by: Wu, Yuning, et al.
Published: (2026)
by: Wu, Yuning, et al.
Published: (2026)
Quantifying Label-Induced Bias in Large Language Model Self- and Cross-Evaluations
by: Saraf, Muskan, et al.
Published: (2025)
by: Saraf, Muskan, et al.
Published: (2025)
Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models
by: Li, Loka, et al.
Published: (2024)
by: Li, Loka, et al.
Published: (2024)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
Skewed Memorization in Large Language Models: Quantification and Decomposition
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
by: Letzelter, Victor, et al.
Published: (2025)
by: Letzelter, Victor, et al.
Published: (2025)
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
by: Tsui, Ken
Published: (2025)
by: Tsui, Ken
Published: (2025)
Large Language Models for Multi-Choice Question Classification of Medical Subjects
by: Ponce-López, Víctor
Published: (2024)
by: Ponce-López, Víctor
Published: (2024)
Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Similar Items
-
A Personalized Framework for Consumer and Producer Group Fairness Optimization in Recommender Systems
by: Rahmani, Hossein A., et al.
Published: (2024) -
Understanding the Role of User Profile in the Personalization of Large Language Models
by: Wu, Bin, et al.
Published: (2024) -
Beyond Output Critique: Self-Correction via Task Distillation
by: Rahmani, Hossein A., et al.
Published: (2026) -
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
by: Chen, Yinzhu, et al.
Published: (2026) -
Clarifying the Path to User Satisfaction: An Investigation into Clarification Usefulness
by: Rahmani, Hossein A., et al.
Published: (2024)