Feedback Indices to Evaluate LLM Responses to Rebuttals for Multiple Choice Type Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Dunlap, Justin C., Parent, Anne-Simone, Widenhorn, Ralf |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Translating the Force Concept Inventory in the age of AI
by: Babayeva, Marina, et al.
Published: (2025)
by: Babayeva, Marina, et al.
Published: (2025)
Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories
by: Kortemeyer, Gerd, et al.
Published: (2025)
by: Kortemeyer, Gerd, et al.
Published: (2025)
Realizing Visual Question Answering for Education: GPT-4V as a Multimodal AI
by: Lee, Gyeong-Geon, et al.
Published: (2024)
by: Lee, Gyeong-Geon, et al.
Published: (2024)
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback
by: Yao, Zonghai, et al.
Published: (2024)
by: Yao, Zonghai, et al.
Published: (2024)
Developing and Evaluating a Large Language Model-Based Automated Feedback System Grounded in Evidence-Centered Design for Supporting Physics Problem Solving
by: Maus, Holger, et al.
Published: (2025)
by: Maus, Holger, et al.
Published: (2025)
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
by: Wu, Sihong, et al.
Published: (2026)
by: Wu, Sihong, et al.
Published: (2026)
Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment
by: Tschisgale, Paul, et al.
Published: (2025)
by: Tschisgale, Paul, et al.
Published: (2025)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind
by: He, Zhitao, et al.
Published: (2026)
by: He, Zhitao, et al.
Published: (2026)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
by: Deng, Wenqing, et al.
Published: (2024)
by: Deng, Wenqing, et al.
Published: (2024)
Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts
by: Bleckmann, Tom, et al.
Published: (2025)
by: Bleckmann, Tom, et al.
Published: (2025)
Perfect score on IPhO 2025 theory by Gemini agent
by: Huang, Yichen
Published: (2026)
by: Huang, Yichen
Published: (2026)
A Personalised Learning Tool for Physics Undergraduate Students Built On a Large Language Model for Symbolic Regression
by: Zhu, Yufan, et al.
Published: (2024)
by: Zhu, Yufan, et al.
Published: (2024)
A simulation-heuristics dual-process model for intuitive physics
by: Li, Shiqian, et al.
Published: (2025)
by: Li, Shiqian, et al.
Published: (2025)
Reliable generation of isomorphic physics problems using Generative AI with prompt-chaining and tool use
by: Chen, Zhongzhou
Published: (2025)
by: Chen, Zhongzhou
Published: (2025)
Challenge-Device-Synthesis: A multi-disciplinary approach for the development of social innovation competences for students of Artificial Intelligence
by: Bilkis, Matías, et al.
Published: (2024)
by: Bilkis, Matías, et al.
Published: (2024)
Scaffold or Crutch? Examining College Students' Use and Views of Generative AI Tools for STEM Education
by: Wang, Karen D., et al.
Published: (2024)
by: Wang, Karen D., et al.
Published: (2024)
Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance
by: Ma, Qianli, et al.
Published: (2026)
by: Ma, Qianli, et al.
Published: (2026)
LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework
by: Rodríguez, Jesse A.
Published: (2026)
by: Rodríguez, Jesse A.
Published: (2026)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Orchestrating LLM Agents for Scientific Research: A Pilot Study of Multiple Choice Question (MCQ) Generation and Evaluation
by: An, Yuan
Published: (2026)
by: An, Yuan
Published: (2026)
Introducing First-Principles Calculations: New Approach to Group Dynamics and Bridging Social Phenomena in TeNP-Chain Based Social Dynamics Simulations
by: Kawahata, Yasuko
Published: (2024)
by: Kawahata, Yasuko
Published: (2024)
SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
by: Xiang, Kun, et al.
Published: (2025)
by: Xiang, Kun, et al.
Published: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024)
by: Palta, Shramay, et al.
Published: (2024)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
by: Lee, Yooseop, et al.
Published: (2025)
by: Lee, Yooseop, et al.
Published: (2025)
Biomedical Entity Linking as Multiple Choice Question Answering
by: Lin, Zhenxi, et al.
Published: (2024)
by: Lin, Zhenxi, et al.
Published: (2024)
Option-ID Based Elimination For Multiple Choice Questions
by: Zhu, Zhenhao, et al.
Published: (2025)
by: Zhu, Zhenhao, et al.
Published: (2025)
Automated Generation and Tagging of Knowledge Components from Multiple-Choice Questions
by: Moore, Steven, et al.
Published: (2024)
by: Moore, Steven, et al.
Published: (2024)
Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance
by: Khatri, Jyotsana, et al.
Published: (2026)
by: Khatri, Jyotsana, et al.
Published: (2026)
Evaluating and Enhancing LLMs for Multi-turn Text-to-SQL with Multiple Question Types
by: Guo, Ziming, et al.
Published: (2024)
by: Guo, Ziming, et al.
Published: (2024)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
by: Mustapha, Ahmad, et al.
Published: (2024)
by: Mustapha, Ahmad, et al.
Published: (2024)
A Dialogue-Based Framework for Correcting Multimodal Errors in AI-Assisted STEM Education
by: Syal, Akshay, et al.
Published: (2026)
by: Syal, Akshay, et al.
Published: (2026)
Assessing Large Language Models in Mechanical Engineering Education: A Study on Mechanics-Focused Conceptual Understanding
by: Tian, Jie, et al.
Published: (2024)
by: Tian, Jie, et al.
Published: (2024)
Developing an AI Course for Synthetic Chemistry Students
by: Zheng, Zhiling
Published: (2025)
by: Zheng, Zhiling
Published: (2025)
Investigation of the effectiveness of applying ChatGPT in Dialogic Teaching Using Electroencephalography
by: Zhang, Jiayue, et al.
Published: (2024)
by: Zhang, Jiayue, et al.
Published: (2024)
Mastering Olympiad-Level Physics with Artificial Intelligence
by: Jian, Dong-Shan, et al.
Published: (2025)
by: Jian, Dong-Shan, et al.
Published: (2025)
Bridging the Digital Divide: Small Language Models as a Pathway for Physics and Photonics Education in Underdeveloped Regions
by: Ghorbani, Asghar, et al.
Published: (2025)
by: Ghorbani, Asghar, et al.
Published: (2025)
Dissecting Physics Reasoning in Small Language Models: A Multi-Dimensional Analysis from an Educational Perspective
by: Scaria, Nicy, et al.
Published: (2025)
by: Scaria, Nicy, et al.
Published: (2025)
Reward Learning from Multiple Feedback Types
by: Metz, Yannick, et al.
Published: (2025)
by: Metz, Yannick, et al.
Published: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
Similar Items
-
Translating the Force Concept Inventory in the age of AI
by: Babayeva, Marina, et al.
Published: (2025) -
Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories
by: Kortemeyer, Gerd, et al.
Published: (2025) -
Realizing Visual Question Answering for Education: GPT-4V as a Multimodal AI
by: Lee, Gyeong-Geon, et al.
Published: (2024) -
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback
by: Yao, Zonghai, et al.
Published: (2024) -
Developing and Evaluating a Large Language Model-Based Automated Feedback System Grounded in Evidence-Centered Design for Supporting Physics Problem Solving
by: Maus, Holger, et al.
Published: (2025)