Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
Fuente:
arXiv
Saved in:
| Main Authors: | Caraeni, Adriana, Scarlatos, Alexander, Lan, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiVERT: Distractor Generation with Variational Errors Represented as Text for Math Multiple-choice Questions
by: Fernandez, Nigel, et al.
Published: (2024)
by: Fernandez, Nigel, et al.
Published: (2024)
Exploring Knowledge Tracing in Tutor-Student Dialogues using LLMs
by: Scarlatos, Alexander, et al.
Published: (2024)
by: Scarlatos, Alexander, et al.
Published: (2024)
SyllabusQA: A Course Logistics Question Answering Dataset
by: Fernandez, Nigel, et al.
Published: (2024)
by: Fernandez, Nigel, et al.
Published: (2024)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
Improving Automated Distractor Generation for Math Multiple-choice Questions with Overgenerate-and-rank
by: Scarlatos, Alexander, et al.
Published: (2024)
by: Scarlatos, Alexander, et al.
Published: (2024)
RetICL: Sequential Retrieval of In-Context Examples with Reinforcement Learning
by: Scarlatos, Alexander, et al.
Published: (2023)
by: Scarlatos, Alexander, et al.
Published: (2023)
Exploring LLMs for Predicting Tutor Strategy and Student Outcomes in Dialogues
by: Ikram, Fareya, et al.
Published: (2025)
by: Ikram, Fareya, et al.
Published: (2025)
Simulated Students in Tutoring Dialogues: Substance or Illusion?
by: Scarlatos, Alexander, et al.
Published: (2026)
by: Scarlatos, Alexander, et al.
Published: (2026)
Test Case-Informed Knowledge Tracing for Open-ended Coding Tasks
by: Duan, Zhangqi, et al.
Published: (2024)
by: Duan, Zhangqi, et al.
Published: (2024)
Training LLM-based Tutors to Improve Student Learning Outcomes in Dialogues
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
Interpreting Latent Student Knowledge Representations in Programming Assignments
by: Fernandez, Nigel, et al.
Published: (2024)
by: Fernandez, Nigel, et al.
Published: (2024)
Improving Socratic Question Generation using Data Augmentation and Preference Optimization
by: Kumar, Nischal Ashok, et al.
Published: (2024)
by: Kumar, Nischal Ashok, et al.
Published: (2024)
PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases
by: Schlegel, Udo, et al.
Published: (2025)
by: Schlegel, Udo, et al.
Published: (2025)
Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues
by: Duan, Zhangqi, et al.
Published: (2026)
by: Duan, Zhangqi, et al.
Published: (2026)
Evaluating the Performance of ChatGPT for Spam Email Detection
by: Si, Shijing, et al.
Published: (2024)
by: Si, Shijing, et al.
Published: (2024)
ChatGPT in Linear Algebra: Strides Forward, Steps to Go
by: Bagno, Eli, et al.
Published: (2024)
by: Bagno, Eli, et al.
Published: (2024)
RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)
by: Cuclea, Luca-Ncolae, et al.
Published: (2026)
by: Cuclea, Luca-Ncolae, et al.
Published: (2026)
The Dark Side of ChatGPT: Legal and Ethical Challenges from Stochastic Parrots and Hallucination
by: Li, Zihao
Published: (2023)
by: Li, Zihao
Published: (2023)
Artificial-Intelligence Grading Assistance for Handwritten Components of a Calculus Exam
by: Kortemeyer, Gerd, et al.
Published: (2025)
by: Kortemeyer, Gerd, et al.
Published: (2025)
KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks
by: Duan, Zhangqi, et al.
Published: (2026)
by: Duan, Zhangqi, et al.
Published: (2026)
Fairness of ChatGPT
by: Li, Yunqi, et al.
Published: (2023)
by: Li, Yunqi, et al.
Published: (2023)
When VLMs 'Fix' Students: Identifying and Penalizing Over-Correction in the Evaluation of Multi-line Handwritten Math OCR
by: Seong, Jin, et al.
Published: (2026)
by: Seong, Jin, et al.
Published: (2026)
ChatGPT Needs SPADE (Sustainability, PrivAcy, Digital divide, and Ethics) Evaluation: A Review
by: Khowaja, Sunder Ali, et al.
Published: (2023)
by: Khowaja, Sunder Ali, et al.
Published: (2023)
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models
by: Era, Jalisha Jashim, et al.
Published: (2025)
by: Era, Jalisha Jashim, et al.
Published: (2025)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
by: Urchs, Stefanie, et al.
Published: (2023)
by: Urchs, Stefanie, et al.
Published: (2023)
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems
by: Saadat, Asir, et al.
Published: (2024)
by: Saadat, Asir, et al.
Published: (2024)
Navigating the Peril of Generated Alternative Facts: A ChatGPT-4 Fabricated Omega Variant Case as a Cautionary Tale in Medical Misinformation
by: Sallam, Malik, et al.
Published: (2024)
by: Sallam, Malik, et al.
Published: (2024)
Towards Explainable Evaluation Metrics for Machine Translation
by: Leiter, Christoph, et al.
Published: (2023)
by: Leiter, Christoph, et al.
Published: (2023)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
by: Anzenberg, Eitan, et al.
Published: (2025)
by: Anzenberg, Eitan, et al.
Published: (2025)
Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
by: Peczuh, Marisa C., et al.
Published: (2025)
by: Peczuh, Marisa C., et al.
Published: (2025)
ViMGuard: A Novel Multi-Modal System for Video Misinformation Guarding
by: Kan, Andrew, et al.
Published: (2024)
by: Kan, Andrew, et al.
Published: (2024)
RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
by: de Wynter, Adrian, et al.
Published: (2024)
by: de Wynter, Adrian, et al.
Published: (2024)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
Generalization in Healthcare AI: Evaluation of a Clinical Large Language Model
by: Rahman, Salman, et al.
Published: (2024)
by: Rahman, Salman, et al.
Published: (2024)
Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
by: Friedrich, Felix, et al.
Published: (2025)
by: Friedrich, Felix, et al.
Published: (2025)
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
by: Wei, Shou'ang, et al.
Published: (2025)
by: Wei, Shou'ang, et al.
Published: (2025)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
by: Pan, Alexander, et al.
Published: (2024)
by: Pan, Alexander, et al.
Published: (2024)
Personalized Auto-Grading and Feedback System for Constructive Geometry Tasks Using Large Language Models on an Online Math Platform
by: Lee, Yong Oh, et al.
Published: (2025)
by: Lee, Yong Oh, et al.
Published: (2025)
FairPair: A Robust Evaluation of Biases in Language Models through Paired Perturbations
by: Dwivedi-Yu, Jane, et al.
Published: (2024)
by: Dwivedi-Yu, Jane, et al.
Published: (2024)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
Similar Items
-
DiVERT: Distractor Generation with Variational Errors Represented as Text for Math Multiple-choice Questions
by: Fernandez, Nigel, et al.
Published: (2024) -
Exploring Knowledge Tracing in Tutor-Student Dialogues using LLMs
by: Scarlatos, Alexander, et al.
Published: (2024) -
SyllabusQA: A Course Logistics Question Answering Dataset
by: Fernandez, Nigel, et al.
Published: (2024) -
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025) -
Improving Automated Distractor Generation for Math Multiple-choice Questions with Overgenerate-and-rank
by: Scarlatos, Alexander, et al.
Published: (2024)