LLM-Driven Rubric-Based Assessment of Algebraic Competence in Multi-Stage Block Coding Tasks with Design and Field Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Yong Oh, Bang, Byeonghun, Oh, Sejun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Personalized Auto-Grading and Feedback System for Constructive Geometry Tasks Using Large Language Models on an Online Math Platform
di: Lee, Yong Oh, et al.
Pubblicazione: (2025)
di: Lee, Yong Oh, et al.
Pubblicazione: (2025)
Modelling Assessment Rubrics through Bayesian Networks: a Pragmatic Approach
di: Mangili, Francesca, et al.
Pubblicazione: (2022)
di: Mangili, Francesca, et al.
Pubblicazione: (2022)
Strategies for Creating Uncertainty in the AI Era to Trigger Students Critical Thinking: Pedagogical Design, Assessment Rubric, and Exam System
di: Wazan, Ahmad Samer
Pubblicazione: (2026)
di: Wazan, Ahmad Samer
Pubblicazione: (2026)
Retrieval Augmented Large Language Model System for Comprehensive Drug Contraindications
di: Bang, Byeonghun, et al.
Pubblicazione: (2025)
di: Bang, Byeonghun, et al.
Pubblicazione: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
di: Shi, Yuzhen, et al.
Pubblicazione: (2026)
di: Shi, Yuzhen, et al.
Pubblicazione: (2026)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
di: Shin, Jisu, et al.
Pubblicazione: (2025)
di: Shin, Jisu, et al.
Pubblicazione: (2025)
The Limits of Goal-Setting Theory in LLM-Driven Assessment
di: Kumar, Mrityunjay
Pubblicazione: (2025)
di: Kumar, Mrityunjay
Pubblicazione: (2025)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
di: Oh, Gyutaek, et al.
Pubblicazione: (2025)
di: Oh, Gyutaek, et al.
Pubblicazione: (2025)
LLM-Driven Personalized Answer Generation and Evaluation
di: Molavi, Mohammadreza, et al.
Pubblicazione: (2025)
di: Molavi, Mohammadreza, et al.
Pubblicazione: (2025)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
An Evaluation of Cultural Value Alignment in LLM
di: Sukiennik, Nicholas, et al.
Pubblicazione: (2025)
di: Sukiennik, Nicholas, et al.
Pubblicazione: (2025)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
di: Ding, Liang
Pubblicazione: (2026)
di: Ding, Liang
Pubblicazione: (2026)
Exploring Teachers' Perception of Artificial Intelligence: The Socio-emotional Deficiency as Opportunities and Challenges in Human-AI Complementarity in K-12 Education
di: Oh, Soon-young, et al.
Pubblicazione: (2024)
di: Oh, Soon-young, et al.
Pubblicazione: (2024)
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
di: Liu, Xue, et al.
Pubblicazione: (2026)
di: Liu, Xue, et al.
Pubblicazione: (2026)
RubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions
di: Goel, Drishti, et al.
Pubblicazione: (2026)
di: Goel, Drishti, et al.
Pubblicazione: (2026)
Exploring Teacher-Chatbot Interaction and Affect in Block-Based Programming
di: Riahi, Bahare, et al.
Pubblicazione: (2026)
di: Riahi, Bahare, et al.
Pubblicazione: (2026)
Developing a Multi-Agent System to Generate Next Generation Science Assessments with Evidence-Centered Design
di: Yang, Yaxuan, et al.
Pubblicazione: (2026)
di: Yang, Yaxuan, et al.
Pubblicazione: (2026)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
CodeGuard: Improving LLM Guardrails in CS Education
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments
di: Vanhoyweghen, Arne, et al.
Pubblicazione: (2026)
di: Vanhoyweghen, Arne, et al.
Pubblicazione: (2026)
Scaling Behavior of Single LLM-Driven Multi-Agent Systems
di: Li, Jialing, et al.
Pubblicazione: (2026)
di: Li, Jialing, et al.
Pubblicazione: (2026)
Classroom AI: Large Language Models as Grade-Specific Teachers
di: Oh, Jio, et al.
Pubblicazione: (2026)
di: Oh, Jio, et al.
Pubblicazione: (2026)
A Large-Scale Real-World Evaluation of LLM-Based Virtual Teaching Assistant
di: Kweon, Sunjun, et al.
Pubblicazione: (2025)
di: Kweon, Sunjun, et al.
Pubblicazione: (2025)
Multimodal Assessment of Classroom Discourse Quality: A Text-Centered Attention-Based Multi-Task Learning Approach
di: Hou, Ruikun, et al.
Pubblicazione: (2025)
di: Hou, Ruikun, et al.
Pubblicazione: (2025)
Elementary School Students' and Teachers' Perceptions Towards Creative Mathematical Writing with Generative AI
di: Song, Yukyeong, et al.
Pubblicazione: (2024)
di: Song, Yukyeong, et al.
Pubblicazione: (2024)
The Impact of AI on Academic Research and Publishing
di: Lund, Brady, et al.
Pubblicazione: (2024)
di: Lund, Brady, et al.
Pubblicazione: (2024)
REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
di: Zhao, Chengshuai, et al.
Pubblicazione: (2026)
di: Zhao, Chengshuai, et al.
Pubblicazione: (2026)
Automated Assessment of Students' Code Comprehension using LLMs
di: Oli, Priti, et al.
Pubblicazione: (2023)
di: Oli, Priti, et al.
Pubblicazione: (2023)
SNAP: A Plan-Driven Framework for Controllable Interactive Narrative Generation
di: Bang, Geonwoo, et al.
Pubblicazione: (2025)
di: Bang, Geonwoo, et al.
Pubblicazione: (2025)
Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials
di: He, Peng, et al.
Pubblicazione: (2026)
di: He, Peng, et al.
Pubblicazione: (2026)
DREsS: Dataset for Rubric-based Essay Scoring on EFL Writing
di: Yoo, Haneul, et al.
Pubblicazione: (2024)
di: Yoo, Haneul, et al.
Pubblicazione: (2024)
Synthesizing High-Quality Programming Tasks with LLM-based Expert and Student Agents
di: Nguyen, Manh Hung, et al.
Pubblicazione: (2025)
di: Nguyen, Manh Hung, et al.
Pubblicazione: (2025)
Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways
di: Atreja, Shubham, et al.
Pubblicazione: (2024)
di: Atreja, Shubham, et al.
Pubblicazione: (2024)
OnlineMate: An LLM-Based Multi-Agent Companion System for Cognitive Support in Online Learning
di: Gao, Xian, et al.
Pubblicazione: (2025)
di: Gao, Xian, et al.
Pubblicazione: (2025)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
Experiences with Content Development and Assessment Design in the Era of GenAI
di: Sharma, Aakanksha, et al.
Pubblicazione: (2025)
di: Sharma, Aakanksha, et al.
Pubblicazione: (2025)
AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
di: Bogart, Christopher, et al.
Pubblicazione: (2025)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
di: Yoon, Sion, et al.
Pubblicazione: (2024)
di: Yoon, Sion, et al.
Pubblicazione: (2024)
Evaluating Code Generation of LLMs in Advanced Computer Science Problems
di: Catir, Emir, et al.
Pubblicazione: (2025)
di: Catir, Emir, et al.
Pubblicazione: (2025)
Code-Driven Law NO, Normware SI!
di: Sileno, Giovanni
Pubblicazione: (2024)
di: Sileno, Giovanni
Pubblicazione: (2024)
Documenti analoghi
-
Personalized Auto-Grading and Feedback System for Constructive Geometry Tasks Using Large Language Models on an Online Math Platform
di: Lee, Yong Oh, et al.
Pubblicazione: (2025) -
Modelling Assessment Rubrics through Bayesian Networks: a Pragmatic Approach
di: Mangili, Francesca, et al.
Pubblicazione: (2022) -
Strategies for Creating Uncertainty in the AI Era to Trigger Students Critical Thinking: Pedagogical Design, Assessment Rubric, and Exam System
di: Wazan, Ahmad Samer
Pubblicazione: (2026) -
Retrieval Augmented Large Language Model System for Comprehensive Drug Contraindications
di: Bang, Byeonghun, et al.
Pubblicazione: (2025) -
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
di: Shi, Yuzhen, et al.
Pubblicazione: (2026)