Criterion-referenceability determines LLM-as-a-judge validity across physics assessment formats
Fuente:
arXiv
Salvato in:
| Autori principali: | Yeadon, Will, Hardy, Tom, Mackay, Paul, Agra, Elise |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating AI and Human Authorship Quality in Academic Writing through Physics Essays
di: Yeadon, Will, et al.
Pubblicazione: (2024)
di: Yeadon, Will, et al.
Pubblicazione: (2024)
The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
di: Yeadon, Will, et al.
Pubblicazione: (2023)
di: Yeadon, Will, et al.
Pubblicazione: (2023)
Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts
di: Bleckmann, Tom, et al.
Pubblicazione: (2025)
di: Bleckmann, Tom, et al.
Pubblicazione: (2025)
Stan: An LLM-based thermodynamics course assistant
di: Furst, Eric M., et al.
Pubblicazione: (2026)
di: Furst, Eric M., et al.
Pubblicazione: (2026)
From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics
di: Geißler, Anna, et al.
Pubblicazione: (2025)
di: Geißler, Anna, et al.
Pubblicazione: (2025)
A comparison of Human, GPT-3.5, and GPT-4 Performance in a University-Level Coding Course
di: Yeadon, Will, et al.
Pubblicazione: (2024)
di: Yeadon, Will, et al.
Pubblicazione: (2024)
The Honorific Effect: Exploring the Impact of Japanese Linguistic Formalities on AI-Generated Physics Explanations
di: Sato, Keisuke
Pubblicazione: (2024)
di: Sato, Keisuke
Pubblicazione: (2024)
Scalable and consistent few-shot classification of survey responses using text embeddings
di: Mjaaland, Jonas Timmann, et al.
Pubblicazione: (2025)
di: Mjaaland, Jonas Timmann, et al.
Pubblicazione: (2025)
Small Language Models Reshape Higher Education: Courses, Textbooks, and Teaching
di: Zhang, Jian, et al.
Pubblicazione: (2025)
di: Zhang, Jian, et al.
Pubblicazione: (2025)
Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
di: Tschisgale, Paul, et al.
Pubblicazione: (2026)
di: Tschisgale, Paul, et al.
Pubblicazione: (2026)
Self-rationalization improves LLM as a fine-grained judge
di: Trivedi, Prapti, et al.
Pubblicazione: (2024)
di: Trivedi, Prapti, et al.
Pubblicazione: (2024)
An LLM-as-a-judge Approach for Scalable Gender-Neutral Translation Evaluation
di: Piergentili, Andrea, et al.
Pubblicazione: (2025)
di: Piergentili, Andrea, et al.
Pubblicazione: (2025)
Understanding interaction network formation across instructional contexts in remote physics courses
di: Sundstrom, Meagan, et al.
Pubblicazione: (2022)
di: Sundstrom, Meagan, et al.
Pubblicazione: (2022)
Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment
di: Kumar, Vanya Bannihatti, et al.
Pubblicazione: (2025)
di: Kumar, Vanya Bannihatti, et al.
Pubblicazione: (2025)
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
di: Li, Dawei, et al.
Pubblicazione: (2024)
di: Li, Dawei, et al.
Pubblicazione: (2024)
Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring
di: Dussert, Jehanne
Pubblicazione: (2026)
di: Dussert, Jehanne
Pubblicazione: (2026)
When Wording Steers the Evaluation: Framing Bias in LLM judges
di: Hwang, Yerin, et al.
Pubblicazione: (2026)
di: Hwang, Yerin, et al.
Pubblicazione: (2026)
Scaling Physical Reasoning with the PHYSICS Dataset
di: Zheng, Shenghe, et al.
Pubblicazione: (2025)
di: Zheng, Shenghe, et al.
Pubblicazione: (2025)
Assessing Large Language Models in Mechanical Engineering Education: A Study on Mechanics-Focused Conceptual Understanding
di: Tian, Jie, et al.
Pubblicazione: (2024)
di: Tian, Jie, et al.
Pubblicazione: (2024)
Translating scientific Latin texts with artificial intelligence: the works of Euler and contemporaries
di: Bistafa, Sylvio R.
Pubblicazione: (2023)
di: Bistafa, Sylvio R.
Pubblicazione: (2023)
Mastering Olympiad-Level Physics with Artificial Intelligence
di: Jian, Dong-Shan, et al.
Pubblicazione: (2025)
di: Jian, Dong-Shan, et al.
Pubblicazione: (2025)
Dissecting Physics Reasoning in Small Language Models: A Multi-Dimensional Analysis from an Educational Perspective
di: Scaria, Nicy, et al.
Pubblicazione: (2025)
di: Scaria, Nicy, et al.
Pubblicazione: (2025)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
di: Li, Dawei, et al.
Pubblicazione: (2025)
di: Li, Dawei, et al.
Pubblicazione: (2025)
Efficacy of a hybrid take home and in class summative assessment for the postsecondary physics classroom
di: Stonaha, Paul, et al.
Pubblicazione: (2024)
di: Stonaha, Paul, et al.
Pubblicazione: (2024)
LLM-as-a-qualitative-judge: automating error analysis in natural language generation
di: Chirkova, Nadezhda, et al.
Pubblicazione: (2025)
di: Chirkova, Nadezhda, et al.
Pubblicazione: (2025)
From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
di: Stephan, Andreas, et al.
Pubblicazione: (2024)
di: Stephan, Andreas, et al.
Pubblicazione: (2024)
QISCIT: A validated concept inventory assessment for quantum information science
di: Durkin, Kelley, et al.
Pubblicazione: (2025)
di: Durkin, Kelley, et al.
Pubblicazione: (2025)
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
di: Sieker, Judith, et al.
Pubblicazione: (2026)
di: Sieker, Judith, et al.
Pubblicazione: (2026)
Concurrent Criterion Validation of a Validity Screen for LLM Confidence Signals via Selective Prediction
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
Development and validation of a conceptual multiple-choice survey instrument to assess student understanding of introductory thermodynamics
di: Brundage, Mary Jane, et al.
Pubblicazione: (2024)
di: Brundage, Mary Jane, et al.
Pubblicazione: (2024)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
di: Cao, Hongliu, et al.
Pubblicazione: (2025)
di: Cao, Hongliu, et al.
Pubblicazione: (2025)
Online antisemitism across platforms
di: De Smedt, Tom
Pubblicazione: (2021)
di: De Smedt, Tom
Pubblicazione: (2021)
Criterion Validity of LLM-as-Judge for Business Outcomes in Conversational Commerce
di: Chen, Liang, et al.
Pubblicazione: (2026)
di: Chen, Liang, et al.
Pubblicazione: (2026)
Comparing student performance on a multi-attempt asynchronous assessment to a single-attempt synchronous assessment in introductory level physics
di: Frederick, Emily, et al.
Pubblicazione: (2024)
di: Frederick, Emily, et al.
Pubblicazione: (2024)
Overview of couplet scoring in content-focused physics assessments
di: Vignal, Michael, et al.
Pubblicazione: (2023)
di: Vignal, Michael, et al.
Pubblicazione: (2023)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
di: Slyman, Eric, et al.
Pubblicazione: (2025)
di: Slyman, Eric, et al.
Pubblicazione: (2025)
Impact of Criterion-Based Reflection on Prospective Physics Teachers' Perceptions of ChatGPT-Generated Content
di: Sadidi, Farahnaz, et al.
Pubblicazione: (2024)
di: Sadidi, Farahnaz, et al.
Pubblicazione: (2024)
Evaluating recognition and recall formats of social network surveys in physics education research
di: Sundstrom, Meagan, et al.
Pubblicazione: (2025)
di: Sundstrom, Meagan, et al.
Pubblicazione: (2025)
Society's educational debts in biology, chemistry, and physics across race, gender, and class
di: Van Dusen, Ben, et al.
Pubblicazione: (2024)
di: Van Dusen, Ben, et al.
Pubblicazione: (2024)
Investigating peer recognition across an introductory physics sequence: Do first impressions last?
di: Sundstrom, Meagan, et al.
Pubblicazione: (2023)
di: Sundstrom, Meagan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Evaluating AI and Human Authorship Quality in Academic Writing through Physics Essays
di: Yeadon, Will, et al.
Pubblicazione: (2024) -
The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
di: Yeadon, Will, et al.
Pubblicazione: (2023) -
Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts
di: Bleckmann, Tom, et al.
Pubblicazione: (2025) -
Stan: An LLM-based thermodynamics course assistant
di: Furst, Eric M., et al.
Pubblicazione: (2026) -
From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics
di: Geißler, Anna, et al.
Pubblicazione: (2025)