Applying IRT to Distinguish Between Human and Generative AI Responses to Multiple-Choice Assessments
Fuente:
arXiv
Guardado en:
| Autores principales: | Strugatski, Alona, Alexandron, Giora |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Assessment Design in the AI Era: A Method for Identifying Items Functioning Differentially for Humans and Chatbots
por: Zeinfeld, Licol, et al.
Publicado: (2026)
por: Zeinfeld, Licol, et al.
Publicado: (2026)
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
por: Schleifer, Abigail Victoria Gurin, et al.
Publicado: (2026)
por: Schleifer, Abigail Victoria Gurin, et al.
Publicado: (2026)
Anna Karenina Strikes Again: Pre-Trained LLM Embeddings May Favor High-Performing Learners
por: Schleifer, Abigail Gurin, et al.
Publicado: (2024)
por: Schleifer, Abigail Gurin, et al.
Publicado: (2024)
Causal‐mechanical explanations in biology: Applying automated assessment for personalized learning in the science classroom
por: Moriah Ariely, et al.
Publicado: (2024)
por: Moriah Ariely, et al.
Publicado: (2024)
SciTextures: Collecting and Connecting Visual Patterns, Models, and Code Across Science and Art
por: Eppel, Sagi, et al.
Publicado: (2025)
por: Eppel, Sagi, et al.
Publicado: (2025)
IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response Theory
por: Song, Wei, et al.
Publicado: (2025)
por: Song, Wei, et al.
Publicado: (2025)
Generative Topological Networks
por: Levy-Jurgenson, Alona, et al.
Publicado: (2024)
por: Levy-Jurgenson, Alona, et al.
Publicado: (2024)
Distinguishing AI-Generated and Human-Written Text Through Psycholinguistic Analysis
por: Opara, Chidimma
Publicado: (2025)
por: Opara, Chidimma
Publicado: (2025)
Automated Generation of Curriculum-Aligned Multiple-Choice Questions for Malaysian Secondary Mathematics Using Generative AI
por: Wahid, Rohaizah Abdul, et al.
Publicado: (2025)
por: Wahid, Rohaizah Abdul, et al.
Publicado: (2025)
Distinguishing Task-Specific and General-Purpose AI in Regulation
por: Wang, Jennifer, et al.
Publicado: (2025)
por: Wang, Jennifer, et al.
Publicado: (2025)
SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions
por: Xu, Weijie, et al.
Publicado: (2025)
por: Xu, Weijie, et al.
Publicado: (2025)
JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
por: Yao, Louie Hong, et al.
Publicado: (2025)
por: Yao, Louie Hong, et al.
Publicado: (2025)
Aligning to Illusions: Choice Blindness in Human and AI Feedback
por: Wu, Wenbin
Publicado: (2026)
por: Wu, Wenbin
Publicado: (2026)
Feedback Indices to Evaluate LLM Responses to Rebuttals for Multiple Choice Type Questions
por: Dunlap, Justin C., et al.
Publicado: (2026)
por: Dunlap, Justin C., et al.
Publicado: (2026)
CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation
por: Duan, Xiaojing, et al.
Publicado: (2026)
por: Duan, Xiaojing, et al.
Publicado: (2026)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
por: Singh, Shrutika, et al.
Publicado: (2025)
por: Singh, Shrutika, et al.
Publicado: (2025)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
por: Zhang, Jingshen, et al.
Publicado: (2024)
por: Zhang, Jingshen, et al.
Publicado: (2024)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
por: Lee, Yooseop, et al.
Publicado: (2025)
por: Lee, Yooseop, et al.
Publicado: (2025)
Resisting Humanization: Ethical Front-End Design Choices in AI for Sensitive Contexts
por: Rossi, Silvia, et al.
Publicado: (2026)
por: Rossi, Silvia, et al.
Publicado: (2026)
When Choices Become Priors: Contrastive Decoding for Scientific Figure Multiple-Choice QA
por: Roh, Taeyun, et al.
Publicado: (2026)
por: Roh, Taeyun, et al.
Publicado: (2026)
Shape and Texture Recognition in Large Vision-Language Models
por: Eppel, Sagi, et al.
Publicado: (2025)
por: Eppel, Sagi, et al.
Publicado: (2025)
Measuring Creativity in the Age of Generative AI: Distinguishing Human and AI-Generated Creative Performance in Hiring and Talent Systems
por: Rosen, Yigal, et al.
Publicado: (2026)
por: Rosen, Yigal, et al.
Publicado: (2026)
StyloAI: Distinguishing AI-Generated Content with Stylometric Analysis
por: Opara, Chidimma
Publicado: (2024)
por: Opara, Chidimma
Publicado: (2024)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
por: Deng, Wenqing, et al.
Publicado: (2024)
por: Deng, Wenqing, et al.
Publicado: (2024)
Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCT
por: Thomas, Danielle R., et al.
Publicado: (2024)
por: Thomas, Danielle R., et al.
Publicado: (2024)
Automated Evaluation can Distinguish the Good and Bad AI Responses to Patient Questions about Hospitalization
por: Soni, Sarvesh, et al.
Publicado: (2025)
por: Soni, Sarvesh, et al.
Publicado: (2025)
Synthetic Student Responses: LLM-Extracted Features for IRT Difficulty Parameter Estimation
por: Hoyl, Matias
Publicado: (2026)
por: Hoyl, Matias
Publicado: (2026)
To accept or not to accept? An IRT-TOE Framework to Understand Educators' Resistance to Generative AI in Higher Education
por: Kalmus, Jan-Erik, et al.
Publicado: (2024)
por: Kalmus, Jan-Erik, et al.
Publicado: (2024)
Automated Generation and Tagging of Knowledge Components from Multiple-Choice Questions
por: Moore, Steven, et al.
Publicado: (2024)
por: Moore, Steven, et al.
Publicado: (2024)
Self-Correcting Large Language Models: Generation vs. Multiple Choice
por: Rahmani, Hossein A., et al.
Publicado: (2025)
por: Rahmani, Hossein A., et al.
Publicado: (2025)
Integrating ESG and AI: A Comprehensive Responsible AI Assessment Framework
por: Lee, Sung Une, et al.
Publicado: (2024)
por: Lee, Sung Une, et al.
Publicado: (2024)
XChoice: Explainable Evaluation of AI-Human Alignment in LLM-based Constrained Choice Decision Making
por: Qi, Weihong, et al.
Publicado: (2026)
por: Qi, Weihong, et al.
Publicado: (2026)
Process Matters more than Output for Distinguishing Humans from Machines
por: Rmus, Milena, et al.
Publicado: (2026)
por: Rmus, Milena, et al.
Publicado: (2026)
Generalization Bounds in Hybrid Quantum-Classical Machine Learning Models
por: Wu, Tongyan, et al.
Publicado: (2025)
por: Wu, Tongyan, et al.
Publicado: (2025)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
por: Mohamadi, Alireza, et al.
Publicado: (2025)
por: Mohamadi, Alireza, et al.
Publicado: (2025)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
por: Cavalin, Paulo, et al.
Publicado: (2025)
por: Cavalin, Paulo, et al.
Publicado: (2025)
AI-Induced Human Responsibility (AIHR) in AI-Human teams
por: Nyilasy, Greg, et al.
Publicado: (2026)
por: Nyilasy, Greg, et al.
Publicado: (2026)
Modeling Human Responses to Multimodal AI Content
por: Shen, Zhiqi, et al.
Publicado: (2025)
por: Shen, Zhiqi, et al.
Publicado: (2025)
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
por: Rad, Mohammad Heydari, et al.
Publicado: (2024)
por: Rad, Mohammad Heydari, et al.
Publicado: (2024)
Multiple-Choice Question Generation Using Large Language Models: Methodology and Educator Insights
por: Biancini, Giorgio, et al.
Publicado: (2025)
por: Biancini, Giorgio, et al.
Publicado: (2025)
Ejemplares similares
-
Assessment Design in the AI Era: A Method for Identifying Items Functioning Differentially for Humans and Chatbots
por: Zeinfeld, Licol, et al.
Publicado: (2026) -
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
por: Schleifer, Abigail Victoria Gurin, et al.
Publicado: (2026) -
Anna Karenina Strikes Again: Pre-Trained LLM Embeddings May Favor High-Performing Learners
por: Schleifer, Abigail Gurin, et al.
Publicado: (2024) -
Causal‐mechanical explanations in biology: Applying automated assessment for personalized learning in the science classroom
por: Moriah Ariely, et al.
Publicado: (2024) -
SciTextures: Collecting and Connecting Visual Patterns, Models, and Code Across Science and Art
por: Eppel, Sagi, et al.
Publicado: (2025)