A Knowledge-Component-Based Methodology for Evaluating AI Assistants
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Laryn, Zamfirescu-Pereira, J. D., Kim, Taehan, Hartmann, Björn, DeNero, John, Norouzi, Narges |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
61A Bot Report: AI Assistants in CS1 Save Students Homework Time and Reduce Demands on Staff. (Now What?)
by: Zamfirescu-Pereira, J. D., et al.
Published: (2024)
by: Zamfirescu-Pereira, J. D., et al.
Published: (2024)
Pensieve Discuss: Scalable Small-Group CS Tutoring System with AI
by: Yang, Yoonseok, et al.
Published: (2024)
by: Yang, Yoonseok, et al.
Published: (2024)
The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
by: Niousha, Rose, et al.
Published: (2026)
by: Niousha, Rose, et al.
Published: (2026)
ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle
by: Miroyan, Mihran, et al.
Published: (2025)
by: Miroyan, Mihran, et al.
Published: (2025)
Personalized Worked Example Generation from Student Code Submissions Using Pattern-based Knowledge Components
by: Pitts, Griffin, et al.
Published: (2026)
by: Pitts, Griffin, et al.
Published: (2026)
Evaluating a Methodology for Increasing AI Transparency: A Case Study
by: Piorkowski, David, et al.
Published: (2022)
by: Piorkowski, David, et al.
Published: (2022)
Retrieval-Augmented Tutoring for Algorithm Tracing and Problem-Solving in AI Education
by: Jain, Mragisha, et al.
Published: (2026)
by: Jain, Mragisha, et al.
Published: (2026)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
by: Costa, Mariana Lins
Published: (2026)
by: Costa, Mariana Lins
Published: (2026)
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
by: Shankar, Shreya, et al.
Published: (2024)
by: Shankar, Shreya, et al.
Published: (2024)
A Large-Scale Real-World Evaluation of LLM-Based Virtual Teaching Assistant
by: Kweon, Sunjun, et al.
Published: (2025)
by: Kweon, Sunjun, et al.
Published: (2025)
Examining Student Interactions with a Pedagogical AI-Assistant for Essay Writing and their Impact on Students Writing Quality
by: Febriantoro, Wicaksono, et al.
Published: (2025)
by: Febriantoro, Wicaksono, et al.
Published: (2025)
Evaluating AI Evaluation: Perils and Prospects
by: Burden, John
Published: (2024)
by: Burden, John
Published: (2024)
EduMod-LLM: A Modular Approach for Designing Flexible and Transparent Educational Assistants
by: Mittal, Meenakshi, et al.
Published: (2025)
by: Mittal, Meenakshi, et al.
Published: (2025)
Reconciling Methodological Paradigms: Employing Large Language Models as Novice Qualitative Research Assistants in Talent Management Research
by: Bhaduri, Sreyoshi, et al.
Published: (2024)
by: Bhaduri, Sreyoshi, et al.
Published: (2024)
AI-driven Personalized Privacy Assistants: a Systematic Literature Review
by: Morel, Victor, et al.
Published: (2025)
by: Morel, Victor, et al.
Published: (2025)
Human or AI? Comparing Design Thinking Assessments by Teaching Assistants and Bots
by: Khan, Sumbul, et al.
Published: (2025)
by: Khan, Sumbul, et al.
Published: (2025)
Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants
by: Reddy, Pavan, et al.
Published: (2025)
by: Reddy, Pavan, et al.
Published: (2025)
Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants
by: Borges, Beatriz, et al.
Published: (2024)
by: Borges, Beatriz, et al.
Published: (2024)
Using AI-Based Coding Assistants in Practice: State of Affairs, Perceptions, and Ways Forward
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
by: Ding, Shi, et al.
Published: (2025)
by: Ding, Shi, et al.
Published: (2025)
Aalap: AI Assistant for Legal & Paralegal Functions in India
by: Tiwari, Aman, et al.
Published: (2024)
by: Tiwari, Aman, et al.
Published: (2024)
Addressing the regulatory gap: moving towards an EU AI audit ecosystem beyond the AI Act by including civil society
by: Hartmann, David, et al.
Published: (2024)
by: Hartmann, David, et al.
Published: (2024)
Digital Companionship: Overlapping Uses of AI Companions and AI Assistants
by: Manoli, Aikaterina, et al.
Published: (2025)
by: Manoli, Aikaterina, et al.
Published: (2025)
A GPU-Accelerated RAG-Based Telegram Assistant for Supporting Parallel Processing Students
by: Tel-Zur, Guy
Published: (2025)
by: Tel-Zur, Guy
Published: (2025)
Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant
by: Ballestero-Ribó, Marc, et al.
Published: (2025)
by: Ballestero-Ribó, Marc, et al.
Published: (2025)
A Study on the Framework for Evaluating the Ethics and Trustworthiness of Generative AI
by: Jeong, Cheonsu, et al.
Published: (2025)
by: Jeong, Cheonsu, et al.
Published: (2025)
Evaluating AI-Powered Learning Assistants in Engineering Higher Education: Student Engagement, Ethical Challenges, and Policy Implications
by: Sajja, Ramteja, et al.
Published: (2025)
by: Sajja, Ramteja, et al.
Published: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
SAE-RNA: A Sparse Autoencoder Model for Interpreting RNA Language Model Representations
by: Kim, Taehan, et al.
Published: (2025)
by: Kim, Taehan, et al.
Published: (2025)
Responsible Evaluation of AI for Mental Health
by: Arnaout, Hiba, et al.
Published: (2026)
by: Arnaout, Hiba, et al.
Published: (2026)
Photons = Tokens: The Physics of AI and the Economics of Knowledge
by: Litowitz, Alec, et al.
Published: (2026)
by: Litowitz, Alec, et al.
Published: (2026)
Customer Service Representative's Perception of the AI Assistant in an Organization's Call Center
by: Qin, Kai, et al.
Published: (2025)
by: Qin, Kai, et al.
Published: (2025)
Designing Skill-Compatible AI: Methodologies and Frameworks in Chess
by: Hamade, Karim, et al.
Published: (2024)
by: Hamade, Karim, et al.
Published: (2024)
Towards Ethical Personal AI Applications: Practical Considerations for AI Assistants with Long-Term Memory
by: Lee, Eunhae
Published: (2024)
by: Lee, Eunhae
Published: (2024)
An Open Knowledge Graph-Based Approach for Mapping Concepts and Requirements between the EU AI Act and International Standards
by: Hernandez, Julio, et al.
Published: (2024)
by: Hernandez, Julio, et al.
Published: (2024)
An AI Teaching Assistant for Motion Picture Engineering
by: O'Regan, Deirdre, et al.
Published: (2026)
by: O'Regan, Deirdre, et al.
Published: (2026)
EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants
by: Cappello, Franck, et al.
Published: (2025)
by: Cappello, Franck, et al.
Published: (2025)
Virtue Ethics For Ethically Tunable Robotic Assistants
by: Ramanayake, Rajitha, et al.
Published: (2024)
by: Ramanayake, Rajitha, et al.
Published: (2024)
Large Language Model-Based Knowledge Graph System Construction for Sustainable Development Goals: An AI-Based Speculative Design Perspective
by: Lin, Yi-De, et al.
Published: (2025)
by: Lin, Yi-De, et al.
Published: (2025)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025)
by: Sturgeon, Benjamin, et al.
Published: (2025)
Similar Items
-
61A Bot Report: AI Assistants in CS1 Save Students Homework Time and Reduce Demands on Staff. (Now What?)
by: Zamfirescu-Pereira, J. D., et al.
Published: (2024) -
Pensieve Discuss: Scalable Small-Group CS Tutoring System with AI
by: Yang, Yoonseok, et al.
Published: (2024) -
The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
by: Niousha, Rose, et al.
Published: (2026) -
ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle
by: Miroyan, Mihran, et al.
Published: (2025) -
Personalized Worked Example Generation from Student Code Submissions Using Pattern-based Knowledge Components
by: Pitts, Griffin, et al.
Published: (2026)