Evaluating LLM-Generated Q&A Test: a Student-Centered Study
Fuente:
arXiv
Saved in:
| Main Authors: | Wróblewska, Anna, Grabek, Bartosz, Świstak, Jakub, Dan, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Make Museums More Interactive? Case Study of Artistic Chatbot
by: Kucia, Filip J., et al.
Published: (2025)
by: Kucia, Filip J., et al.
Published: (2025)
Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models
by: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Published: (2025)
by: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Published: (2025)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
by: Martin-Boyle, Anna, et al.
Published: (2026)
by: Martin-Boyle, Anna, et al.
Published: (2026)
Evaluating LLM-Generated Lessons from the Language Learning Students' Perspective: A Short Case Study on Duolingo
by: Catalan, Carlos Rafael, et al.
Published: (2026)
by: Catalan, Carlos Rafael, et al.
Published: (2026)
MEGAnno+: A Human-LLM Collaborative Annotation System
by: Kim, Hannah, et al.
Published: (2024)
by: Kim, Hannah, et al.
Published: (2024)
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
QuaLLM: An LLM-based Framework to Extract Quantitative Insights from Online Forums
by: Rao, Varun Nagaraj, et al.
Published: (2024)
by: Rao, Varun Nagaraj, et al.
Published: (2024)
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
PRAISE: Enhancing Product Descriptions with LLM-Driven Structured Insights
by: Qidwai, Adnan, et al.
Published: (2025)
by: Qidwai, Adnan, et al.
Published: (2025)
From Surface Learning to Deep Understanding: A Grounded AI Tutoring System for Moodle
by: Ostrowska, Anna, et al.
Published: (2026)
by: Ostrowska, Anna, et al.
Published: (2026)
A Formative Study of Brief Affective Text as a Complement to Wearable Sensing for Longitudinal Student Health Monitoring
by: Harry, Tamunotonye, et al.
Published: (2026)
by: Harry, Tamunotonye, et al.
Published: (2026)
BADGE: BADminton report Generation and Evaluation with LLM
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
by: Shaikh, Omar, et al.
Published: (2025)
by: Shaikh, Omar, et al.
Published: (2025)
A Comparative Study on Annotation Quality of Crowdsourcing and LLM via Label Aggregation
by: Li, Jiyi
Published: (2024)
by: Li, Jiyi
Published: (2024)
VegaChat: A Robust Framework for LLM-Based Chart Generation and Assessment
by: Hostnik, Marko, et al.
Published: (2026)
by: Hostnik, Marko, et al.
Published: (2026)
Is Passive Expertise-Based Personalization Enough? A Case Study in AI-Assisted Test-Taking
by: Siyan, Li, et al.
Published: (2025)
by: Siyan, Li, et al.
Published: (2025)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
Telephone Surveys Meet Conversational AI: Evaluating a LLM-Based Telephone Survey System at Scale
by: Lang, Max M., et al.
Published: (2025)
by: Lang, Max M., et al.
Published: (2025)
From Chatbots to Confidants: A Cross-Cultural Study of LLM Adoption for Emotional Support
by: Amat-Lefort, Natalia, et al.
Published: (2026)
by: Amat-Lefort, Natalia, et al.
Published: (2026)
Conversations in Space: Structuring Non-Linear LLM Interactions on a Canvas
by: Amin, Rifat Mehreen, et al.
Published: (2026)
by: Amin, Rifat Mehreen, et al.
Published: (2026)
Using Generative Text Models to Create Qualitative Codebooks for Student Evaluations of Teaching
by: Katz, Andrew, et al.
Published: (2024)
by: Katz, Andrew, et al.
Published: (2024)
Grounding Gaps in Language Model Generations
by: Shaikh, Omar, et al.
Published: (2023)
by: Shaikh, Omar, et al.
Published: (2023)
Product vs. Process: Exploring EFL Students' Editing of AI-Generated Text for Expository Writing
by: Woo, David James, et al.
Published: (2025)
by: Woo, David James, et al.
Published: (2025)
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
by: Zengaffinen, Yanick, et al.
Published: (2026)
by: Zengaffinen, Yanick, et al.
Published: (2026)
Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
by: Amin, Danial, et al.
Published: (2025)
by: Amin, Danial, et al.
Published: (2025)
Sketch Then Generate: Providing Incremental User Feedback and Guiding LLM Code Generation through Language-Oriented Code Sketches
by: Zhu-Tian, Chen, et al.
Published: (2024)
by: Zhu-Tian, Chen, et al.
Published: (2024)
Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
by: He, Gaole, et al.
Published: (2025)
by: He, Gaole, et al.
Published: (2025)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
by: Afzoon, Saleh, et al.
Published: (2026)
by: Afzoon, Saleh, et al.
Published: (2026)
Digital assistant in a point of sales
by: Lesiak, Emilia, et al.
Published: (2024)
by: Lesiak, Emilia, et al.
Published: (2024)
Online vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots
by: Svikhnushina, Ekaterina, et al.
Published: (2024)
by: Svikhnushina, Ekaterina, et al.
Published: (2024)
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
by: Martin-Boyle, Anna, et al.
Published: (2026)
by: Martin-Boyle, Anna, et al.
Published: (2026)
What Makes LLM Agent Simulations Useful for Policy Practice? An Iterative Design Study in Emergency Preparedness
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Evaluating the Application of ChatGPT in Outpatient Triage Guidance: A Comparative Study
by: Liu, Dou, et al.
Published: (2024)
by: Liu, Dou, et al.
Published: (2024)
Simulating Classroom Education with LLM-Empowered Agents
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
One-Topic-Doesn't-Fit-All: Transcreating Reading Comprehension Test for Personalized Learning
by: Han, Jieun, et al.
Published: (2025)
by: Han, Jieun, et al.
Published: (2025)
A-MEM: Agentic Memory for LLM Agents
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
Retrieval-Augmented Generation of Pediatric Speech-Language Pathology vignettes: A Proof-of-Concept Study
by: Liu, Yilan
Published: (2025)
by: Liu, Yilan
Published: (2025)
Similar Items
-
How to Make Museums More Interactive? Case Study of Artistic Chatbot
by: Kucia, Filip J., et al.
Published: (2025) -
Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models
by: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Published: (2025) -
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
by: Sung, Yoo Yeon, et al.
Published: (2025) -
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
by: Martin-Boyle, Anna, et al.
Published: (2026) -
Evaluating LLM-Generated Lessons from the Language Learning Students' Perspective: A Short Case Study on Duolingo
by: Catalan, Carlos Rafael, et al.
Published: (2026)