Domain-Grounded Evaluation of LLMs in International Student Knowledge
Fuente:
arXiv
Guardado en:
| Autores principales: | Daitx, Claudinei, Amar, Haitham |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Knowledge-Data Fusion Based Source-Free Semi-Supervised Domain Adaptation for Seizure Subtype Classification
por: Peng, Ruimin, et al.
Publicado: (2024)
por: Peng, Ruimin, et al.
Publicado: (2024)
LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems
por: Li, Chu, et al.
Publicado: (2024)
por: Li, Chu, et al.
Publicado: (2024)
OAK -- Onboarding with Actionable Knowledge
por: Devènes, Steve, et al.
Publicado: (2025)
por: Devènes, Steve, et al.
Publicado: (2025)
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
por: Park, Joon Sung, et al.
Publicado: (2024)
por: Park, Joon Sung, et al.
Publicado: (2024)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
por: Fu, Xiao, et al.
Publicado: (2025)
por: Fu, Xiao, et al.
Publicado: (2025)
Teaching According to Students' Aptitude: Personalized Mathematics Tutoring via Persona-, Memory-, and Forgetting-Aware LLMs
por: Wu, Yang, et al.
Publicado: (2025)
por: Wu, Yang, et al.
Publicado: (2025)
Enabling On-Device LLMs Personalization with Smartphone Sensing
por: Zhang, Shiquan, et al.
Publicado: (2024)
por: Zhang, Shiquan, et al.
Publicado: (2024)
Assessing the Real-World Utility of Explainable AI for Arousal Diagnostics: An Application-Grounded User Study
por: Kraft, Stefan, et al.
Publicado: (2025)
por: Kraft, Stefan, et al.
Publicado: (2025)
ConvoLearn: A Learning Sciences Grounded Dataset for Fine-Tuning Dialogic AI Tutors
por: Sharma, Mayank, et al.
Publicado: (2026)
por: Sharma, Mayank, et al.
Publicado: (2026)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
por: Zeng, Qiuhai, et al.
Publicado: (2025)
por: Zeng, Qiuhai, et al.
Publicado: (2025)
The Hardness of Achieving Impact in AI for Social Impact Research: A Ground-Level View of Challenges & Opportunities
por: Majumdar, Aditya, et al.
Publicado: (2025)
por: Majumdar, Aditya, et al.
Publicado: (2025)
Detection of adversarial intent in Human-AI teams using LLMs
por: Musaffar, Abed K., et al.
Publicado: (2026)
por: Musaffar, Abed K., et al.
Publicado: (2026)
Sentiment Analysis in Learning Management Systems Understanding Student Feedback at Scale
por: Almutairi, Mohammed
Publicado: (2025)
por: Almutairi, Mohammed
Publicado: (2025)
Domain-Adversarial Anatomical Graph Networks for Cross-User Human Activity Recognition
por: Ye, Xiaozhou, et al.
Publicado: (2025)
por: Ye, Xiaozhou, et al.
Publicado: (2025)
Diagrammatization and Abduction to Improve AI Interpretability With Domain-Aligned Explanations for Medical Diagnosis
por: Lim, Brian Y., et al.
Publicado: (2023)
por: Lim, Brian Y., et al.
Publicado: (2023)
STDA-Net: Spectrogram-Based Domain Adaptation for cross-dataset Sleep Stage Classification
por: Tallal, Unaza, et al.
Publicado: (2026)
por: Tallal, Unaza, et al.
Publicado: (2026)
Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
por: Xu, Songlin, et al.
Publicado: (2025)
por: Xu, Songlin, et al.
Publicado: (2025)
Adversarial Domain Adaptation for Cross-user Activity Recognition Using Diffusion-based Noise-centred Learning
por: Ye, Xiaozhou, et al.
Publicado: (2024)
por: Ye, Xiaozhou, et al.
Publicado: (2024)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
por: Soni, Nikita, et al.
Publicado: (2025)
por: Soni, Nikita, et al.
Publicado: (2025)
LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents
por: Xiao, Chang, et al.
Publicado: (2024)
por: Xiao, Chang, et al.
Publicado: (2024)
Benchmarking System Dynamics AI Assistants: Cloud Versus Local LLMs on CLD Extraction and Discussion
por: Leitch, Terry
Publicado: (2026)
por: Leitch, Terry
Publicado: (2026)
LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams
por: Justus, Aju Ani, et al.
Publicado: (2025)
por: Justus, Aju Ani, et al.
Publicado: (2025)
InFiConD: Interactive No-code Fine-tuning with Concept-based Knowledge Distillation
por: Huang, Jinbin, et al.
Publicado: (2024)
por: Huang, Jinbin, et al.
Publicado: (2024)
Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
por: Albers, Nele, et al.
Publicado: (2025)
por: Albers, Nele, et al.
Publicado: (2025)
Efficient Speech Command Recognition Leveraging Spiking Neural Network and Curriculum Learning-based Knowledge Distillation
por: Wang, Jiaqi, et al.
Publicado: (2024)
por: Wang, Jiaqi, et al.
Publicado: (2024)
TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health
por: Fan, Yuang, et al.
Publicado: (2026)
por: Fan, Yuang, et al.
Publicado: (2026)
A Systematic Review of User-Centred Evaluation of Explainable AI in Healthcare
por: Donoso-Guzmán, Ivania, et al.
Publicado: (2025)
por: Donoso-Guzmán, Ivania, et al.
Publicado: (2025)
Evaluation of Human-Understandability of Global Model Explanations using Decision Tree
por: Sivaprasad, Adarsa, et al.
Publicado: (2023)
por: Sivaprasad, Adarsa, et al.
Publicado: (2023)
Harmonic LLMs are Trustworthy
por: Kersting, Nicholas S., et al.
Publicado: (2024)
por: Kersting, Nicholas S., et al.
Publicado: (2024)
Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding
por: Baer, Gregor, et al.
Publicado: (2026)
por: Baer, Gregor, et al.
Publicado: (2026)
Evaluating Deep Networks for Detecting User Familiarity with VR from Hand Interactions
por: Li, Mingjun, et al.
Publicado: (2024)
por: Li, Mingjun, et al.
Publicado: (2024)
Roamify: Designing and Evaluating an LLM Based Google Chrome Extension for Personalised Itinerary Planning
por: Udandarao, Vikranth, et al.
Publicado: (2025)
por: Udandarao, Vikranth, et al.
Publicado: (2025)
Measuring What Matters: Connecting AI Ethics Evaluations to System Attributes, Hazards, and Harms
por: Rismani, Shalaleh, et al.
Publicado: (2025)
por: Rismani, Shalaleh, et al.
Publicado: (2025)
Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression Experiments
por: Boggust, Angie, et al.
Publicado: (2024)
por: Boggust, Angie, et al.
Publicado: (2024)
A Systematic Evaluation of Self-Supervised Learning for Label-Efficient Sleep Staging with Wearable EEG
por: Estevan, Emilio, et al.
Publicado: (2025)
por: Estevan, Emilio, et al.
Publicado: (2025)
VisMoDAl: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language Models
por: Wang, Huanchen, et al.
Publicado: (2025)
por: Wang, Huanchen, et al.
Publicado: (2025)
Reassessing Evaluation Functions in Algorithmic Recourse: An Empirical Study from a Human-Centered Perspective
por: Tominaga, Tomu, et al.
Publicado: (2024)
por: Tominaga, Tomu, et al.
Publicado: (2024)
Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings
por: Silva, Inês Oliveira e, et al.
Publicado: (2026)
por: Silva, Inês Oliveira e, et al.
Publicado: (2026)
Vi(E)va LLM! A Conceptual Stack for Evaluating and Interpreting Generative AI-based Visualizations
por: Podo, Luca, et al.
Publicado: (2024)
por: Podo, Luca, et al.
Publicado: (2024)
Evaluating Interactive 2D Visualization as a Sample Selection Strategy for Biomedical Time-Series Data Annotation
por: Vaaras, Einari, et al.
Publicado: (2026)
por: Vaaras, Einari, et al.
Publicado: (2026)
Ejemplares similares
-
Knowledge-Data Fusion Based Source-Free Semi-Supervised Domain Adaptation for Seizure Subtype Classification
por: Peng, Ruimin, et al.
Publicado: (2024) -
LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems
por: Li, Chu, et al.
Publicado: (2024) -
OAK -- Onboarding with Actionable Knowledge
por: Devènes, Steve, et al.
Publicado: (2025) -
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
por: Park, Joon Sung, et al.
Publicado: (2024) -
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
por: Fu, Xiao, et al.
Publicado: (2025)