EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Buyuan, Hu, Shiyu, Ma, Yiping, Zhang, Yuanming, Cheong, Kang Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EduVerse: A User-Defined Multi-Agent Simulation Space for Education Scenario
by: Ma, Yiping, et al.
Published: (2025)
by: Ma, Yiping, et al.
Published: (2025)
When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction
by: Ma, Yiping, et al.
Published: (2024)
by: Ma, Yiping, et al.
Published: (2024)
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
by: Lucy, Li, et al.
Published: (2026)
by: Lucy, Li, et al.
Published: (2026)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
by: Hu, Shiyu, et al.
Published: (2024)
by: Hu, Shiyu, et al.
Published: (2024)
Consensus and Subjectivity of Skin Tone Annotation for ML Fairness
by: Schumann, Candice, et al.
Published: (2023)
by: Schumann, Candice, et al.
Published: (2023)
Synthesizing Images on Perceptual Boundaries of ANNs for Uncovering Human Perceptual Variability on Facial Expressions
by: Deng, Haotian, et al.
Published: (2025)
by: Deng, Haotian, et al.
Published: (2025)
AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images
by: Zhang, Bo, et al.
Published: (2026)
by: Zhang, Bo, et al.
Published: (2026)
ViMU: Benchmarking Video Metaphorical Understanding
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
by: Ji, Haonian, et al.
Published: (2025)
by: Ji, Haonian, et al.
Published: (2025)
Student Classroom Behavior Recognition Based on Improved YOLOv8s
by: Gao, Xiang, et al.
Published: (2026)
by: Gao, Xiang, et al.
Published: (2026)
Exploring Student Perception on Gen AI Adoption in Higher Education: A Descriptive Study
by: Singh, Harpreet, et al.
Published: (2026)
by: Singh, Harpreet, et al.
Published: (2026)
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
ChaosBench: A Multi-Channel, Physics-Based Benchmark for Subseasonal-to-Seasonal Climate Prediction
by: Nathaniel, Juan, et al.
Published: (2024)
by: Nathaniel, Juan, et al.
Published: (2024)
Breaking the Global North Stereotype: A Global South-centric Benchmark Dataset for Auditing and Mitigating Biases in Facial Recognition Systems
by: Jaiswal, Siddharth D, et al.
Published: (2024)
by: Jaiswal, Siddharth D, et al.
Published: (2024)
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
FairFedMed: Benchmarking Group Fairness in Federated Medical Imaging with FairLoRA
by: Li, Minghan, et al.
Published: (2025)
by: Li, Minghan, et al.
Published: (2025)
Determining the Difficulties of Students With Dyslexia via Virtual Reality and Artificial Intelligence: An Exploratory Analysis
by: Yeguas-Bolívar, Enrique, et al.
Published: (2024)
by: Yeguas-Bolívar, Enrique, et al.
Published: (2024)
EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos
by: Chen, Baoliang, et al.
Published: (2026)
by: Chen, Baoliang, et al.
Published: (2026)
WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
by: Sultan, Rafi Ibn, et al.
Published: (2026)
by: Sultan, Rafi Ibn, et al.
Published: (2026)
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection
by: Hu, Xiaobin, et al.
Published: (2026)
by: Hu, Xiaobin, et al.
Published: (2026)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
by: Jin, Yiqiao, et al.
Published: (2024)
by: Jin, Yiqiao, et al.
Published: (2024)
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
by: Bu, Wendong, et al.
Published: (2025)
by: Bu, Wendong, et al.
Published: (2025)
SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
by: Wang, Yipei, et al.
Published: (2025)
by: Wang, Yipei, et al.
Published: (2025)
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
by: Baral, Sami, et al.
Published: (2025)
by: Baral, Sami, et al.
Published: (2025)
CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models
by: Zhu, Yiqi, et al.
Published: (2025)
by: Zhu, Yiqi, et al.
Published: (2025)
STAMP: Multi-pattern Attention-aware Multiple Instance Learning for STAS Diagnosis in Multi-center Histopathology Images
by: Pan, Liangrui, et al.
Published: (2025)
by: Pan, Liangrui, et al.
Published: (2025)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
Open-Vocabulary X-ray Prohibited Item Detection via Fine-tuning CLIP
by: Lin, Shuyang, et al.
Published: (2024)
by: Lin, Shuyang, et al.
Published: (2024)
CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
OrthoEraser: Coupled-Neuron Orthogonal Projection for Concept Erasure
by: Shi, Chuancheng, et al.
Published: (2026)
by: Shi, Chuancheng, et al.
Published: (2026)
Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
by: Chandra, Nuria Alina, et al.
Published: (2025)
by: Chandra, Nuria Alina, et al.
Published: (2025)
MLLM-as-a-Judge for Image Safety without Human Labeling
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models
by: Xiaobin, Hu, et al.
Published: (2025)
by: Xiaobin, Hu, et al.
Published: (2025)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
by: Shen, Yixuan, et al.
Published: (2026)
by: Shen, Yixuan, et al.
Published: (2026)
OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
by: Chang, Haochen, et al.
Published: (2025)
by: Chang, Haochen, et al.
Published: (2025)
Face4FairShifts: A Large Image Benchmark for Fairness and Robust Learning across Visual Domains
by: Lin, Yumeng, et al.
Published: (2025)
by: Lin, Yumeng, et al.
Published: (2025)
EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions
by: Sun, Weiyu, et al.
Published: (2026)
by: Sun, Weiyu, et al.
Published: (2026)
Similar Items
-
EduVerse: A User-Defined Multi-Agent Simulation Space for Education Scenario
by: Ma, Yiping, et al.
Published: (2025) -
When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction
by: Ma, Yiping, et al.
Published: (2024) -
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
by: Lucy, Li, et al.
Published: (2026) -
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
by: Hu, Shiyu, et al.
Published: (2024) -
Consensus and Subjectivity of Skin Tone Annotation for ML Fairness
by: Schumann, Candice, et al.
Published: (2023)