Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Apartsin, Alexander, Sason, Omri, Aperstein, Yehudit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis
by: Aperstein, Yehudit, et al.
Published: (2026)
by: Aperstein, Yehudit, et al.
Published: (2026)
SeaAlert: Critical Information Extraction From Maritime Distress Communications with Large Language Models
by: Atia, Tomer, et al.
Published: (2026)
by: Atia, Tomer, et al.
Published: (2026)
Reliable Extraction of Clinical Follow-Up Instructions: A Hybrid Neural-Symbolic Pipeline
by: Laufer, Michal, et al.
Published: (2026)
by: Laufer, Michal, et al.
Published: (2026)
IRC-Bench: Recognizing Entities from Contextual Cues in First-Person Reminiscences
by: Aperstein, Yehudit, et al.
Published: (2026)
by: Aperstein, Yehudit, et al.
Published: (2026)
From Joy to Fear: A Benchmark of Emotion Estimation in Pop Song Lyrics
by: Dahary, Shay, et al.
Published: (2025)
by: Dahary, Shay, et al.
Published: (2025)
LLM-guided headline rewriting for clickability enhancement without clickbait
by: Aperstein, Yehudit, et al.
Published: (2026)
by: Aperstein, Yehudit, et al.
Published: (2026)
DGLD: Domain-Gated Latent Diffusion for the Discovery of Novel Energetic Materials
by: Aperstein, Yehudit, et al.
Published: (2026)
by: Aperstein, Yehudit, et al.
Published: (2026)
Acting on the Unseen: Communication-Free Collaborative Filtering for Decentralized Multi-Robot Task Allocation
by: Apartsin, Alexander, et al.
Published: (2026)
by: Apartsin, Alexander, et al.
Published: (2026)
From Fuzzy Speech to Medical Insight: Benchmarking LLMs on Noisy Patient Narratives
by: Mama, Eden, et al.
Published: (2025)
by: Mama, Eden, et al.
Published: (2025)
Reading Between the Lines: Classifying Resume Seniority with Large Language Models
by: Cohen, Matan, et al.
Published: (2025)
by: Cohen, Matan, et al.
Published: (2025)
Mapping License Plate Recoverability Under Extreme Viewing Angles for Oppor-tunistic Urban Sensing
by: Adamenko, Igor, et al.
Published: (2026)
by: Adamenko, Igor, et al.
Published: (2026)
Do Large Language Models Need Intent? Revisiting Response Generation Strategies for Service Assistant
by: Bolshinsky, Inbal, et al.
Published: (2025)
by: Bolshinsky, Inbal, et al.
Published: (2025)
An Interpretable Benchmark for Clickbait Detection and Tactic Attribution
by: Nofar, Lihi, et al.
Published: (2025)
by: Nofar, Lihi, et al.
Published: (2025)
CalexNet: Soft Cascade-Aligned Training and Calibration for Lightweight Early-Exit Branches
by: Aperstein, Yehudit, et al.
Published: (2025)
by: Aperstein, Yehudit, et al.
Published: (2025)
When Curiosity Signals Danger: Predicting Health Crises Through Online Medication Inquiries
by: Goncharok, Dvora, et al.
Published: (2025)
by: Goncharok, Dvora, et al.
Published: (2025)
Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
by: Aperstein, Yehudit, et al.
Published: (2025)
by: Aperstein, Yehudit, et al.
Published: (2025)
Code Review Without Borders: Evaluating Synthetic vs. Real Data for Review Recommendation
by: Cohen, Yogev, et al.
Published: (2025)
by: Cohen, Yogev, et al.
Published: (2025)
PTEENet: Post-Trained Early-Exit Neural Networks Augmentation for Inference Cost Optimization
by: Lahiany, Assaf, et al.
Published: (2025)
by: Lahiany, Assaf, et al.
Published: (2025)
Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
by: Geng, Yilin, et al.
Published: (2025)
by: Geng, Yilin, et al.
Published: (2025)
Measuring Pragmatic Influence in Large Language Model Instructions
by: Geng, Yilin, et al.
Published: (2026)
by: Geng, Yilin, et al.
Published: (2026)
Stitching the Story: Creating Panoramic Incident Summaries from Body-Worn Footage
by: Cohen, Dor, et al.
Published: (2025)
by: Cohen, Dor, et al.
Published: (2025)
Beyond Words: Interjection Classification for Improved Human-Computer Interaction
by: Goren, Yaniv, et al.
Published: (2025)
by: Goren, Yaniv, et al.
Published: (2025)
Towards a Benchmark for Large Language Models for Business Process Management Tasks
by: Busch, Kiran, et al.
Published: (2024)
by: Busch, Kiran, et al.
Published: (2024)
CONTESTS: a Framework for Consistency Testing of Span Probabilities in Language Models
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
by: Tang, Weizhi, et al.
Published: (2024)
by: Tang, Weizhi, et al.
Published: (2024)
Generating Benchmarks for Factuality Evaluation of Language Models
by: Muhlgay, Dor, et al.
Published: (2023)
by: Muhlgay, Dor, et al.
Published: (2023)
The Imperfect Learner: Incorporating Developmental Trajectories in Memory-based Student Simulation
by: Liu, Zhengyuan, et al.
Published: (2025)
by: Liu, Zhengyuan, et al.
Published: (2025)
Exploring the Learning Capabilities of Language Models using LEVERWORLDS
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
Codenames as a Benchmark for Large Language Models
by: Stephenson, Matthew, et al.
Published: (2024)
by: Stephenson, Matthew, et al.
Published: (2024)
Enhancing Commentary Strategies for Imperfect Information Card Games: A Study of Large Language Models in Guandan Commentary
by: Tao, Meiling, et al.
Published: (2024)
by: Tao, Meiling, et al.
Published: (2024)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
by: Ball, Sarah, et al.
Published: (2025)
by: Ball, Sarah, et al.
Published: (2025)
Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models
by: Tuan, Yi-Lin, et al.
Published: (2024)
by: Tuan, Yi-Lin, et al.
Published: (2024)
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
by: Zhong, Chengzhi, et al.
Published: (2025)
by: Zhong, Chengzhi, et al.
Published: (2025)
Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Benchmarking Distributional Alignment of Large Language Models
by: Meister, Nicole, et al.
Published: (2024)
by: Meister, Nicole, et al.
Published: (2024)
Benchmarking the Pedagogical Knowledge of Large Language Models
by: Lelièvre, Maxime, et al.
Published: (2025)
by: Lelièvre, Maxime, et al.
Published: (2025)
Large Language Model Benchmarks in Medical Tasks
by: Yan, Lawrence K. Q., et al.
Published: (2024)
by: Yan, Lawrence K. Q., et al.
Published: (2024)
BELL: Benchmarking the Explainability of Large Language Models
by: Ahmed, Syed Quiser, et al.
Published: (2025)
by: Ahmed, Syed Quiser, et al.
Published: (2025)
Large Language Models for In-Context Student Modeling: Synthesizing Student's Behavior in Visual Programming
by: Nguyen, Manh Hung, et al.
Published: (2023)
by: Nguyen, Manh Hung, et al.
Published: (2023)
Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs
by: Wagner, Eitan, et al.
Published: (2025)
by: Wagner, Eitan, et al.
Published: (2025)
Similar Items
-
A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis
by: Aperstein, Yehudit, et al.
Published: (2026) -
SeaAlert: Critical Information Extraction From Maritime Distress Communications with Large Language Models
by: Atia, Tomer, et al.
Published: (2026) -
Reliable Extraction of Clinical Follow-Up Instructions: A Hybrid Neural-Symbolic Pipeline
by: Laufer, Michal, et al.
Published: (2026) -
IRC-Bench: Recognizing Entities from Contextual Cues in First-Person Reminiscences
by: Aperstein, Yehudit, et al.
Published: (2026) -
From Joy to Fear: A Benchmark of Emotion Estimation in Pop Song Lyrics
by: Dahary, Shay, et al.
Published: (2025)