S-GRADES -- Studying Generalization of Student Response Assessments in Diverse Evaluative Settings
Fuente:
arXiv
Guardado en:
| Autores principales: | Seuti, Tasfia, Choudhury, Sagnik Ray |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
por: Sun, Jingyi, et al.
Publicado: (2025)
por: Sun, Jingyi, et al.
Publicado: (2025)
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
por: Jahan, Rafid Ishrak, et al.
Publicado: (2026)
por: Jahan, Rafid Ishrak, et al.
Publicado: (2026)
FutureGen: A RAG-based Approach to Generate the Future Work of Scientific Article
por: Azher, Ibrahim Al, et al.
Publicado: (2025)
por: Azher, Ibrahim Al, et al.
Publicado: (2025)
ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
por: Anik, Anirban Saha, et al.
Publicado: (2025)
por: Anik, Anirban Saha, et al.
Publicado: (2025)
Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
por: Nemkova, Apollinaire Poli, et al.
Publicado: (2025)
por: Nemkova, Apollinaire Poli, et al.
Publicado: (2025)
A Detailed Factor Analysis for the Political Compass Test: Navigating Ideologies of Large Language Models
por: Kamal, Sadia, et al.
Publicado: (2025)
por: Kamal, Sadia, et al.
Publicado: (2025)
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
por: Kabir, Mohsinul, et al.
Publicado: (2025)
por: Kabir, Mohsinul, et al.
Publicado: (2025)
Evaluating Human-LLM Representation Alignment: A Case Study on Affective Sentence Generation for Augmentative and Alternative Communication
por: Choudhury, Shadab, et al.
Publicado: (2025)
por: Choudhury, Shadab, et al.
Publicado: (2025)
A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science
por: Cohn, Clayton, et al.
Publicado: (2024)
por: Cohn, Clayton, et al.
Publicado: (2024)
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
por: Ray, Sourjyadip, et al.
Publicado: (2025)
por: Ray, Sourjyadip, et al.
Publicado: (2025)
Evaluating the Evaluation of Diversity in Commonsense Generation
por: Zhang, Tianhui, et al.
Publicado: (2025)
por: Zhang, Tianhui, et al.
Publicado: (2025)
Towards Automated Situation Awareness: A RAG-Based Framework for Peacebuilding Reports
por: Nemkova, Poli A., et al.
Publicado: (2025)
por: Nemkova, Poli A., et al.
Publicado: (2025)
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
por: Saha, Sougata, et al.
Publicado: (2025)
por: Saha, Sougata, et al.
Publicado: (2025)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
por: Mukherjee, Sagnik, et al.
Publicado: (2024)
por: Mukherjee, Sagnik, et al.
Publicado: (2024)
Evaluating Diversity in Automatic Poetry Generation
por: Chen, Yanran, et al.
Publicado: (2024)
por: Chen, Yanran, et al.
Publicado: (2024)
Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?
por: Knipper, R. Alexander, et al.
Publicado: (2025)
por: Knipper, R. Alexander, et al.
Publicado: (2025)
Evaluating LLM-Generated Q&A Test: a Student-Centered Study
por: Wróblewska, Anna, et al.
Publicado: (2025)
por: Wróblewska, Anna, et al.
Publicado: (2025)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
por: Aisyah, Nurul, et al.
Publicado: (2025)
por: Aisyah, Nurul, et al.
Publicado: (2025)
The Responsible Development of Automated Student Feedback with Generative AI
por: Lindsay, Euan D, et al.
Publicado: (2023)
por: Lindsay, Euan D, et al.
Publicado: (2023)
Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets
por: Gupta, Vatsal, et al.
Publicado: (2023)
por: Gupta, Vatsal, et al.
Publicado: (2023)
Evaluating the Diversity and Quality of LLM Generated Content
por: Shypula, Alexander, et al.
Publicado: (2025)
por: Shypula, Alexander, et al.
Publicado: (2025)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
por: Liu, Yongkang, et al.
Publicado: (2023)
por: Liu, Yongkang, et al.
Publicado: (2023)
Interpretability Framework for LLMs in Undergraduate Calculus
por: Dakshit, Sagnik, et al.
Publicado: (2025)
por: Dakshit, Sagnik, et al.
Publicado: (2025)
DiaCDM: Cognitive Diagnosis in Teacher-Student Dialogues using the Initiation-Response-Evaluation Framework
por: Jia, Rui, et al.
Publicado: (2025)
por: Jia, Rui, et al.
Publicado: (2025)
Asking a Language Model for Diverse Responses
por: Troshin, Sergey, et al.
Publicado: (2025)
por: Troshin, Sergey, et al.
Publicado: (2025)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
por: Schaeffer, Rylan, et al.
Publicado: (2026)
por: Schaeffer, Rylan, et al.
Publicado: (2026)
D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Model
por: Byun, Grace, et al.
Publicado: (2025)
por: Byun, Grace, et al.
Publicado: (2025)
Reproducing the Metric-Based Evaluation of a Set of Controllable Text Generation Techniques
por: Lorandi, Michela, et al.
Publicado: (2024)
por: Lorandi, Michela, et al.
Publicado: (2024)
From NLG Evaluation to Modern Student Assessment in the Era of ChatGPT: The Great Misalignment Problem and Pedagogical Multi-Factor Assessment (P-MFA)
por: Hämäläinen, Mika, et al.
Publicado: (2025)
por: Hämäläinen, Mika, et al.
Publicado: (2025)
CHORUS: Zero-shot Hierarchical Retrieval and Orchestration for Generating Linear Programming Code
por: Ahmed, Tasnim, et al.
Publicado: (2025)
por: Ahmed, Tasnim, et al.
Publicado: (2025)
MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning
por: Bai, Yi, et al.
Publicado: (2026)
por: Bai, Yi, et al.
Publicado: (2026)
To Generate or Discriminate? Methodological Considerations for Measuring Cultural Alignment in LLMs
por: Pandey, Saurabh Kumar, et al.
Publicado: (2026)
por: Pandey, Saurabh Kumar, et al.
Publicado: (2026)
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
por: Hao, Jiangang
Publicado: (2026)
por: Hao, Jiangang
Publicado: (2026)
Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses
por: Peng, Kerui, et al.
Publicado: (2026)
por: Peng, Kerui, et al.
Publicado: (2026)
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
por: Mou, Yutao, et al.
Publicado: (2024)
por: Mou, Yutao, et al.
Publicado: (2024)
GAICo: A Deployed and Extensible Framework for Evaluating Diverse and Multimodal Generative AI Outputs
por: Gupta, Nitin, et al.
Publicado: (2025)
por: Gupta, Nitin, et al.
Publicado: (2025)
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability
por: Takami, Kyosuke, et al.
Publicado: (2026)
por: Takami, Kyosuke, et al.
Publicado: (2026)
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets
por: Samardzic, Tanja, et al.
Publicado: (2024)
por: Samardzic, Tanja, et al.
Publicado: (2024)
Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs
por: Chatterjee, Sagnik, et al.
Publicado: (2026)
por: Chatterjee, Sagnik, et al.
Publicado: (2026)
Can Public LLMs be used for Self-Diagnosis of Medical Conditions ?
por: Balasubramanian, Nikil Sharan Prabahar, et al.
Publicado: (2024)
por: Balasubramanian, Nikil Sharan Prabahar, et al.
Publicado: (2024)
Ejemplares similares
-
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
por: Sun, Jingyi, et al.
Publicado: (2025) -
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
por: Jahan, Rafid Ishrak, et al.
Publicado: (2026) -
FutureGen: A RAG-based Approach to Generate the Future Work of Scientific Article
por: Azher, Ibrahim Al, et al.
Publicado: (2025) -
ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
por: Anik, Anirban Saha, et al.
Publicado: (2025) -
Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
por: Nemkova, Apollinaire Poli, et al.
Publicado: (2025)