S-GRADES -- Studying Generalization of Student Response Assessments in Diverse Evaluative Settings
Fuente:
arXiv
Salvato in:
| Autori principali: | Seuti, Tasfia, Choudhury, Sagnik Ray |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
di: Sun, Jingyi, et al.
Pubblicazione: (2025)
di: Sun, Jingyi, et al.
Pubblicazione: (2025)
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
di: Jahan, Rafid Ishrak, et al.
Pubblicazione: (2026)
di: Jahan, Rafid Ishrak, et al.
Pubblicazione: (2026)
FutureGen: A RAG-based Approach to Generate the Future Work of Scientific Article
di: Azher, Ibrahim Al, et al.
Pubblicazione: (2025)
di: Azher, Ibrahim Al, et al.
Pubblicazione: (2025)
ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
di: Anik, Anirban Saha, et al.
Pubblicazione: (2025)
di: Anik, Anirban Saha, et al.
Pubblicazione: (2025)
Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
di: Nemkova, Apollinaire Poli, et al.
Pubblicazione: (2025)
di: Nemkova, Apollinaire Poli, et al.
Pubblicazione: (2025)
A Detailed Factor Analysis for the Political Compass Test: Navigating Ideologies of Large Language Models
di: Kamal, Sadia, et al.
Pubblicazione: (2025)
di: Kamal, Sadia, et al.
Pubblicazione: (2025)
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
Evaluating Human-LLM Representation Alignment: A Case Study on Affective Sentence Generation for Augmentative and Alternative Communication
di: Choudhury, Shadab, et al.
Pubblicazione: (2025)
di: Choudhury, Shadab, et al.
Pubblicazione: (2025)
A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science
di: Cohn, Clayton, et al.
Pubblicazione: (2024)
di: Cohn, Clayton, et al.
Pubblicazione: (2024)
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
di: Ray, Sourjyadip, et al.
Pubblicazione: (2025)
di: Ray, Sourjyadip, et al.
Pubblicazione: (2025)
Evaluating the Evaluation of Diversity in Commonsense Generation
di: Zhang, Tianhui, et al.
Pubblicazione: (2025)
di: Zhang, Tianhui, et al.
Pubblicazione: (2025)
Towards Automated Situation Awareness: A RAG-Based Framework for Peacebuilding Reports
di: Nemkova, Poli A., et al.
Pubblicazione: (2025)
di: Nemkova, Poli A., et al.
Pubblicazione: (2025)
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
di: Saha, Sougata, et al.
Pubblicazione: (2025)
di: Saha, Sougata, et al.
Pubblicazione: (2025)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2024)
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2024)
Evaluating Diversity in Automatic Poetry Generation
di: Chen, Yanran, et al.
Pubblicazione: (2024)
di: Chen, Yanran, et al.
Pubblicazione: (2024)
Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?
di: Knipper, R. Alexander, et al.
Pubblicazione: (2025)
di: Knipper, R. Alexander, et al.
Pubblicazione: (2025)
Evaluating LLM-Generated Q&A Test: a Student-Centered Study
di: Wróblewska, Anna, et al.
Pubblicazione: (2025)
di: Wróblewska, Anna, et al.
Pubblicazione: (2025)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
di: Aisyah, Nurul, et al.
Pubblicazione: (2025)
di: Aisyah, Nurul, et al.
Pubblicazione: (2025)
The Responsible Development of Automated Student Feedback with Generative AI
di: Lindsay, Euan D, et al.
Pubblicazione: (2023)
di: Lindsay, Euan D, et al.
Pubblicazione: (2023)
Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets
di: Gupta, Vatsal, et al.
Pubblicazione: (2023)
di: Gupta, Vatsal, et al.
Pubblicazione: (2023)
Evaluating the Diversity and Quality of LLM Generated Content
di: Shypula, Alexander, et al.
Pubblicazione: (2025)
di: Shypula, Alexander, et al.
Pubblicazione: (2025)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
di: Liu, Yongkang, et al.
Pubblicazione: (2023)
di: Liu, Yongkang, et al.
Pubblicazione: (2023)
Interpretability Framework for LLMs in Undergraduate Calculus
di: Dakshit, Sagnik, et al.
Pubblicazione: (2025)
di: Dakshit, Sagnik, et al.
Pubblicazione: (2025)
DiaCDM: Cognitive Diagnosis in Teacher-Student Dialogues using the Initiation-Response-Evaluation Framework
di: Jia, Rui, et al.
Pubblicazione: (2025)
di: Jia, Rui, et al.
Pubblicazione: (2025)
Asking a Language Model for Diverse Responses
di: Troshin, Sergey, et al.
Pubblicazione: (2025)
di: Troshin, Sergey, et al.
Pubblicazione: (2025)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
di: Schaeffer, Rylan, et al.
Pubblicazione: (2026)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2026)
D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Model
di: Byun, Grace, et al.
Pubblicazione: (2025)
di: Byun, Grace, et al.
Pubblicazione: (2025)
Reproducing the Metric-Based Evaluation of a Set of Controllable Text Generation Techniques
di: Lorandi, Michela, et al.
Pubblicazione: (2024)
di: Lorandi, Michela, et al.
Pubblicazione: (2024)
From NLG Evaluation to Modern Student Assessment in the Era of ChatGPT: The Great Misalignment Problem and Pedagogical Multi-Factor Assessment (P-MFA)
di: Hämäläinen, Mika, et al.
Pubblicazione: (2025)
di: Hämäläinen, Mika, et al.
Pubblicazione: (2025)
CHORUS: Zero-shot Hierarchical Retrieval and Orchestration for Generating Linear Programming Code
di: Ahmed, Tasnim, et al.
Pubblicazione: (2025)
di: Ahmed, Tasnim, et al.
Pubblicazione: (2025)
MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning
di: Bai, Yi, et al.
Pubblicazione: (2026)
di: Bai, Yi, et al.
Pubblicazione: (2026)
To Generate or Discriminate? Methodological Considerations for Measuring Cultural Alignment in LLMs
di: Pandey, Saurabh Kumar, et al.
Pubblicazione: (2026)
di: Pandey, Saurabh Kumar, et al.
Pubblicazione: (2026)
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
di: Hao, Jiangang
Pubblicazione: (2026)
di: Hao, Jiangang
Pubblicazione: (2026)
Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses
di: Peng, Kerui, et al.
Pubblicazione: (2026)
di: Peng, Kerui, et al.
Pubblicazione: (2026)
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
di: Mou, Yutao, et al.
Pubblicazione: (2024)
di: Mou, Yutao, et al.
Pubblicazione: (2024)
GAICo: A Deployed and Extensible Framework for Evaluating Diverse and Multimodal Generative AI Outputs
di: Gupta, Nitin, et al.
Pubblicazione: (2025)
di: Gupta, Nitin, et al.
Pubblicazione: (2025)
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability
di: Takami, Kyosuke, et al.
Pubblicazione: (2026)
di: Takami, Kyosuke, et al.
Pubblicazione: (2026)
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets
di: Samardzic, Tanja, et al.
Pubblicazione: (2024)
di: Samardzic, Tanja, et al.
Pubblicazione: (2024)
Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs
di: Chatterjee, Sagnik, et al.
Pubblicazione: (2026)
di: Chatterjee, Sagnik, et al.
Pubblicazione: (2026)
Can Public LLMs be used for Self-Diagnosis of Medical Conditions ?
di: Balasubramanian, Nikil Sharan Prabahar, et al.
Pubblicazione: (2024)
di: Balasubramanian, Nikil Sharan Prabahar, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
di: Sun, Jingyi, et al.
Pubblicazione: (2025) -
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
di: Jahan, Rafid Ishrak, et al.
Pubblicazione: (2026) -
FutureGen: A RAG-based Approach to Generate the Future Work of Scientific Article
di: Azher, Ibrahim Al, et al.
Pubblicazione: (2025) -
ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
di: Anik, Anirban Saha, et al.
Pubblicazione: (2025) -
Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
di: Nemkova, Apollinaire Poli, et al.
Pubblicazione: (2025)