Saved in:
| Main Authors: | Seuti, Tasfia, Choudhury, Sagnik Ray |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.10233 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
by: Sun, Jingyi, et al.
Published: (2025)
by: Sun, Jingyi, et al.
Published: (2025)
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
by: Jahan, Rafid Ishrak, et al.
Published: (2026)
by: Jahan, Rafid Ishrak, et al.
Published: (2026)
FutureGen: A RAG-based Approach to Generate the Future Work of Scientific Article
by: Azher, Ibrahim Al, et al.
Published: (2025)
by: Azher, Ibrahim Al, et al.
Published: (2025)
ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
by: Anik, Anirban Saha, et al.
Published: (2025)
by: Anik, Anirban Saha, et al.
Published: (2025)
Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
by: Nemkova, Apollinaire Poli, et al.
Published: (2025)
by: Nemkova, Apollinaire Poli, et al.
Published: (2025)
A Detailed Factor Analysis for the Political Compass Test: Navigating Ideologies of Large Language Models
by: Kamal, Sadia, et al.
Published: (2025)
by: Kamal, Sadia, et al.
Published: (2025)
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
Towards Automated Situation Awareness: A RAG-Based Framework for Peacebuilding Reports
by: Nemkova, Poli A., et al.
Published: (2025)
by: Nemkova, Poli A., et al.
Published: (2025)
Evaluating Human-LLM Representation Alignment: A Case Study on Affective Sentence Generation for Augmentative and Alternative Communication
by: Choudhury, Shadab, et al.
Published: (2025)
by: Choudhury, Shadab, et al.
Published: (2025)
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
by: Ray, Sourjyadip, et al.
Published: (2025)
by: Ray, Sourjyadip, et al.
Published: (2025)
A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science
by: Cohn, Clayton, et al.
Published: (2024)
by: Cohn, Clayton, et al.
Published: (2024)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
by: Mukherjee, Sagnik, et al.
Published: (2024)
by: Mukherjee, Sagnik, et al.
Published: (2024)
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Evaluating the Evaluation of Diversity in Commonsense Generation
by: Zhang, Tianhui, et al.
Published: (2025)
by: Zhang, Tianhui, et al.
Published: (2025)
Evaluating Diversity in Automatic Poetry Generation
by: Chen, Yanran, et al.
Published: (2024)
by: Chen, Yanran, et al.
Published: (2024)
Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?
by: Knipper, R. Alexander, et al.
Published: (2025)
by: Knipper, R. Alexander, et al.
Published: (2025)
Evaluating LLM-Generated Q&A Test: a Student-Centered Study
by: Wróblewska, Anna, et al.
Published: (2025)
by: Wróblewska, Anna, et al.
Published: (2025)
Interpretability Framework for LLMs in Undergraduate Calculus
by: Dakshit, Sagnik, et al.
Published: (2025)
by: Dakshit, Sagnik, et al.
Published: (2025)
Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets
by: Gupta, Vatsal, et al.
Published: (2023)
by: Gupta, Vatsal, et al.
Published: (2023)
The Responsible Development of Automated Student Feedback with Generative AI
by: Lindsay, Euan D, et al.
Published: (2023)
by: Lindsay, Euan D, et al.
Published: (2023)
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms
by: Aisyah, Nurul, et al.
Published: (2025)
by: Aisyah, Nurul, et al.
Published: (2025)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
Evaluating the Diversity and Quality of LLM Generated Content
by: Shypula, Alexander, et al.
Published: (2025)
by: Shypula, Alexander, et al.
Published: (2025)
CHORUS: Zero-shot Hierarchical Retrieval and Orchestration for Generating Linear Programming Code
by: Ahmed, Tasnim, et al.
Published: (2025)
by: Ahmed, Tasnim, et al.
Published: (2025)
Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs
by: Chatterjee, Sagnik, et al.
Published: (2026)
by: Chatterjee, Sagnik, et al.
Published: (2026)
Can Public LLMs be used for Self-Diagnosis of Medical Conditions ?
by: Balasubramanian, Nikil Sharan Prabahar, et al.
Published: (2024)
by: Balasubramanian, Nikil Sharan Prabahar, et al.
Published: (2024)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
by: Liu, Yongkang, et al.
Published: (2023)
by: Liu, Yongkang, et al.
Published: (2023)
To Generate or Discriminate? Methodological Considerations for Measuring Cultural Alignment in LLMs
by: Pandey, Saurabh Kumar, et al.
Published: (2026)
by: Pandey, Saurabh Kumar, et al.
Published: (2026)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
by: Schaeffer, Rylan, et al.
Published: (2026)
by: Schaeffer, Rylan, et al.
Published: (2026)
DiaCDM: Cognitive Diagnosis in Teacher-Student Dialogues using the Initiation-Response-Evaluation Framework
by: Jia, Rui, et al.
Published: (2025)
by: Jia, Rui, et al.
Published: (2025)
Asking a Language Model for Diverse Responses
by: Troshin, Sergey, et al.
Published: (2025)
by: Troshin, Sergey, et al.
Published: (2025)
From NLG Evaluation to Modern Student Assessment in the Era of ChatGPT: The Great Misalignment Problem and Pedagogical Multi-Factor Assessment (P-MFA)
by: Hämäläinen, Mika, et al.
Published: (2025)
by: Hämäläinen, Mika, et al.
Published: (2025)
D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Model
by: Byun, Grace, et al.
Published: (2025)
by: Byun, Grace, et al.
Published: (2025)
Reproducing the Metric-Based Evaluation of a Set of Controllable Text Generation Techniques
by: Lorandi, Michela, et al.
Published: (2024)
by: Lorandi, Michela, et al.
Published: (2024)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning
by: Bai, Yi, et al.
Published: (2026)
by: Bai, Yi, et al.
Published: (2026)
From Hallucinations to Facts: Enhancing Language Models with Curated Knowledge Graphs
by: Joshi, Ratnesh Kumar, et al.
Published: (2024)
by: Joshi, Ratnesh Kumar, et al.
Published: (2024)
Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses
by: Peng, Kerui, et al.
Published: (2026)
by: Peng, Kerui, et al.
Published: (2026)
ReasoningFlow: Semantic Structure of Complex Reasoning Traces
by: Lee, Jinu, et al.
Published: (2025)
by: Lee, Jinu, et al.
Published: (2025)
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
by: Hao, Jiangang
Published: (2026)
by: Hao, Jiangang
Published: (2026)
Similar Items
-
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
by: Sun, Jingyi, et al.
Published: (2025) -
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
by: Jahan, Rafid Ishrak, et al.
Published: (2026) -
FutureGen: A RAG-based Approach to Generate the Future Work of Scientific Article
by: Azher, Ibrahim Al, et al.
Published: (2025) -
ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
by: Anik, Anirban Saha, et al.
Published: (2025) -
Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
by: Nemkova, Apollinaire Poli, et al.
Published: (2025)