Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bonthu, Sridevi, Sree, S. Rama, Prasad, M. H. M. Krishna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ASAG2024: A Combined Benchmark for Short Answer Grading
von: Meyer, Gérôme, et al.
Veröffentlicht: (2024)
von: Meyer, Gérôme, et al.
Veröffentlicht: (2024)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
von: Bennion, Jonathan, et al.
Veröffentlicht: (2025)
von: Bennion, Jonathan, et al.
Veröffentlicht: (2025)
Exploring the Performance of ML/DL Architectures on the MNIST-1D Dataset
von: Beebe, Michael, et al.
Veröffentlicht: (2026)
von: Beebe, Michael, et al.
Veröffentlicht: (2026)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
von: Puri, Bruno, et al.
Veröffentlicht: (2025)
von: Puri, Bruno, et al.
Veröffentlicht: (2025)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
von: Jeong, Hyejun, et al.
Veröffentlicht: (2024)
von: Jeong, Hyejun, et al.
Veröffentlicht: (2024)
Mitigate Negative Transfer with Similarity Heuristic Lifelong Prompt Tuning
von: Wu, Chenyuan, et al.
Veröffentlicht: (2024)
von: Wu, Chenyuan, et al.
Veröffentlicht: (2024)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
von: Tian, Yijun, et al.
Veröffentlicht: (2024)
von: Tian, Yijun, et al.
Veröffentlicht: (2024)
Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
von: Muneeb, Muhammad, et al.
Veröffentlicht: (2025)
von: Muneeb, Muhammad, et al.
Veröffentlicht: (2025)
Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation
von: Rahman, A B M Ashikur, et al.
Veröffentlicht: (2024)
von: Rahman, A B M Ashikur, et al.
Veröffentlicht: (2024)
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
von: Nikitin, Alexander, et al.
Veröffentlicht: (2024)
von: Nikitin, Alexander, et al.
Veröffentlicht: (2024)
Efficient data selection employing Semantic Similarity-based Graph Structures for model training
von: Petcu, Roxana, et al.
Veröffentlicht: (2024)
von: Petcu, Roxana, et al.
Veröffentlicht: (2024)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models
von: Jha, Abha, et al.
Veröffentlicht: (2026)
von: Jha, Abha, et al.
Veröffentlicht: (2026)
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
von: Qi, Jirui, et al.
Veröffentlicht: (2024)
von: Qi, Jirui, et al.
Veröffentlicht: (2024)
Linguistic Patterns in Pandemic-Related Content: A Comparative Analysis of COVID-19, Constraint, and Monkeypox Datasets
von: Sikosana, Mkululi, et al.
Veröffentlicht: (2025)
von: Sikosana, Mkululi, et al.
Veröffentlicht: (2025)
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
von: Mackraz, Natalie, et al.
Veröffentlicht: (2024)
von: Mackraz, Natalie, et al.
Veröffentlicht: (2024)
MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers
von: Cho, Nicole, et al.
Veröffentlicht: (2025)
von: Cho, Nicole, et al.
Veröffentlicht: (2025)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
Explicit Diversity Conditions for Effective Question Answer Generation with Large Language Models
von: Yadav, Vikas, et al.
Veröffentlicht: (2024)
von: Yadav, Vikas, et al.
Veröffentlicht: (2024)
Extraction of Research Objectives, Machine Learning Model Names, and Dataset Names from Academic Papers and Analysis of Their Interrelationships Using LLM and Network Analysis
von: Nishio, S., et al.
Veröffentlicht: (2024)
von: Nishio, S., et al.
Veröffentlicht: (2024)
Proving that Cryptic Crossword Clue Answers are Correct
von: Andrews, Martin, et al.
Veröffentlicht: (2024)
von: Andrews, Martin, et al.
Veröffentlicht: (2024)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset
von: Henkel, Owen, et al.
Veröffentlicht: (2023)
von: Henkel, Owen, et al.
Veröffentlicht: (2023)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
von: Seddik, Mohamed El Amine, et al.
Veröffentlicht: (2024)
von: Seddik, Mohamed El Amine, et al.
Veröffentlicht: (2024)
When Does a Language Model Commit? A Finite-Answer Theory of Pre-Verbalization Commitment
von: Zhang, Long, et al.
Veröffentlicht: (2026)
von: Zhang, Long, et al.
Veröffentlicht: (2026)
Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
Measuring and Reducing LLM Hallucination without Gold-Standard Answers
von: Wei, Jiaheng, et al.
Veröffentlicht: (2024)
von: Wei, Jiaheng, et al.
Veröffentlicht: (2024)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
von: Baan, Joris, et al.
Veröffentlicht: (2026)
von: Baan, Joris, et al.
Veröffentlicht: (2026)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
von: Schiekiera, Louis, et al.
Veröffentlicht: (2026)
von: Schiekiera, Louis, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
von: Méloux, Maxime, et al.
Veröffentlicht: (2025)
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
von: Thornton, Scott
Veröffentlicht: (2025)
von: Thornton, Scott
Veröffentlicht: (2025)
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
von: Zhang, Hugh, et al.
Veröffentlicht: (2024)
von: Zhang, Hugh, et al.
Veröffentlicht: (2024)
Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications
von: Balne, Charith Chandra Sai, et al.
Veröffentlicht: (2024)
von: Balne, Charith Chandra Sai, et al.
Veröffentlicht: (2024)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ASAG2024: A Combined Benchmark for Short Answer Grading
von: Meyer, Gérôme, et al.
Veröffentlicht: (2024) -
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
von: Bennion, Jonathan, et al.
Veröffentlicht: (2025) -
Exploring the Performance of ML/DL Architectures on the MNIST-1D Dataset
von: Beebe, Michael, et al.
Veröffentlicht: (2026) -
Atlas-Alignment: Making Interpretability Transferable Across Language Models
von: Puri, Bruno, et al.
Veröffentlicht: (2025) -
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)