Comparison of Scoring Rationales Between Large Language Models and Human Raters
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hua, Haowei, Jiao, Hong, Song, Dan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Empirical Comparison of Encoder-Based Language Models and Feature-Based Supervised Machine Learning Approaches to Automated Scoring of Long Essays
par: Wang, Kuo, et autres
Publié: (2026)
par: Wang, Kuo, et autres
Publié: (2026)
Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model
par: Jiao, Hong, et autres
Publié: (2025)
par: Jiao, Hong, et autres
Publié: (2025)
Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays
par: Hua, Haowei, et autres
Publié: (2025)
par: Hua, Haowei, et autres
Publié: (2025)
Exploring the Utilities of the Rationales from Large Language Models to Enhance Automated Essay Scoring
par: Jiao, Hong, et autres
Publié: (2025)
par: Jiao, Hong, et autres
Publié: (2025)
Teaching Large Language Models Number-Focused Headline Generation With Key Element Rationales
par: Qian, Zhen, et autres
Publié: (2025)
par: Qian, Zhen, et autres
Publié: (2025)
Exploring the Trade-off Between Model Performance and Explanation Plausibility of Text Classifiers Using Human Rationales
par: Resck, Lucas E., et autres
Publié: (2024)
par: Resck, Lucas E., et autres
Publié: (2024)
How Ambiguous Are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty
par: Kim, Hazel H.
Publié: (2024)
par: Kim, Hazel H.
Publié: (2024)
Large Language Model Can Be a Foundation for Hidden Rationale-Based Retrieval
par: Ji, Luo, et autres
Publié: (2024)
par: Ji, Luo, et autres
Publié: (2024)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
par: Nirmal, Ayushi, et autres
Publié: (2024)
par: Nirmal, Ayushi, et autres
Publié: (2024)
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
par: Eilertsen, Brage, et autres
Publié: (2025)
par: Eilertsen, Brage, et autres
Publié: (2025)
Investigating Automatic Scoring and Feedback using Large Language Models
par: Katuka, Gloria Ashiya, et autres
Publié: (2024)
par: Katuka, Gloria Ashiya, et autres
Publié: (2024)
Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?
par: Zhou, Zhanke, et autres
Publié: (2024)
par: Zhou, Zhanke, et autres
Publié: (2024)
Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
par: Kim, Bo-Kyeong, et autres
Publié: (2024)
par: Kim, Bo-Kyeong, et autres
Publié: (2024)
Shuttle Between the Instructions and the Parameters of Large Language Models
par: Sun, Wangtao, et autres
Publié: (2025)
par: Sun, Wangtao, et autres
Publié: (2025)
Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
par: Xu, Qingshu, et autres
Publié: (2025)
par: Xu, Qingshu, et autres
Publié: (2025)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
par: Hua, Peichun, et autres
Publié: (2025)
par: Hua, Peichun, et autres
Publié: (2025)
QA-Calibration of Language Model Confidence Scores
par: Manggala, Putra, et autres
Publié: (2024)
par: Manggala, Putra, et autres
Publié: (2024)
A Survey of On-Policy Distillation for Large Language Models
par: Song, Mingyang, et autres
Publié: (2026)
par: Song, Mingyang, et autres
Publié: (2026)
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
par: Eisape, Tiwalayo, et autres
Publié: (2023)
par: Eisape, Tiwalayo, et autres
Publié: (2023)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
par: Tang, Shengkun, et autres
Publié: (2025)
par: Tang, Shengkun, et autres
Publié: (2025)
Beyond the Prompt in Large Language Models: Comprehension, In-Context Learning, and Chain-of-Thought
par: Jiao, Yuling, et autres
Publié: (2026)
par: Jiao, Yuling, et autres
Publié: (2026)
Extreme Compression of Large Language Models via Additive Quantization
par: Egiazarian, Vage, et autres
Publié: (2024)
par: Egiazarian, Vage, et autres
Publié: (2024)
Generative Evaluation of Complex Reasoning in Large Language Models
par: Lin, Haowei, et autres
Publié: (2025)
par: Lin, Haowei, et autres
Publié: (2025)
Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages
par: Zhao, Yue, et autres
Publié: (2026)
par: Zhao, Yue, et autres
Publié: (2026)
Hansel: Output Length Controlling Framework for Large Language Models
par: Song, Seoha, et autres
Publié: (2024)
par: Song, Seoha, et autres
Publié: (2024)
ActTail: Global Activation Sparsity in Large Language Models
par: Hou, Wenwen, et autres
Publié: (2026)
par: Hou, Wenwen, et autres
Publié: (2026)
Self-Comparison for Dataset-Level Membership Inference in Large (Vision-)Language Models
par: Ren, Jie, et autres
Publié: (2024)
par: Ren, Jie, et autres
Publié: (2024)
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
par: Phan, Buu, et autres
Publié: (2025)
par: Phan, Buu, et autres
Publié: (2025)
CoLa: Learning to Interactively Collaborate with Large Language Models
par: Sharma, Abhishek, et autres
Publié: (2025)
par: Sharma, Abhishek, et autres
Publié: (2025)
Model Hemorrhage and the Robustness Limits of Large Language Models
par: Ma, Ziyang, et autres
Publié: (2025)
par: Ma, Ziyang, et autres
Publié: (2025)
A Survey on Symbolic Knowledge Distillation of Large Language Models
par: Acharya, Kamal, et autres
Publié: (2024)
par: Acharya, Kamal, et autres
Publié: (2024)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
par: Hengle, Amey, et autres
Publié: (2024)
par: Hengle, Amey, et autres
Publié: (2024)
Dissecting Fine-Tuning Unlearning in Large Language Models
par: Hong, Yihuai, et autres
Publié: (2024)
par: Hong, Yihuai, et autres
Publié: (2024)
DetoxBench: Benchmarking Large Language Models for Multitask Fraud & Abuse Detection
par: Chakraborty, Joymallya, et autres
Publié: (2024)
par: Chakraborty, Joymallya, et autres
Publié: (2024)
Model Interpretability and Rationale Extraction by Input Mask Optimization
par: Brinner, Marc, et autres
Publié: (2025)
par: Brinner, Marc, et autres
Publié: (2025)
Surveying Attitudinal Alignment Between Large Language Models Vs. Humans Towards 17 Sustainable Development Goals
par: Wu, Qingyang, et autres
Publié: (2024)
par: Wu, Qingyang, et autres
Publié: (2024)
SORSA: Singular Values and Orthonormal Regularized Singular Vectors Adaptation of Large Language Models
par: Cao, Yang, et autres
Publié: (2024)
par: Cao, Yang, et autres
Publié: (2024)
Adaptive Task Vectors for Large Language Models
par: Kang, Joonseong, et autres
Publié: (2025)
par: Kang, Joonseong, et autres
Publié: (2025)
LLMs Plagiarize: Ensuring Responsible Sourcing of Large Language Model Training Data Through Knowledge Graph Comparison
par: Mondal, Devam, et autres
Publié: (2024)
par: Mondal, Devam, et autres
Publié: (2024)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
par: Hong, Yinrong, et autres
Publié: (2025)
par: Hong, Yinrong, et autres
Publié: (2025)
Documents similaires
-
Empirical Comparison of Encoder-Based Language Models and Feature-Based Supervised Machine Learning Approaches to Automated Scoring of Long Essays
par: Wang, Kuo, et autres
Publié: (2026) -
Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model
par: Jiao, Hong, et autres
Publié: (2025) -
Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays
par: Hua, Haowei, et autres
Publié: (2025) -
Exploring the Utilities of the Rationales from Large Language Models to Enhance Automated Essay Scoring
par: Jiao, Hong, et autres
Publié: (2025) -
Teaching Large Language Models Number-Focused Headline Generation With Key Element Rationales
par: Qian, Zhen, et autres
Publié: (2025)