Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model
Fuente:
arXiv
Saved in:
| Main Authors: | Jiao, Hong, Song, Dan, Lee, Won-Chan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comparison of Scoring Rationales Between Large Language Models and Human Raters
by: Hua, Haowei, et al.
Published: (2025)
by: Hua, Haowei, et al.
Published: (2025)
Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory
by: Song, Dan, et al.
Published: (2025)
by: Song, Dan, et al.
Published: (2025)
Out of One, Many: Using Language Models to Simulate Human Samples
by: Argyle, Lisa P., et al.
Published: (2022)
by: Argyle, Lisa P., et al.
Published: (2022)
How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation
by: Lee, Wilson Y.
Published: (2026)
by: Lee, Wilson Y.
Published: (2026)
Empirical Comparison of Encoder-Based Language Models and Feature-Based Supervised Machine Learning Approaches to Automated Scoring of Long Essays
by: Wang, Kuo, et al.
Published: (2026)
by: Wang, Kuo, et al.
Published: (2026)
Task Facet Learning: A Structured Approach to Prompt Optimization
by: Juneja, Gurusha, et al.
Published: (2024)
by: Juneja, Gurusha, et al.
Published: (2024)
Evaluating Rater Effects of Large Language Models in Automated Essay Scoring: GPT, Claude, Gemini, and DeepSeek
by: Hong Jiao, et al.
Published: (2026)
by: Hong Jiao, et al.
Published: (2026)
Bridging Human and Model Perspectives: A Comparative Analysis of Political Bias Detection in News Media Using Large Language Models
by: Banik, Shreya Adrita, et al.
Published: (2025)
by: Banik, Shreya Adrita, et al.
Published: (2025)
Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays
by: Hua, Haowei, et al.
Published: (2025)
by: Hua, Haowei, et al.
Published: (2025)
Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
by: Xu, Qingshu, et al.
Published: (2025)
by: Xu, Qingshu, et al.
Published: (2025)
Multi-Faceted Question Complexity Estimation Targeting Topic Domain-Specificity
by: R, Sujay, et al.
Published: (2024)
by: R, Sujay, et al.
Published: (2024)
Flextron: Many-in-One Flexible Large Language Model
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
Adversarial Attacks on AI-Generated Text Detection Models: A Token Probability-Based Approach Using Embeddings
by: Kadhim, Ahmed K., et al.
Published: (2025)
by: Kadhim, Ahmed K., et al.
Published: (2025)
Traditional Readability Formulas Compared for English
by: Lee, Bruce W., et al.
Published: (2023)
by: Lee, Bruce W., et al.
Published: (2023)
Many-Shot In-Context Learning
by: Agarwal, Rishabh, et al.
Published: (2024)
by: Agarwal, Rishabh, et al.
Published: (2024)
Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence
by: Resnick, Paul, et al.
Published: (2021)
by: Resnick, Paul, et al.
Published: (2021)
The Few Govern the Many:Unveiling Few-Layer Dominance for Time Series Models
by: Qiu, Xin, et al.
Published: (2025)
by: Qiu, Xin, et al.
Published: (2025)
Many Minds from One Model: Bayesian-Inspired Transformers for Population Diversity
by: Yang, Diji, et al.
Published: (2025)
by: Yang, Diji, et al.
Published: (2025)
Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service Dialogues
by: Hong, Mengze, et al.
Published: (2024)
by: Hong, Mengze, et al.
Published: (2024)
From Small to Large Language Models: Revisiting the Federalist Papers
by: Jeong, So Won, et al.
Published: (2025)
by: Jeong, So Won, et al.
Published: (2025)
BIG5-TPoT: Predicting BIG Five Personality Traits, Facets, and Items Through Targeted Preselection of Texts
by: Le, Triet M., et al.
Published: (2025)
by: Le, Triet M., et al.
Published: (2025)
A Multi-Faceted Evaluation Framework for Assessing Synthetic Data Generated by Large Language Models
by: Yuan, Yefeng, et al.
Published: (2024)
by: Yuan, Yefeng, et al.
Published: (2024)
On the Calibration of Multilingual Question Answering LLMs
by: Yang, Yahan, et al.
Published: (2023)
by: Yang, Yahan, et al.
Published: (2023)
Base Models Look Human To AI Detectors
by: Xu, Yixuan Even, et al.
Published: (2026)
by: Xu, Yixuan Even, et al.
Published: (2026)
Hansel: Output Length Controlling Framework for Large Language Models
by: Song, Seoha, et al.
Published: (2024)
by: Song, Seoha, et al.
Published: (2024)
Comparing Pre-trained Human Language Models: Is it Better with Human Context as Groups, Individual Traits, or Both?
by: Soni, Nikita, et al.
Published: (2024)
by: Soni, Nikita, et al.
Published: (2024)
Comparing Neighbors Together Makes it Easy: Jointly Comparing Multiple Candidates for Efficient and Effective Retrieval
by: Song, Jonghyun, et al.
Published: (2024)
by: Song, Jonghyun, et al.
Published: (2024)
AI-Enhanced Cognitive Behavioral Therapy: Deep Learning and Large Language Models for Extracting Cognitive Pathways from Social Media Texts
by: Jiang, Meng, et al.
Published: (2024)
by: Jiang, Meng, et al.
Published: (2024)
ULMA: Unified Language Model Alignment with Human Demonstration and Point-wise Preference
by: Cai, Tianchi, et al.
Published: (2023)
by: Cai, Tianchi, et al.
Published: (2023)
Assessing Gender Bias in LLMs: Comparing LLM Outputs with Human Perceptions and Official Statistics
by: Bas, Tetiana
Published: (2024)
by: Bas, Tetiana
Published: (2024)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
by: Kim, Jane Paik
Published: (2026)
by: Kim, Jane Paik
Published: (2026)
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
by: Hong, Yuzhong, et al.
Published: (2024)
by: Hong, Yuzhong, et al.
Published: (2024)
Many-to-English Machine Translation Tools, Data, and Pretrained Models
by: Gowda, Thamme, et al.
Published: (2021)
by: Gowda, Thamme, et al.
Published: (2021)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
by: Lee, Tzu-Yun, et al.
Published: (2025)
by: Lee, Tzu-Yun, et al.
Published: (2025)
Towards Compute-Optimal Many-Shot In-Context Learning
by: Golchin, Shahriar, et al.
Published: (2025)
by: Golchin, Shahriar, et al.
Published: (2025)
Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
by: Wu, Xixi, et al.
Published: (2026)
by: Wu, Xixi, et al.
Published: (2026)
Early Linguistic Pattern of Anxiety from Social Media Using Interpretable Linguistic Features: A Multi-Faceted Validation Study with Author-Disjoint Evaluation
by: Utsa, Arnab Das
Published: (2026)
by: Utsa, Arnab Das
Published: (2026)
Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
by: Tutnov, Rasul, et al.
Published: (2025)
by: Tutnov, Rasul, et al.
Published: (2025)
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
by: Upasani, Shubhangi, et al.
Published: (2026)
by: Upasani, Shubhangi, et al.
Published: (2026)
Efficient Process Reward Modeling via Contrastive Mutual Information
by: Lee, Nakyung, et al.
Published: (2026)
by: Lee, Nakyung, et al.
Published: (2026)
Similar Items
-
Comparison of Scoring Rationales Between Large Language Models and Human Raters
by: Hua, Haowei, et al.
Published: (2025) -
Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory
by: Song, Dan, et al.
Published: (2025) -
Out of One, Many: Using Language Models to Simulate Human Samples
by: Argyle, Lisa P., et al.
Published: (2022) -
How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation
by: Lee, Wilson Y.
Published: (2026) -
Empirical Comparison of Encoder-Based Language Models and Feature-Based Supervised Machine Learning Approaches to Automated Scoring of Long Essays
by: Wang, Kuo, et al.
Published: (2026)