How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jung, Julie, Lu, Max, Benker, Sina Chole, Darici, Dogus |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Maximizing Signal in Human-Model Preference Alignment
von: Kraus, Kelsey, et al.
Veröffentlicht: (2025)
von: Kraus, Kelsey, et al.
Veröffentlicht: (2025)
Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
von: Dobariya, Om, et al.
Veröffentlicht: (2025)
von: Dobariya, Om, et al.
Veröffentlicht: (2025)
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
von: Kern, Christoph, et al.
Veröffentlicht: (2023)
von: Kern, Christoph, et al.
Veröffentlicht: (2023)
Segmenting Human-LLM Co-authored Text via Change Point Detection
von: Li, Mengchu, et al.
Veröffentlicht: (2026)
von: Li, Mengchu, et al.
Veröffentlicht: (2026)
A Finite-Calibration Regime Map for LLM Judge Panels
von: Zhu, Bin, et al.
Veröffentlicht: (2026)
von: Zhu, Bin, et al.
Veröffentlicht: (2026)
Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
von: Li, Yanran
Veröffentlicht: (2026)
von: Li, Yanran
Veröffentlicht: (2026)
Human-AI Co-design for Clinical Prediction Models
von: Feng, Jean, et al.
Veröffentlicht: (2026)
von: Feng, Jean, et al.
Veröffentlicht: (2026)
Optimizing Language Models for Human Preferences is a Causal Inference Problem
von: Lin, Victoria, et al.
Veröffentlicht: (2024)
von: Lin, Victoria, et al.
Veröffentlicht: (2024)
MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
Syntax-Guided Diffusion Language Models with User-Integrated Personalization
von: Zhang, Ruqian, et al.
Veröffentlicht: (2025)
von: Zhang, Ruqian, et al.
Veröffentlicht: (2025)
A New Semisupervised Technique for Polarity Analysis using Masked Language Models
von: Watanabe, Kohei
Veröffentlicht: (2026)
von: Watanabe, Kohei
Veröffentlicht: (2026)
Distributed Asymmetric Allocation: A Topic Model for Large Imbalanced Corpora in Social Sciences
von: Watanabe, Kohei
Veröffentlicht: (2025)
von: Watanabe, Kohei
Veröffentlicht: (2025)
An Embedded Diachronic Sense Change Model with a Case Study from Ancient Greek
von: Zafar, Schyan, et al.
Veröffentlicht: (2023)
von: Zafar, Schyan, et al.
Veröffentlicht: (2023)
Aligning NLP Models with Target Population Perspectives using PAIR: Population-Aligned Instance Replication
von: Eckman, Stephanie, et al.
Veröffentlicht: (2025)
von: Eckman, Stephanie, et al.
Veröffentlicht: (2025)
LIDS: LLM Summary Inference Under the Layered Lens
von: Park, Dylan, et al.
Veröffentlicht: (2026)
von: Park, Dylan, et al.
Veröffentlicht: (2026)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
Optimal Estimation of Watermark Proportions in Hybrid AI-Human Texts
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
How to Evaluate Entity Resolution Systems: An Entity-Centric Framework with Application to Inventor Name Disambiguation
von: Binette, Olivier, et al.
Veröffentlicht: (2024)
von: Binette, Olivier, et al.
Veröffentlicht: (2024)
TWIN-GPT: Digital Twins for Clinical Trials via Large Language Model
von: Wang, Yue, et al.
Veröffentlicht: (2024)
von: Wang, Yue, et al.
Veröffentlicht: (2024)
Robust Detection of Watermarks for Large Language Models Under Human Edits
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
von: Hobelsberger, Christian, et al.
Veröffentlicht: (2025)
von: Hobelsberger, Christian, et al.
Veröffentlicht: (2025)
Improving Probabilistic Models in Text Classification via Active Learning
von: Bosley, Mitchell, et al.
Veröffentlicht: (2022)
von: Bosley, Mitchell, et al.
Veröffentlicht: (2022)
A chart review process aided by natural language processing and multi-wave adaptive sampling to expedite validation of code-based algorithms for large database studies
von: Wang, Shirley V, et al.
Veröffentlicht: (2025)
von: Wang, Shirley V, et al.
Veröffentlicht: (2025)
Geological Inference from Textual Data using Word Embeddings
von: Linphrachaya, Nanmanas, et al.
Veröffentlicht: (2025)
von: Linphrachaya, Nanmanas, et al.
Veröffentlicht: (2025)
Transforming Sensitive Documents into Quantitative Data: An AI-Based Preprocessing Toolchain for Structured and Privacy-Conscious Analysis
von: Ledberg, Anders, et al.
Veröffentlicht: (2025)
von: Ledberg, Anders, et al.
Veröffentlicht: (2025)
Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models
von: Lyu, Y., et al.
Veröffentlicht: (2025)
von: Lyu, Y., et al.
Veröffentlicht: (2025)
A Position Paper on the Automatic Generation of Machine Learning Leaderboards
von: Timmer, Roelien C, et al.
Veröffentlicht: (2025)
von: Timmer, Roelien C, et al.
Veröffentlicht: (2025)
Differential contributions of machine learning and statistical analysis to language and cognitive sciences
von: Sun, Kun, et al.
Veröffentlicht: (2024)
von: Sun, Kun, et al.
Veröffentlicht: (2024)
Exploring Intra and Inter-language Consistency in Embeddings with ICA
von: Li, Rongzhi, et al.
Veröffentlicht: (2024)
von: Li, Rongzhi, et al.
Veröffentlicht: (2024)
Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study
von: Okada, Kensuke, et al.
Veröffentlicht: (2026)
von: Okada, Kensuke, et al.
Veröffentlicht: (2026)
Isolated Causal Effects of Natural Language
von: Lin, Victoria, et al.
Veröffentlicht: (2024)
von: Lin, Victoria, et al.
Veröffentlicht: (2024)
Omitted Variable Bias in Language Models Under Distribution Shift
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
von: Lin, Victoria, et al.
Veröffentlicht: (2026)
ALCM: Autonomous LLM-Augmented Causal Discovery Framework
von: Khatibi, Elahe, et al.
Veröffentlicht: (2024)
von: Khatibi, Elahe, et al.
Veröffentlicht: (2024)
Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks
von: Hua, Wenyue, et al.
Veröffentlicht: (2024)
von: Hua, Wenyue, et al.
Veröffentlicht: (2024)
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
von: Tu, Ruibo, et al.
Veröffentlicht: (2024)
von: Tu, Ruibo, et al.
Veröffentlicht: (2024)
Using Large Language Models to Suggest Informative Prior Distributions in Bayesian Statistics
von: Riegler, Michael A., et al.
Veröffentlicht: (2025)
von: Riegler, Michael A., et al.
Veröffentlicht: (2025)
Causal Graph Discovery with Retrieval-Augmented Generation based Large Language Models
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2024)
CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring
von: Cohn, Clayton, et al.
Veröffentlicht: (2025)
von: Cohn, Clayton, et al.
Veröffentlicht: (2025)
CLEAR: Can Language Models Really Understand Causal Graphs?
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
von: Chew, Robert, et al.
Veröffentlicht: (2026)
von: Chew, Robert, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Maximizing Signal in Human-Model Preference Alignment
von: Kraus, Kelsey, et al.
Veröffentlicht: (2025) -
Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
von: Dobariya, Om, et al.
Veröffentlicht: (2025) -
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
von: Kern, Christoph, et al.
Veröffentlicht: (2023) -
Segmenting Human-LLM Co-authored Text via Change Point Detection
von: Li, Mengchu, et al.
Veröffentlicht: (2026) -
A Finite-Calibration Regime Map for LLM Judge Panels
von: Zhu, Bin, et al.
Veröffentlicht: (2026)