Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Kern, Christoph, Eckman, Stephanie, Beck, Jacob, Chew, Rob, Ma, Bolei, Kreuter, Frauke |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aligning NLP Models with Target Population Perspectives using PAIR: Population-Aligned Instance Replication
by: Eckman, Stephanie, et al.
Published: (2025)
by: Eckman, Stephanie, et al.
Published: (2025)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
by: Chew, Robert, et al.
Published: (2026)
by: Chew, Robert, et al.
Published: (2026)
Bias in the Loop: How Humans Evaluate AI-Generated Suggestions
by: Beck, Jacob, et al.
Published: (2025)
by: Beck, Jacob, et al.
Published: (2025)
Position: Insights from Survey Methodology can Improve Training Data
by: Eckman, Stephanie, et al.
Published: (2024)
by: Eckman, Stephanie, et al.
Published: (2024)
The Missing Link: Allocation Performance in Causal Machine Learning
by: Fischer-Abaigar, Unai, et al.
Published: (2024)
by: Fischer-Abaigar, Unai, et al.
Published: (2024)
Bridging the gap: Towards an Expanded Toolkit for AI-driven Decision-Making in the Public Sector
by: Fischer-Abaigar, Unai, et al.
Published: (2023)
by: Fischer-Abaigar, Unai, et al.
Published: (2023)
An Agglomerative Clustering of Simulation Output Distributions Using Regularized Wasserstein Distance
by: Ghasemloo, Mohammadmahdi, et al.
Published: (2024)
by: Ghasemloo, Mohammadmahdi, et al.
Published: (2024)
Quantifying and Attributing Submodel Uncertainty in Stochastic Simulation Models and Digital Twins
by: Ghasemloo, Mohammadmahdi, et al.
Published: (2026)
by: Ghasemloo, Mohammadmahdi, et al.
Published: (2026)
Proximal Causal Inference With Text Data
by: Chen, Jacob M., et al.
Published: (2024)
by: Chen, Jacob M., et al.
Published: (2024)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
by: Ball, Sarah, et al.
Published: (2024)
by: Ball, Sarah, et al.
Published: (2024)
Connecting Algorithmic Fairness to Quality Dimensions in Machine Learning in Official Statistics and Survey Production
by: Schenk, Patrick Oliver, et al.
Published: (2024)
by: Schenk, Patrick Oliver, et al.
Published: (2024)
Multi-CATE: Multi-Accurate Conditional Average Treatment Effect Estimation Robust to Unknown Covariate Shifts
by: Kern, Christoph, et al.
Published: (2024)
by: Kern, Christoph, et al.
Published: (2024)
Bias Begins with Data: The FairGround Corpus for Robust and Reproducible Research on Algorithmic Fairness
by: Simson, Jan, et al.
Published: (2025)
by: Simson, Jan, et al.
Published: (2025)
TWIN-GPT: Digital Twins for Clinical Trials via Large Language Model
by: Wang, Yue, et al.
Published: (2024)
by: Wang, Yue, et al.
Published: (2024)
End-To-End Causal Effect Estimation from Unstructured Natural Language Data
by: Dhawan, Nikita, et al.
Published: (2024)
by: Dhawan, Nikita, et al.
Published: (2024)
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
by: Nie, Ercong, et al.
Published: (2024)
by: Nie, Ercong, et al.
Published: (2024)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
by: Yuan, Chenchen, et al.
Published: (2026)
by: Yuan, Chenchen, et al.
Published: (2026)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
by: Boyeau, Pierre, et al.
Published: (2024)
by: Boyeau, Pierre, et al.
Published: (2024)
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
by: Jung, Julie, et al.
Published: (2025)
by: Jung, Julie, et al.
Published: (2025)
Omitted Variable Bias in Language Models Under Distribution Shift
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Optimizing Language Models for Human Preferences is a Causal Inference Problem
by: Lin, Victoria, et al.
Published: (2024)
by: Lin, Victoria, et al.
Published: (2024)
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
by: Tu, Ruibo, et al.
Published: (2024)
by: Tu, Ruibo, et al.
Published: (2024)
Causal Graph Discovery with Retrieval-Augmented Generation based Large Language Models
by: Zhang, Yuzhe, et al.
Published: (2024)
by: Zhang, Yuzhe, et al.
Published: (2024)
Transforming Sensitive Documents into Quantitative Data: An AI-Based Preprocessing Toolchain for Structured and Privacy-Conscious Analysis
by: Ledberg, Anders, et al.
Published: (2025)
by: Ledberg, Anders, et al.
Published: (2025)
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
by: Ma, Bolei, et al.
Published: (2025)
by: Ma, Bolei, et al.
Published: (2025)
Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models
by: Egami, Naoki, et al.
Published: (2023)
by: Egami, Naoki, et al.
Published: (2023)
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks
by: Ma, Bolei, et al.
Published: (2024)
by: Ma, Bolei, et al.
Published: (2024)
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models
by: Ma, Bolei, et al.
Published: (2024)
by: Ma, Bolei, et al.
Published: (2024)
Kernel Treatment Effects with Adaptively Collected Data
by: Zenati, Houssam, et al.
Published: (2025)
by: Zenati, Houssam, et al.
Published: (2025)
Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
by: Dobariya, Om, et al.
Published: (2025)
by: Dobariya, Om, et al.
Published: (2025)
A Design-based Solution for Causal Inference with Text: Can a Language Model Be Too Large?
by: Tierney, Graham, et al.
Published: (2025)
by: Tierney, Graham, et al.
Published: (2025)
A Sensitivity Approach to Causal Inference Under Limited Overlap
by: Ma, Yuanzhe, et al.
Published: (2025)
by: Ma, Yuanzhe, et al.
Published: (2025)
MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment
by: Wang, Tianze, et al.
Published: (2025)
by: Wang, Tianze, et al.
Published: (2025)
Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality Missingness
by: Liang, Zihan, et al.
Published: (2025)
by: Liang, Zihan, et al.
Published: (2025)
MuPlon: Multi-Path Causal Optimization for Claim Verification through Controlling Confounding
by: Guo, Hanghui, et al.
Published: (2025)
by: Guo, Hanghui, et al.
Published: (2025)
How to Evaluate Entity Resolution Systems: An Entity-Centric Framework with Application to Inventor Name Disambiguation
by: Binette, Olivier, et al.
Published: (2024)
by: Binette, Olivier, et al.
Published: (2024)
LLMs as Implicit Imputers: Uncertainty Should Scale with Missing Information
by: van Buuren, Stef
Published: (2026)
by: van Buuren, Stef
Published: (2026)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Spiking the training data to correct for test set contamination
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026)
LIDS: LLM Summary Inference Under the Layered Lens
by: Park, Dylan, et al.
Published: (2026)
by: Park, Dylan, et al.
Published: (2026)
Similar Items
-
Aligning NLP Models with Target Population Perspectives using PAIR: Population-Aligned Instance Replication
by: Eckman, Stephanie, et al.
Published: (2025) -
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
by: Chew, Robert, et al.
Published: (2026) -
Bias in the Loop: How Humans Evaluate AI-Generated Suggestions
by: Beck, Jacob, et al.
Published: (2025) -
Position: Insights from Survey Methodology can Improve Training Data
by: Eckman, Stephanie, et al.
Published: (2024) -
The Missing Link: Allocation Performance in Causal Machine Learning
by: Fischer-Abaigar, Unai, et al.
Published: (2024)