The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
Fuente:
arXiv
Saved in:
| Main Authors: | Orlikowski, Matthias, Röttger, Paul, Cimiano, Philipp, Hovy, Dirk |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
by: Orlikowski, Matthias, et al.
Published: (2025)
by: Orlikowski, Matthias, et al.
Published: (2025)
Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It's Best to Relate Perspectives!
by: Heinisch, Philipp, et al.
Published: (2023)
by: Heinisch, Philipp, et al.
Published: (2023)
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
by: Fleisig, Eve, et al.
Published: (2025)
by: Fleisig, Eve, et al.
Published: (2025)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
by: Pernisi, Fabio, et al.
Published: (2024)
by: Pernisi, Fabio, et al.
Published: (2024)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
by: Russo, Giuseppe, et al.
Published: (2025)
by: Russo, Giuseppe, et al.
Published: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
by: Rooein, Donya, et al.
Published: (2024)
by: Rooein, Donya, et al.
Published: (2024)
Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement
by: Abercrombie, Gavin, et al.
Published: (2023)
by: Abercrombie, Gavin, et al.
Published: (2023)
Fine-grained Fallacy Detection with Human Label Variation
by: Ramponi, Alan, et al.
Published: (2025)
by: Ramponi, Alan, et al.
Published: (2025)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2023)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2023)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
Diffusion Language Models Are Natively Length-Aware
by: Rossi, Vittorio, et al.
Published: (2026)
by: Rossi, Vittorio, et al.
Published: (2026)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
by: Baumann, Joachim, et al.
Published: (2025)
by: Baumann, Joachim, et al.
Published: (2025)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2025)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2025)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
by: Röttger, Paul, et al.
Published: (2025)
by: Röttger, Paul, et al.
Published: (2025)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
Addressing the Ecological Fallacy in Larger LMs with Human Context
by: Soni, Nikita, et al.
Published: (2026)
by: Soni, Nikita, et al.
Published: (2026)
Do Large Language Models Adapt to Language Variation across Socioeconomic Status?
by: Bassignana, Elisa, et al.
Published: (2026)
by: Bassignana, Elisa, et al.
Published: (2026)
Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset
by: Goldzycher, Janis, et al.
Published: (2024)
by: Goldzycher, Janis, et al.
Published: (2024)
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
by: Rooein, Donya, et al.
Published: (2024)
by: Rooein, Donya, et al.
Published: (2024)
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns
by: Sinelnik, Antonina, et al.
Published: (2024)
by: Sinelnik, Antonina, et al.
Published: (2024)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
VariErr NLI: Separating Annotation Error from Human Label Variation
by: Weber-Genzel, Leon, et al.
Published: (2024)
by: Weber-Genzel, Leon, et al.
Published: (2024)
Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations
by: Peng, Siyao, et al.
Published: (2024)
by: Peng, Siyao, et al.
Published: (2024)
Pointing out the Shortcomings of Relation Extraction Models with Semantically Motivated Adversarials
by: Nolano, Gennaro, et al.
Published: (2024)
by: Nolano, Gennaro, et al.
Published: (2024)
A Factorized Probabilistic Model of the Semantics of Vague Temporal Adverbials Relative to Different Event Types
by: Kenneweg, Svenja, et al.
Published: (2025)
by: Kenneweg, Svenja, et al.
Published: (2025)
Comparing Pre-trained Human Language Models: Is it Better with Human Context as Groups, Individual Traits, or Both?
by: Soni, Nikita, et al.
Published: (2024)
by: Soni, Nikita, et al.
Published: (2024)
Flee the Flaw: Annotating the Underlying Logic of Fallacious Arguments Through Templates and Slot-filling
by: Robbani, Irfan, et al.
Published: (2024)
by: Robbani, Irfan, et al.
Published: (2024)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
by: Gonzalez-Gutierrez, Cesar, et al.
Published: (2025)
by: Gonzalez-Gutierrez, Cesar, et al.
Published: (2025)
Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps
by: Attanasio, Giuseppe, et al.
Published: (2024)
by: Attanasio, Giuseppe, et al.
Published: (2024)
Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
by: Richardson, Andrew Keenan, et al.
Published: (2025)
by: Richardson, Andrew Keenan, et al.
Published: (2025)
TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References
by: Kenneweg, Svenja, et al.
Published: (2025)
by: Kenneweg, Svenja, et al.
Published: (2025)
CompoST: A Benchmark for Analyzing the Ability of LLMs To Compositionally Interpret Questions in a QALD Setting
by: Schmidt, David Maria, et al.
Published: (2025)
by: Schmidt, David Maria, et al.
Published: (2025)
The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
by: Bassignana, Elisa, et al.
Published: (2025)
by: Bassignana, Elisa, et al.
Published: (2025)
On the Interplay between Human Label Variation and Model Fairness
by: Kurniawan, Kemal, et al.
Published: (2025)
by: Kurniawan, Kemal, et al.
Published: (2025)
DADIT: A Dataset for Demographic Classification of Italian Twitter Users and a Comparison of Prediction Methods
by: Lupo, Lorenzo, et al.
Published: (2024)
by: Lupo, Lorenzo, et al.
Published: (2024)
Similar Items
-
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
by: Orlikowski, Matthias, et al.
Published: (2025) -
Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It's Best to Relate Perspectives!
by: Heinisch, Philipp, et al.
Published: (2023) -
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
by: Fleisig, Eve, et al.
Published: (2025) -
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
by: Pernisi, Fabio, et al.
Published: (2024) -
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
by: Russo, Giuseppe, et al.
Published: (2025)