Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM
Fuente:
arXiv
Saved in:
| Main Authors: | Shetty, Samay U., Weerasooriya, Tharindu Cyril, Pandita, Deepak, Homan, Christopher M. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LPI-RIT at LeWiDi-2025: Improving Distributional Predictions via Metadata and Loss Reweighting with DisCo
by: Sawkar, Mandira, et al.
Published: (2025)
by: Sawkar, Mandira, et al.
Published: (2025)
ProRefine: Inference-Time Prompt Refinement with Textual Feedback
by: Pandita, Deepak, et al.
Published: (2025)
by: Pandita, Deepak, et al.
Published: (2025)
ARTICLE: Annotator Reliability Through In-Context Learning
by: Dutta, Sujan, et al.
Published: (2024)
by: Dutta, Sujan, et al.
Published: (2024)
Forest vs Tree: The $(N, K)$ Trade-off in Reproducible ML Evaluation
by: Pandita, Deepak, et al.
Published: (2025)
by: Pandita, Deepak, et al.
Published: (2025)
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
by: Pandita, Deepak, et al.
Published: (2026)
by: Pandita, Deepak, et al.
Published: (2026)
Rater Cohesion and Quality from a Vicarious Perspective
by: Pandita, Deepak, et al.
Published: (2024)
by: Pandita, Deepak, et al.
Published: (2024)
Subasa - Adapting Language Models for Low-resourced Offensive Language Detection in Sinhala
by: Haturusinghe, Shanilka, et al.
Published: (2025)
by: Haturusinghe, Shanilka, et al.
Published: (2025)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
HumanLM: Simulating Users with State Alignment Beats Response Imitation
by: Wu, Shirley, et al.
Published: (2026)
by: Wu, Shirley, et al.
Published: (2026)
Iterative Prompt Refinement for Dyslexia-Friendly Text Summarization Using GPT-4o
by: Bhojwani, Samay, et al.
Published: (2026)
by: Bhojwani, Samay, et al.
Published: (2026)
RadAnnotate: Large Language Models for Efficient and Reliable Radiology Report Annotation
by: Shetty, Saisha Pradeep, et al.
Published: (2026)
by: Shetty, Saisha Pradeep, et al.
Published: (2026)
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
by: Madusanka, Tharindu, et al.
Published: (2025)
by: Madusanka, Tharindu, et al.
Published: (2025)
MultiGA: Leveraging Multi-Source Seeding in Genetic Algorithms
by: Ng, Isabelle Diana May-Xin, et al.
Published: (2025)
by: Ng, Isabelle Diana May-Xin, et al.
Published: (2025)
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
Importance Weighting Can Help Large Language Models Self-Improve
by: Jiang, Chunyang, et al.
Published: (2024)
by: Jiang, Chunyang, et al.
Published: (2024)
"Let's Agree to Disagree": Investigating the Disagreement Problem in Explainable AI for Text Summarization
by: Aswani, Seema, et al.
Published: (2024)
by: Aswani, Seema, et al.
Published: (2024)
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
by: Kunilovskaya, Maria, et al.
Published: (2026)
by: Kunilovskaya, Maria, et al.
Published: (2026)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
by: Kumarage, Tharindu, et al.
Published: (2024)
by: Kumarage, Tharindu, et al.
Published: (2024)
When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews
by: Kumar, Sandeep, et al.
Published: (2026)
by: Kumar, Sandeep, et al.
Published: (2026)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)
by: Giorgi, Tommaso, et al.
Published: (2024)
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
by: Kasprova, Vira, et al.
Published: (2026)
by: Kasprova, Vira, et al.
Published: (2026)
Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
by: Plum, Alistair, et al.
Published: (2026)
by: Plum, Alistair, et al.
Published: (2026)
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification
by: North, Kai, et al.
Published: (2022)
by: North, Kai, et al.
Published: (2022)
Exploring the Performance of Large Language Models on Subjective Span Identification Tasks
by: Dmonte, Alphaeus, et al.
Published: (2026)
by: Dmonte, Alphaeus, et al.
Published: (2026)
Why They Disagree: Decoding Differences in Opinions about AI Risk on the Lex Fridman Podcast
by: Truong, Nghi, et al.
Published: (2025)
by: Truong, Nghi, et al.
Published: (2025)
Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias
by: Ogunnubi, Tomisin, et al.
Published: (2026)
by: Ogunnubi, Tomisin, et al.
Published: (2026)
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
by: Ormerod, Christopher
Published: (2025)
by: Ormerod, Christopher
Published: (2025)
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
by: Singhvi, Arnav, et al.
Published: (2023)
by: Singhvi, Arnav, et al.
Published: (2023)
Towards Generalized Offensive Language Identification
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
MultiLS: A Multi-task Lexical Simplification Framework
by: North, Kai, et al.
Published: (2024)
by: North, Kai, et al.
Published: (2024)
StaRPO: Stability-Augmented Reinforcement Policy Optimization
by: Zhang, Jinghan, et al.
Published: (2026)
by: Zhang, Jinghan, et al.
Published: (2026)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Who Reasons in the Large Language Models?
by: Shao, Jie, et al.
Published: (2025)
by: Shao, Jie, et al.
Published: (2025)
Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment
by: Liu, Jun, et al.
Published: (2024)
by: Liu, Jun, et al.
Published: (2024)
BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
by: Ruiz, Tomas, et al.
Published: (2025)
by: Ruiz, Tomas, et al.
Published: (2025)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
by: Gu, Feng, et al.
Published: (2025)
by: Gu, Feng, et al.
Published: (2025)
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
by: Braslavski, Pavel, et al.
Published: (2026)
by: Braslavski, Pavel, et al.
Published: (2026)
DiLA: Enhancing LLM Tool Learning with Differential Logic Layer
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics
by: Nair, Jishnu Sethumadhavan, et al.
Published: (2026)
by: Nair, Jishnu Sethumadhavan, et al.
Published: (2026)
Similar Items
-
LPI-RIT at LeWiDi-2025: Improving Distributional Predictions via Metadata and Loss Reweighting with DisCo
by: Sawkar, Mandira, et al.
Published: (2025) -
ProRefine: Inference-Time Prompt Refinement with Textual Feedback
by: Pandita, Deepak, et al.
Published: (2025) -
ARTICLE: Annotator Reliability Through In-Context Learning
by: Dutta, Sujan, et al.
Published: (2024) -
Forest vs Tree: The $(N, K)$ Trade-off in Reproducible ML Evaluation
by: Pandita, Deepak, et al.
Published: (2025) -
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
by: Pandita, Deepak, et al.
Published: (2026)