Training and Evaluating with Human Label Variation: An Empirical Study
Fuente:
arXiv
Salvato in:
| Autori principali: | Kurniawan, Kemal, Mistica, Meladel, Baldwin, Timothy, Lau, Jey Han |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the Interplay between Human Label Variation and Model Fairness
di: Kurniawan, Kemal, et al.
Pubblicazione: (2025)
di: Kurniawan, Kemal, et al.
Pubblicazione: (2025)
To Aggregate or Not to Aggregate. That is the Question: A Case Study on Annotation Subjectivity in Span Prediction
di: Kurniawan, Kemal, et al.
Pubblicazione: (2024)
di: Kurniawan, Kemal, et al.
Pubblicazione: (2024)
MoDEM: Mixture of Domain Expert Models
di: Simonds, Toby, et al.
Pubblicazione: (2024)
di: Simonds, Toby, et al.
Pubblicazione: (2024)
Evaluating Evidence Attribution in Generated Fact Checking Explanations
di: Xing, Rui, et al.
Pubblicazione: (2024)
di: Xing, Rui, et al.
Pubblicazione: (2024)
WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks
di: Shetty, Anudeex, et al.
Pubblicazione: (2024)
di: Shetty, Anudeex, et al.
Pubblicazione: (2024)
A Joint Multitask Model for Morpho-Syntactic Parsing
di: Inostroza, Demian, et al.
Pubblicazione: (2025)
di: Inostroza, Demian, et al.
Pubblicazione: (2025)
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
di: Xing, Rui, et al.
Pubblicazione: (2025)
di: Xing, Rui, et al.
Pubblicazione: (2025)
Revisiting Active Learning under (Human) Label Variation
di: Gruber, Cornelia, et al.
Pubblicazione: (2025)
di: Gruber, Cornelia, et al.
Pubblicazione: (2025)
ThermoQA: A Three-Tier Benchmark for Evaluating Thermodynamic Reasoning in Large Language Models
di: Düzkar, Kemal
Pubblicazione: (2026)
di: Düzkar, Kemal
Pubblicazione: (2026)
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
di: Kaneko, Masahiro, et al.
Pubblicazione: (2026)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2026)
Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
di: Dasgupta, Sayantan, et al.
Pubblicazione: (2026)
di: Dasgupta, Sayantan, et al.
Pubblicazione: (2026)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
RAGProbe: An Automated Approach for Evaluating RAG Applications
di: Sivasothy, Shangeetha, et al.
Pubblicazione: (2024)
di: Sivasothy, Shangeetha, et al.
Pubblicazione: (2024)
Benchmarking Gender and Political Bias in Large Language Models
di: Yang, Jinrui, et al.
Pubblicazione: (2025)
di: Yang, Jinrui, et al.
Pubblicazione: (2025)
A Unified Study of LoRA Variants: Taxonomy, Review, Codebase, and Empirical Evaluation
di: He, Haonan, et al.
Pubblicazione: (2026)
di: He, Haonan, et al.
Pubblicazione: (2026)
Topics as Entity Clusters: Entity-based Topics from Large Language Models and Graph Neural Networks
di: Loureiro, Manuel V., et al.
Pubblicazione: (2023)
di: Loureiro, Manuel V., et al.
Pubblicazione: (2023)
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
di: Zhou, Yujun, et al.
Pubblicazione: (2025)
di: Zhou, Yujun, et al.
Pubblicazione: (2025)
Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents
di: Li, Xirui, et al.
Pubblicazione: (2026)
di: Li, Xirui, et al.
Pubblicazione: (2026)
An Empirical Study of Mamba-based Language Models
di: Waleffe, Roger, et al.
Pubblicazione: (2024)
di: Waleffe, Roger, et al.
Pubblicazione: (2024)
Interaction Matters: An Evaluation Framework for Interactive Dialogue Assessment on English Second Language Conversations
di: Gao, Rena, et al.
Pubblicazione: (2024)
di: Gao, Rena, et al.
Pubblicazione: (2024)
Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval
di: Yang, Jinrui, et al.
Pubblicazione: (2023)
di: Yang, Jinrui, et al.
Pubblicazione: (2023)
An Analytical Emotion Framework of Rumour Threads on Social Media
di: Xing, Rui, et al.
Pubblicazione: (2025)
di: Xing, Rui, et al.
Pubblicazione: (2025)
Empirical Analysis of Efficient Fine-Tuning Methods for Large Pre-Trained Language Models
di: Doering, Nigel, et al.
Pubblicazione: (2024)
di: Doering, Nigel, et al.
Pubblicazione: (2024)
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
di: Ruan, Jie, et al.
Pubblicazione: (2024)
di: Ruan, Jie, et al.
Pubblicazione: (2024)
A Multi-Label Dataset of French Fake News: Human and Machine Insights
di: Icard, Benjamin, et al.
Pubblicazione: (2024)
di: Icard, Benjamin, et al.
Pubblicazione: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
di: Pangakis, Nicholas, et al.
Pubblicazione: (2024)
di: Pangakis, Nicholas, et al.
Pubblicazione: (2024)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
di: Chen, Zui, et al.
Pubblicazione: (2024)
di: Chen, Zui, et al.
Pubblicazione: (2024)
Dissecting Long-Chain-of-Thought Reasoning Models: An Empirical Study
di: Mu, Yongyu, et al.
Pubblicazione: (2025)
di: Mu, Yongyu, et al.
Pubblicazione: (2025)
Groundedness in Retrieval-augmented Long-form Generation: An Empirical Study
di: Stolfo, Alessandro
Pubblicazione: (2024)
di: Stolfo, Alessandro
Pubblicazione: (2024)
Assessing Adversarial Robustness of Large Language Models: An Empirical Study
di: Yang, Zeyu, et al.
Pubblicazione: (2024)
di: Yang, Zeyu, et al.
Pubblicazione: (2024)
HMS-BERT: Hybrid Multi-Task Self-Training for Multilingual and Multi-Label Cyberbullying Detection
di: Feng, Zixin, et al.
Pubblicazione: (2026)
di: Feng, Zixin, et al.
Pubblicazione: (2026)
SPICED: News Similarity Detection Dataset with Multiple Topics and Complexity Levels
di: Shushkevich, Elena, et al.
Pubblicazione: (2023)
di: Shushkevich, Elena, et al.
Pubblicazione: (2023)
Cross-Lingual Empirical Evaluation of Large Language Models for Arabic Medical Tasks
di: Abouzahir, Chaimae, et al.
Pubblicazione: (2026)
di: Abouzahir, Chaimae, et al.
Pubblicazione: (2026)
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
di: Jain, Samyak, et al.
Pubblicazione: (2024)
di: Jain, Samyak, et al.
Pubblicazione: (2024)
Leveraging Label Semantics and Meta-Label Refinement for Multi-Label Question Classification
di: Dong, Shi, et al.
Pubblicazione: (2024)
di: Dong, Shi, et al.
Pubblicazione: (2024)
GLaPE: Gold Label-agnostic Prompt Evaluation and Optimization for Large Language Model
di: Zhang, Xuanchang, et al.
Pubblicazione: (2024)
di: Zhang, Xuanchang, et al.
Pubblicazione: (2024)
Generalization Gaps in Political Fake News Detection: An Empirical Study on the LIAR Dataset
di: Hasan, S Mahmudul, et al.
Pubblicazione: (2025)
di: Hasan, S Mahmudul, et al.
Pubblicazione: (2025)
An Empirical Study of Multi-Generation Sampling for Jailbreak Detection in Large Language Models
di: Luo, Hanrui, et al.
Pubblicazione: (2026)
di: Luo, Hanrui, et al.
Pubblicazione: (2026)
To Label or Not to Label: Hybrid Active Learning for Neural Machine Translation
di: Azeemi, Abdul Hameed, et al.
Pubblicazione: (2024)
di: Azeemi, Abdul Hameed, et al.
Pubblicazione: (2024)
Documenti analoghi
-
On the Interplay between Human Label Variation and Model Fairness
di: Kurniawan, Kemal, et al.
Pubblicazione: (2025) -
To Aggregate or Not to Aggregate. That is the Question: A Case Study on Annotation Subjectivity in Span Prediction
di: Kurniawan, Kemal, et al.
Pubblicazione: (2024) -
MoDEM: Mixture of Domain Expert Models
di: Simonds, Toby, et al.
Pubblicazione: (2024) -
Evaluating Evidence Attribution in Generated Fact Checking Explanations
di: Xing, Rui, et al.
Pubblicazione: (2024) -
WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks
di: Shetty, Anudeex, et al.
Pubblicazione: (2024)