Enregistré dans:
| Auteurs principaux: | Ruiz, Tomas, Agustoslu, Tanalp, Schwemmer, Carsten |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2603.19744 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Sign Language Sense Disambiguation
par: Grimm, Jana, et autres
Publié: (2024)
par: Grimm, Jana, et autres
Publié: (2024)
BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
par: Ruiz, Tomas, et autres
Publié: (2025)
par: Ruiz, Tomas, et autres
Publié: (2025)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
par: Chew, Robert, et autres
Publié: (2026)
par: Chew, Robert, et autres
Publié: (2026)
MLLM-as-a-Judge for Image Safety without Human Labeling
par: Wang, Zhenting, et autres
Publié: (2024)
par: Wang, Zhenting, et autres
Publié: (2024)
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
par: Sun, Lin, et autres
Publié: (2025)
par: Sun, Lin, et autres
Publié: (2025)
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
par: Xiong, Shengwu., et autres
Publié: (2025)
par: Xiong, Shengwu., et autres
Publié: (2025)
Information Extraction from Heterogeneous Documents without Ground Truth Labels using Synthetic Label Generation and Knowledge Distillation
par: Bhattacharyya, Aniket, et autres
Publié: (2024)
par: Bhattacharyya, Aniket, et autres
Publié: (2024)
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
par: Gautam, Sushant, et autres
Publié: (2026)
par: Gautam, Sushant, et autres
Publié: (2026)
Training and Evaluating with Human Label Variation: An Empirical Study
par: Kurniawan, Kemal, et autres
Publié: (2025)
par: Kurniawan, Kemal, et autres
Publié: (2025)
Information Density Principle for MLLM Benchmarks
par: Li, Chunyi, et autres
Publié: (2025)
par: Li, Chunyi, et autres
Publié: (2025)
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
par: Shaikh, Omar, et autres
Publié: (2025)
par: Shaikh, Omar, et autres
Publié: (2025)
Improve MLLM Benchmark Efficiency through Interview
par: Wen, Farong, et autres
Publié: (2025)
par: Wen, Farong, et autres
Publié: (2025)
Human Label Variation in Implicit Discourse Relation Recognition
par: Yung, Frances, et autres
Publié: (2026)
par: Yung, Frances, et autres
Publié: (2026)
On the Interplay between Human Label Variation and Model Fairness
par: Kurniawan, Kemal, et autres
Publié: (2025)
par: Kurniawan, Kemal, et autres
Publié: (2025)
Fine-grained Fallacy Detection with Human Label Variation
par: Ramponi, Alan, et autres
Publié: (2025)
par: Ramponi, Alan, et autres
Publié: (2025)
Rethinking Visual Neglect: Steering via Context-Preference for MLLM Hallucination Mitigation
par: Wu, Jingwen, et autres
Publié: (2026)
par: Wu, Jingwen, et autres
Publié: (2026)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
par: Liu, Runzhou, et autres
Publié: (2026)
par: Liu, Runzhou, et autres
Publié: (2026)
Rethinking Scientific Summarization Evaluation: Grounding Explainable Metrics on Facet-aware Benchmark
par: Chen, Xiuying, et autres
Publié: (2024)
par: Chen, Xiuying, et autres
Publié: (2024)
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
par: Baan, Joris, et autres
Publié: (2024)
par: Baan, Joris, et autres
Publié: (2024)
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
par: Chen, Beiduo, et autres
Publié: (2026)
par: Chen, Beiduo, et autres
Publié: (2026)
Revisiting Active Learning under (Human) Label Variation
par: Gruber, Cornelia, et autres
Publié: (2025)
par: Gruber, Cornelia, et autres
Publié: (2025)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
par: Krumdick, Michael, et autres
Publié: (2025)
par: Krumdick, Michael, et autres
Publié: (2025)
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
par: Horych, Tomas, et autres
Publié: (2024)
par: Horych, Tomas, et autres
Publié: (2024)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
par: Gur-Arieh, Yoav, et autres
Publié: (2026)
par: Gur-Arieh, Yoav, et autres
Publié: (2026)
VariErr NLI: Separating Annotation Error from Human Label Variation
par: Weber-Genzel, Leon, et autres
Publié: (2024)
par: Weber-Genzel, Leon, et autres
Publié: (2024)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
par: Orlikowski, Matthias, et autres
Publié: (2023)
par: Orlikowski, Matthias, et autres
Publié: (2023)
Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
par: Chen, Beiduo, et autres
Publié: (2025)
par: Chen, Beiduo, et autres
Publié: (2025)
Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations
par: Peng, Siyao, et autres
Publié: (2024)
par: Peng, Siyao, et autres
Publié: (2024)
UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages
par: Abdullahi, Tassallah, et autres
Publié: (2026)
par: Abdullahi, Tassallah, et autres
Publié: (2026)
GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
par: Lan, Xiang, et autres
Publié: (2025)
par: Lan, Xiang, et autres
Publié: (2025)
Ground Truth Generation for Multilingual Historical NLP using LLMs
par: Gladstone, Clovis, et autres
Publié: (2025)
par: Gladstone, Clovis, et autres
Publié: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
par: Kil, Jihyung, et autres
Publié: (2024)
par: Kil, Jihyung, et autres
Publié: (2024)
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
par: Hong, Pingjun, et autres
Publié: (2025)
par: Hong, Pingjun, et autres
Publié: (2025)
Beyond Single Ground Truth: Reference Monism as Epistemic Injustice in ASR Evaluation
par: Choi, Anna Seo Gyeong, et autres
Publié: (2026)
par: Choi, Anna Seo Gyeong, et autres
Publié: (2026)
MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
par: Guo, Haiyun, et autres
Publié: (2025)
par: Guo, Haiyun, et autres
Publié: (2025)
Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency
par: Morbiato, Filippo, et autres
Publié: (2025)
par: Morbiato, Filippo, et autres
Publié: (2025)
Ranking Large Language Models without Ground Truth
par: Dhurandhar, Amit, et autres
Publié: (2024)
par: Dhurandhar, Amit, et autres
Publié: (2024)
VIVA: A Benchmark for Vision-Grounded Decision-Making with Human Values
par: Hu, Zhe, et autres
Publié: (2024)
par: Hu, Zhe, et autres
Publié: (2024)
The Consensus Trap: Dissecting Subjectivity and the "Ground Truth" Illusion in Data Annotation
par: Munir, Sheza, et autres
Publié: (2026)
par: Munir, Sheza, et autres
Publié: (2026)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
par: Zhang, Xinran
Publié: (2026)
par: Zhang, Xinran
Publié: (2026)
Documents similaires
-
Sign Language Sense Disambiguation
par: Grimm, Jana, et autres
Publié: (2024) -
BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
par: Ruiz, Tomas, et autres
Publié: (2025) -
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
par: Chew, Robert, et autres
Publié: (2026) -
MLLM-as-a-Judge for Image Safety without Human Labeling
par: Wang, Zhenting, et autres
Publié: (2024) -
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
par: Sun, Lin, et autres
Publié: (2025)