BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
Fuente:
arXiv
Salvato in:
| Autori principali: | Ruiz, Tomas, Peng, Siyao, Plank, Barbara, Schwemmer, Carsten |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task
di: Leonardelli, Elisa, et al.
Pubblicazione: (2025)
di: Leonardelli, Elisa, et al.
Pubblicazione: (2025)
DeMeVa at LeWiDi-2025: Modeling Perspectives with In-Context Learning and Label Distribution Learning
di: Ignatev, Daniil, et al.
Pubblicazione: (2025)
di: Ignatev, Daniil, et al.
Pubblicazione: (2025)
LPI-RIT at LeWiDi-2025: Improving Distributional Predictions via Metadata and Loss Reweighting with DisCo
di: Sawkar, Mandira, et al.
Pubblicazione: (2025)
di: Sawkar, Mandira, et al.
Pubblicazione: (2025)
Opt-ICL at LeWiDi-2025: Maximizing In-Context Signal from Rater Examples via Meta-Learning
di: Sorensen, Taylor, et al.
Pubblicazione: (2025)
di: Sorensen, Taylor, et al.
Pubblicazione: (2025)
MaiBaam Annotation Guidelines
di: Blaschke, Verena, et al.
Pubblicazione: (2024)
di: Blaschke, Verena, et al.
Pubblicazione: (2024)
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
di: Tang, Yung-Chen, et al.
Pubblicazione: (2025)
di: Tang, Yung-Chen, et al.
Pubblicazione: (2025)
RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
di: Geuter, Jonathan, et al.
Pubblicazione: (2025)
di: Geuter, Jonathan, et al.
Pubblicazione: (2025)
AdaBoN: Adaptive Best-of-N Alignment
di: Raman, Vinod, et al.
Pubblicazione: (2025)
di: Raman, Vinod, et al.
Pubblicazione: (2025)
Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations
di: Peng, Siyao, et al.
Pubblicazione: (2024)
di: Peng, Siyao, et al.
Pubblicazione: (2024)
Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking
di: Ruiz, Tomas, et al.
Pubblicazione: (2026)
di: Ruiz, Tomas, et al.
Pubblicazione: (2026)
EEVEE: An Easy Annotation Tool for Natural Language Processing
di: Sorensen, Axel, et al.
Pubblicazione: (2024)
di: Sorensen, Axel, et al.
Pubblicazione: (2024)
VariErr NLI: Separating Annotation Error from Human Label Variation
di: Weber-Genzel, Leon, et al.
Pubblicazione: (2024)
di: Weber-Genzel, Leon, et al.
Pubblicazione: (2024)
CLIMATELI: Evaluating Entity Linking on Climate Change Data
di: Zhou, Shijia, et al.
Pubblicazione: (2024)
di: Zhou, Shijia, et al.
Pubblicazione: (2024)
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
di: Zuo, Longfei, et al.
Pubblicazione: (2025)
di: Zuo, Longfei, et al.
Pubblicazione: (2025)
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
di: Lan, Jian, et al.
Pubblicazione: (2024)
di: Lan, Jian, et al.
Pubblicazione: (2024)
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
di: Qiu, Jiahao, et al.
Pubblicazione: (2024)
di: Qiu, Jiahao, et al.
Pubblicazione: (2024)
Team Disagreement and Productive Persuasion
di: Bonomi, Giampaolo
Pubblicazione: (2025)
di: Bonomi, Giampaolo
Pubblicazione: (2025)
BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
di: Gui, Lin, et al.
Pubblicazione: (2024)
di: Gui, Lin, et al.
Pubblicazione: (2024)
JuniperLiu at CoMeDi Shared Task: Models as Annotators in Lexical Semantics Disagreements
di: Liu, Zhu, et al.
Pubblicazione: (2024)
di: Liu, Zhu, et al.
Pubblicazione: (2024)
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
di: Khurana, Urja, et al.
Pubblicazione: (2024)
di: Khurana, Urja, et al.
Pubblicazione: (2024)
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
di: Weber-Genzel, Leon, et al.
Pubblicazione: (2023)
di: Weber-Genzel, Leon, et al.
Pubblicazione: (2023)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
di: Chen, Beiduo, et al.
Pubblicazione: (2024)
di: Chen, Beiduo, et al.
Pubblicazione: (2024)
MultiClimate: Multimodal Stance Detection on Climate Change Videos
di: Wang, Jiawen, et al.
Pubblicazione: (2024)
di: Wang, Jiawen, et al.
Pubblicazione: (2024)
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
di: Chen, Beiduo, et al.
Pubblicazione: (2024)
di: Chen, Beiduo, et al.
Pubblicazione: (2024)
Funzac at CoMeDi Shared Task: Modeling Annotator Disagreement from Word-In-Context Perspectives
di: Sarumi, Olufunke O., et al.
Pubblicazione: (2025)
di: Sarumi, Olufunke O., et al.
Pubblicazione: (2025)
MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank
di: Blaschke, Verena, et al.
Pubblicazione: (2024)
di: Blaschke, Verena, et al.
Pubblicazione: (2024)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
di: Ni, Jingwei, et al.
Pubblicazione: (2025)
di: Ni, Jingwei, et al.
Pubblicazione: (2025)
Through the Lens of Split Vote: Exploring Disagreement, Difficulty and Calibration in Legal Case Outcome Classification
di: Xu, Shanshan, et al.
Pubblicazione: (2024)
di: Xu, Shanshan, et al.
Pubblicazione: (2024)
From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome Classification
di: Xu, Shanshan, et al.
Pubblicazione: (2023)
di: Xu, Shanshan, et al.
Pubblicazione: (2023)
The Best is Yet to Come: Graph Convolution in the Testing Phase for Multimodal Recommendation
di: Xu, Jinfeng, et al.
Pubblicazione: (2025)
di: Xu, Jinfeng, et al.
Pubblicazione: (2025)
Leveraging Annotator Disagreement for Text Classification
di: Xu, Jin, et al.
Pubblicazione: (2024)
di: Xu, Jin, et al.
Pubblicazione: (2024)
Information Asymmetry across Language Varieties: A Case Study on Cantonese-Mandarin and Bavarian-German QA
di: Pei, Renhao, et al.
Pubblicazione: (2026)
di: Pei, Renhao, et al.
Pubblicazione: (2026)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
di: Wang, Yiming, et al.
Pubblicazione: (2025)
di: Wang, Yiming, et al.
Pubblicazione: (2025)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
di: Huang, Audrey, et al.
Pubblicazione: (2025)
di: Huang, Audrey, et al.
Pubblicazione: (2025)
Calibrating Probabilistic Object Detectors with Annotator Disagreement
di: Tan, Zhi Qin, et al.
Pubblicazione: (2026)
di: Tan, Zhi Qin, et al.
Pubblicazione: (2026)
Dealing with Annotator Disagreement in Hate Speech Classification
di: Dehghan, Somaiyeh, et al.
Pubblicazione: (2025)
di: Dehghan, Somaiyeh, et al.
Pubblicazione: (2025)
DiZiNER: Disagreement-guided Instruction Refinement via Pilot Annotation Simulation for Zero-shot Named Entity Recognition
di: Kim, Siun, et al.
Pubblicazione: (2026)
di: Kim, Siun, et al.
Pubblicazione: (2026)
Hitchcock's Appetites
di: McKittrick, Casey
Pubblicazione: (2016)
di: McKittrick, Casey
Pubblicazione: (2016)
Display Week 2026 Promises to be the Best Yet
di: Ioannis (John) Kymissis
Pubblicazione: (2026)
di: Ioannis (John) Kymissis
Pubblicazione: (2026)
Documenti analoghi
-
LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task
di: Leonardelli, Elisa, et al.
Pubblicazione: (2025) -
DeMeVa at LeWiDi-2025: Modeling Perspectives with In-Context Learning and Label Distribution Learning
di: Ignatev, Daniil, et al.
Pubblicazione: (2025) -
LPI-RIT at LeWiDi-2025: Improving Distributional Predictions via Metadata and Loss Reweighting with DisCo
di: Sawkar, Mandira, et al.
Pubblicazione: (2025) -
Opt-ICL at LeWiDi-2025: Maximizing In-Context Signal from Rater Examples via Meta-Learning
di: Sorensen, Taylor, et al.
Pubblicazione: (2025) -
MaiBaam Annotation Guidelines
di: Blaschke, Verena, et al.
Pubblicazione: (2024)