Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Radharapu, Bhaktipriya, Revel, Manon, Ung, Megan, Ruder, Sebastian, Williams, Adina |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Chained Tuning Leads to Biased Forgetting
di: Ung, Megan, et al.
Pubblicazione: (2024)
di: Ung, Megan, et al.
Pubblicazione: (2024)
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
di: Sturman, Olivia, et al.
Pubblicazione: (2024)
di: Sturman, Olivia, et al.
Pubblicazione: (2024)
Changing Answer Order Can Decrease MMLU Accuracy
di: Gupta, Vipul, et al.
Pubblicazione: (2024)
di: Gupta, Vipul, et al.
Pubblicazione: (2024)
RealSeal: Revolutionizing Media Authentication with Real-Time Realism Scoring
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2024)
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2024)
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
di: Ovalle, Anaelia, et al.
Pubblicazione: (2025)
di: Ovalle, Anaelia, et al.
Pubblicazione: (2025)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
di: Gupta, Vipul, et al.
Pubblicazione: (2024)
di: Gupta, Vipul, et al.
Pubblicazione: (2024)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
Understanding and Mitigating Language Confusion in LLMs
di: Marchisio, Kelly, et al.
Pubblicazione: (2024)
di: Marchisio, Kelly, et al.
Pubblicazione: (2024)
Domain Regeneration: How well do LLMs match syntactic properties of text domains?
di: Ju, Da, et al.
Pubblicazione: (2025)
di: Ju, Da, et al.
Pubblicazione: (2025)
Language and Task Arithmetic with Parameter-Efficient Layers for Zero-Shot Summarization
di: Chronopoulou, Alexandra, et al.
Pubblicazione: (2023)
di: Chronopoulou, Alexandra, et al.
Pubblicazione: (2023)
SEAL: Systematic Error Analysis for Value ALignment
di: Revel, Manon, et al.
Pubblicazione: (2024)
di: Revel, Manon, et al.
Pubblicazione: (2024)
ShieldGemma: Generative AI Content Moderation Based on Gemma
di: Zeng, Wenjun, et al.
Pubblicazione: (2024)
di: Zeng, Wenjun, et al.
Pubblicazione: (2024)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
How Does Quantization Affect Multilingual LLMs?
di: Marchisio, Kelly, et al.
Pubblicazione: (2024)
di: Marchisio, Kelly, et al.
Pubblicazione: (2024)
AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic
di: Robinson, Nathaniel R., et al.
Pubblicazione: (2024)
di: Robinson, Nathaniel R., et al.
Pubblicazione: (2024)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
di: Shimabucoro, Luisa, et al.
Pubblicazione: (2025)
di: Shimabucoro, Luisa, et al.
Pubblicazione: (2025)
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction
di: Ruder, Sebastian, et al.
Pubblicazione: (2018)
di: Ruder, Sebastian, et al.
Pubblicazione: (2018)
Are Female Carpenters like Blue Bananas? A Corpus Investigation of Occupation Gender Typicality
di: Ju, Da, et al.
Pubblicazione: (2024)
di: Ju, Da, et al.
Pubblicazione: (2024)
Assessing the Reliability and Validity of GPT-4 in Annotating Emotion Appraisal Ratings
di: Ruder, Deniss, et al.
Pubblicazione: (2025)
di: Ruder, Deniss, et al.
Pubblicazione: (2025)
FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
di: Chen, Mingda, et al.
Pubblicazione: (2025)
di: Chen, Mingda, et al.
Pubblicazione: (2025)
AI-Enhanced Deliberative Democracy and the Future of the Collective Will
di: Revel, Manon, et al.
Pubblicazione: (2025)
di: Revel, Manon, et al.
Pubblicazione: (2025)
Actor Identification in Discourse: A Challenge for LLMs?
di: Barić, Ana, et al.
Pubblicazione: (2024)
di: Barić, Ana, et al.
Pubblicazione: (2024)
Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs
di: Reusens, Manon, et al.
Pubblicazione: (2025)
di: Reusens, Manon, et al.
Pubblicazione: (2025)
What makes a good metric? Evaluating automatic metrics for text-to-image consistency
di: Ross, Candace, et al.
Pubblicazione: (2024)
di: Ross, Candace, et al.
Pubblicazione: (2024)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
di: Shimabucoro, Luísa, et al.
Pubblicazione: (2024)
di: Shimabucoro, Luísa, et al.
Pubblicazione: (2024)
The Causal Influence of Grammatical Gender on Distributional Semantics
di: Stańczak, Karolina, et al.
Pubblicazione: (2023)
di: Stańczak, Karolina, et al.
Pubblicazione: (2023)
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
di: Warstadt, Alex, et al.
Pubblicazione: (2025)
di: Warstadt, Alex, et al.
Pubblicazione: (2025)
Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs
di: Williams, Tristan, et al.
Pubblicazione: (2026)
di: Williams, Tristan, et al.
Pubblicazione: (2026)
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
di: Hu, Michael Y., et al.
Pubblicazione: (2024)
di: Hu, Michael Y., et al.
Pubblicazione: (2024)
User Modeling Challenges in Interactive AI Assistant Systems
di: Su, Megan, et al.
Pubblicazione: (2024)
di: Su, Megan, et al.
Pubblicazione: (2024)
Arbiter: Detecting Interference in LLM Agent System Prompts
di: Mason, Tony
Pubblicazione: (2026)
di: Mason, Tony
Pubblicazione: (2026)
A Latent-Variable Model for Intrinsic Probing
di: Stańczak, Karolina, et al.
Pubblicazione: (2022)
di: Stańczak, Karolina, et al.
Pubblicazione: (2022)
[Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL
di: Wu, Zhaofeng, et al.
Pubblicazione: (2026)
di: Wu, Zhaofeng, et al.
Pubblicazione: (2026)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
di: Li, Xingxuan, et al.
Pubblicazione: (2024)
di: Li, Xingxuan, et al.
Pubblicazione: (2024)
Do different prompting methods yield a common task representation in language models?
di: Davidson, Guy, et al.
Pubblicazione: (2025)
di: Davidson, Guy, et al.
Pubblicazione: (2025)
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
di: LeVi, Amit, et al.
Pubblicazione: (2025)
di: LeVi, Amit, et al.
Pubblicazione: (2025)
Cost-Performance Optimization for Processing Low-Resource Language Tasks Using Commercial LLMs
di: Nag, Arijit, et al.
Pubblicazione: (2024)
di: Nag, Arijit, et al.
Pubblicazione: (2024)
Finetuning LLMs for Comparative Assessment Tasks
di: Raina, Vatsal, et al.
Pubblicazione: (2024)
di: Raina, Vatsal, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Chained Tuning Leads to Biased Forgetting
di: Ung, Megan, et al.
Pubblicazione: (2024) -
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
di: Sturman, Olivia, et al.
Pubblicazione: (2024) -
Changing Answer Order Can Decrease MMLU Accuracy
di: Gupta, Vipul, et al.
Pubblicazione: (2024) -
RealSeal: Revolutionizing Media Authentication with Real-Time Realism Scoring
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2024) -
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
di: Ovalle, Anaelia, et al.
Pubblicazione: (2025)