Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | von Däniken, Pius, Deriu, Jan, Tuggener, Don, Cieliebak, Mark |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Measure of the System Dependence of Automated Metrics
di: von Däniken, Pius, et al.
Pubblicazione: (2024)
di: von Däniken, Pius, et al.
Pubblicazione: (2024)
ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videos
di: Giedemann, Patrick, et al.
Pubblicazione: (2025)
di: Giedemann, Patrick, et al.
Pubblicazione: (2025)
Error-preserving Automatic Speech Recognition of Young English Learners' Language
di: Michot, Janick, et al.
Pubblicazione: (2024)
di: Michot, Janick, et al.
Pubblicazione: (2024)
Voice Adaptation for Swiss German
di: Stucki, Samuel, et al.
Pubblicazione: (2025)
di: Stucki, Samuel, et al.
Pubblicazione: (2025)
SwissGPC v1.0 -- The Swiss German Podcasts Corpus
di: Stucki, Samuel, et al.
Pubblicazione: (2025)
di: Stucki, Samuel, et al.
Pubblicazione: (2025)
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
di: Spiliopoulou, Evangelia, et al.
Pubblicazione: (2025)
di: Spiliopoulou, Evangelia, et al.
Pubblicazione: (2025)
COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
di: Taoudi-Benchekroun, Yassine, et al.
Pubblicazione: (2025)
di: Taoudi-Benchekroun, Yassine, et al.
Pubblicazione: (2025)
Beyond String Matching: Semantic Evaluation of PDF Table Extraction
di: Horn, Pius, et al.
Pubblicazione: (2026)
di: Horn, Pius, et al.
Pubblicazione: (2026)
Automated Text Scoring in the Age of Generative AI for the GPU-poor
di: Ormerod, Christopher Michael, et al.
Pubblicazione: (2024)
di: Ormerod, Christopher Michael, et al.
Pubblicazione: (2024)
Decision and Gender Biases in Large Language Models: A Behavioral Economic Perspective
di: Corazzini, Luca, et al.
Pubblicazione: (2025)
di: Corazzini, Luca, et al.
Pubblicazione: (2025)
A+AI: Threats to Society, Remedies, and Governance
di: Byrd, Don
Pubblicazione: (2024)
di: Byrd, Don
Pubblicazione: (2024)
When Algorithms Play Favorites: Lookism in the Generation and Perception of Faces
di: Doh, Miriam, et al.
Pubblicazione: (2025)
di: Doh, Miriam, et al.
Pubblicazione: (2025)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
di: Testini, Irene, et al.
Pubblicazione: (2025)
di: Testini, Irene, et al.
Pubblicazione: (2025)
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
di: Oleson, Jon
Pubblicazione: (2024)
di: Oleson, Jon
Pubblicazione: (2024)
Automating Forecasting Question Generation and Resolution for AI Evaluation
di: Bosse, Nikos I., et al.
Pubblicazione: (2026)
di: Bosse, Nikos I., et al.
Pubblicazione: (2026)
Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems
di: Johnson, Rebecca L.
Pubblicazione: (2026)
di: Johnson, Rebecca L.
Pubblicazione: (2026)
A Study of LLMs' Preferences for Libraries and Programming Languages
di: Twist, Lukas, et al.
Pubblicazione: (2025)
di: Twist, Lukas, et al.
Pubblicazione: (2025)
Pastiche Novel Generation Creating: Fan Fiction You Love in Your Favorite Author's Style
di: Han, Xueran, et al.
Pubblicazione: (2025)
di: Han, Xueran, et al.
Pubblicazione: (2025)
Evaluating Embedding Models and Pipeline Optimization for AI Search Quality
di: Zhong, Philip, et al.
Pubblicazione: (2025)
di: Zhong, Philip, et al.
Pubblicazione: (2025)
Measuring AI R&D Automation
di: Chan, Alan, et al.
Pubblicazione: (2026)
di: Chan, Alan, et al.
Pubblicazione: (2026)
Evaluating AI Meeting Summaries with a Reusable Cross-Domain Pipeline
di: Zhong, Philip, et al.
Pubblicazione: (2026)
di: Zhong, Philip, et al.
Pubblicazione: (2026)
AI-generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity
di: Zhong, Yang, et al.
Pubblicazione: (2024)
di: Zhong, Yang, et al.
Pubblicazione: (2024)
Truth or Tribe: How In-group Favoritism Prioritize Facts in Persona Agents
di: Lei, Shijun, et al.
Pubblicazione: (2026)
di: Lei, Shijun, et al.
Pubblicazione: (2026)
Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
Branching Out: Broadening AI Measurement and Evaluation with Measurement Trees
di: Greenberg, Craig, et al.
Pubblicazione: (2025)
di: Greenberg, Craig, et al.
Pubblicazione: (2025)
Evaluating AI-Driven Automated Map Digitization in QGIS
di: Febrita, Diana
Pubblicazione: (2025)
di: Febrita, Diana
Pubblicazione: (2025)
Evaluating AI Recruitment Sourcing Tools by Human Preference
di: Slaykovskiy, Vladimir, et al.
Pubblicazione: (2025)
di: Slaykovskiy, Vladimir, et al.
Pubblicazione: (2025)
Measuring Political Preferences in AI Systems: An Integrative Approach
di: Rozado, David
Pubblicazione: (2025)
di: Rozado, David
Pubblicazione: (2025)
SynthAI: A Multi Agent Generative AI Framework for Automated Modular HLS Design Generation
di: Sheikholeslam, Seyed Arash, et al.
Pubblicazione: (2024)
di: Sheikholeslam, Seyed Arash, et al.
Pubblicazione: (2024)
Don't Play Favorites: Minority Guidance for Diffusion Models
di: Um, Soobin, et al.
Pubblicazione: (2023)
di: Um, Soobin, et al.
Pubblicazione: (2023)
Evaluating Austrian A-Level German Essays with Large Language Models for Automated Essay Scoring
di: Kubesch, Jonas, et al.
Pubblicazione: (2026)
di: Kubesch, Jonas, et al.
Pubblicazione: (2026)
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
di: Horn, Pius, et al.
Pubblicazione: (2025)
di: Horn, Pius, et al.
Pubblicazione: (2025)
Beyond Preferences in AI Alignment
di: Zhi-Xuan, Tan, et al.
Pubblicazione: (2024)
di: Zhi-Xuan, Tan, et al.
Pubblicazione: (2024)
Sampling Preferences Yields Simple Trustworthiness Scores
di: Steinle, Sean
Pubblicazione: (2025)
di: Steinle, Sean
Pubblicazione: (2025)
Open-World Evaluations for Measuring Frontier AI Capabilities
di: Kapoor, Sayash, et al.
Pubblicazione: (2026)
di: Kapoor, Sayash, et al.
Pubblicazione: (2026)
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
di: Cai, Yida, et al.
Pubblicazione: (2025)
di: Cai, Yida, et al.
Pubblicazione: (2025)
Towards Prompt Generalization: Grammar-aware Cross-Prompt Automated Essay Scoring
di: Do, Heejin, et al.
Pubblicazione: (2025)
di: Do, Heejin, et al.
Pubblicazione: (2025)
HelpSteer2-Preference: Complementing Ratings with Preferences
di: Wang, Zhilin, et al.
Pubblicazione: (2024)
di: Wang, Zhilin, et al.
Pubblicazione: (2024)
Scale over Preference: The Impact of AI-Generated Content on Online Content Ecology
di: Shi, Tianhao, et al.
Pubblicazione: (2026)
di: Shi, Tianhao, et al.
Pubblicazione: (2026)
Automating Computational Design with Generative AI
di: Ploennigs, Joern, et al.
Pubblicazione: (2023)
di: Ploennigs, Joern, et al.
Pubblicazione: (2023)
Documenti analoghi
-
A Measure of the System Dependence of Automated Metrics
di: von Däniken, Pius, et al.
Pubblicazione: (2024) -
ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videos
di: Giedemann, Patrick, et al.
Pubblicazione: (2025) -
Error-preserving Automatic Speech Recognition of Young English Learners' Language
di: Michot, Janick, et al.
Pubblicazione: (2024) -
Voice Adaptation for Swiss German
di: Stucki, Samuel, et al.
Pubblicazione: (2025) -
SwissGPC v1.0 -- The Swiss German Podcasts Corpus
di: Stucki, Samuel, et al.
Pubblicazione: (2025)