Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity
Fuente:
arXiv
Salvato in:
| Autori principali: | Mishra, Pushkar, Rastogi, Charvi, Pfohl, Stephen R., Parrish, Alicia, Teh, Tian Huey, Patel, Roma, Diaz, Mark, Wang, Ding, Paganini, Michela, Prabhakaran, Vinodkumar, Aroyo, Lora, Rieser, Verena |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives
di: Wang, Ding, et al.
Pubblicazione: (2025)
di: Wang, Ding, et al.
Pubblicazione: (2025)
Insights on Disagreement Patterns in Multimodal Safety Perception across Diverse Rater Groups
di: Rastogi, Charvi, et al.
Pubblicazione: (2024)
di: Rastogi, Charvi, et al.
Pubblicazione: (2024)
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
di: Rastogi, Charvi, et al.
Pubblicazione: (2025)
di: Rastogi, Charvi, et al.
Pubblicazione: (2025)
GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives
di: Prabhakaran, Vinodkumar, et al.
Pubblicazione: (2023)
di: Prabhakaran, Vinodkumar, et al.
Pubblicazione: (2023)
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
di: Quaye, Jessica, et al.
Pubblicazione: (2025)
di: Quaye, Jessica, et al.
Pubblicazione: (2025)
Value Profiles for Encoding Human Variation
di: Sorensen, Taylor, et al.
Pubblicazione: (2025)
di: Sorensen, Taylor, et al.
Pubblicazione: (2025)
Taxonomy of User Needs and Actions
di: Shelby, Renee, et al.
Pubblicazione: (2025)
di: Shelby, Renee, et al.
Pubblicazione: (2025)
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
di: Rastogi, Charvi, et al.
Pubblicazione: (2026)
di: Rastogi, Charvi, et al.
Pubblicazione: (2026)
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation
di: Quaye, Jessica, et al.
Pubblicazione: (2024)
di: Quaye, Jessica, et al.
Pubblicazione: (2024)
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
di: Ibrahim, Lujain, et al.
Pubblicazione: (2025)
di: Ibrahim, Lujain, et al.
Pubblicazione: (2025)
A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluations
di: Davani, Aida, et al.
Pubblicazione: (2025)
di: Davani, Aida, et al.
Pubblicazione: (2025)
D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation
di: Davani, Aida Mostafazadeh, et al.
Pubblicazione: (2024)
di: Davani, Aida Mostafazadeh, et al.
Pubblicazione: (2024)
SeeGULL Multilingual: a Dataset of Geo-Culturally Situated Stereotypes
di: Bhutani, Mukul, et al.
Pubblicazione: (2024)
di: Bhutani, Mukul, et al.
Pubblicazione: (2024)
Risks of Cultural Erasure in Large Language Models
di: Qadri, Rida, et al.
Pubblicazione: (2025)
di: Qadri, Rida, et al.
Pubblicazione: (2025)
GeniL: A Multilingual Dataset on Generalizing Language
di: Davani, Aida Mostafazadeh, et al.
Pubblicazione: (2024)
di: Davani, Aida Mostafazadeh, et al.
Pubblicazione: (2024)
Towards Geo-Culturally Grounded LLM Generations
di: Lertvittayakumjorn, Piyawat, et al.
Pubblicazione: (2025)
di: Lertvittayakumjorn, Piyawat, et al.
Pubblicazione: (2025)
Yesterday's News: Benchmarking Multi-Dimensional Out-of-Distribution Generalization of Misinformation Detection Models
di: Verhoeven, Ivo, et al.
Pubblicazione: (2024)
di: Verhoeven, Ivo, et al.
Pubblicazione: (2024)
SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context
di: Verma, Aishwarya, et al.
Pubblicazione: (2026)
di: Verma, Aishwarya, et al.
Pubblicazione: (2026)
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
di: Jha, Akshita, et al.
Pubblicazione: (2024)
di: Jha, Akshita, et al.
Pubblicazione: (2024)
Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement
di: Abercrombie, Gavin, et al.
Pubblicazione: (2023)
di: Abercrombie, Gavin, et al.
Pubblicazione: (2023)
DecoderLens: Layerwise Interpretation of Encoder-Decoder Transformers
di: Langedijk, Anna, et al.
Pubblicazione: (2023)
di: Langedijk, Anna, et al.
Pubblicazione: (2023)
Understanding and evaluating computer vision models through the lens of counterfactuals
di: Shukla, Pushkar
Pubblicazione: (2025)
di: Shukla, Pushkar
Pubblicazione: (2025)
Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence
di: Resnick, Paul, et al.
Pubblicazione: (2021)
di: Resnick, Paul, et al.
Pubblicazione: (2021)
A Randomized Controlled Trial on Anonymizing Reviewers to Each Other in Peer Review Discussions
di: Rastogi, Charvi, et al.
Pubblicazione: (2024)
di: Rastogi, Charvi, et al.
Pubblicazione: (2024)
Cultural Compass: A Framework for Organizing Societal Norms to Detect Violations in Human-AI Conversations
di: Cheng, Myra, et al.
Pubblicazione: (2026)
di: Cheng, Myra, et al.
Pubblicazione: (2026)
Multi-Rater Calibrated Segmentation Models
di: Riera-Marín, Meritxell, et al.
Pubblicazione: (2026)
di: Riera-Marín, Meritxell, et al.
Pubblicazione: (2026)
Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES)
di: Ivetta, Guido, et al.
Pubblicazione: (2025)
di: Ivetta, Guido, et al.
Pubblicazione: (2025)
Rater Cohesion and Quality from a Vicarious Perspective
di: Pandita, Deepak, et al.
Pubblicazione: (2024)
di: Pandita, Deepak, et al.
Pubblicazione: (2024)
Positive Alignment: Artificial Intelligence for Human Flourishing
di: Laukkonen, Ruben, et al.
Pubblicazione: (2026)
di: Laukkonen, Ruben, et al.
Pubblicazione: (2026)
A Multivariate to Bivariate Reduction for Noncommutative Rank and Related Results
di: Arvind, Vikraman, et al.
Pubblicazione: (2024)
di: Arvind, Vikraman, et al.
Pubblicazione: (2024)
On Efficient Noncommutative Polynomial Factorization via Higman Linearization
di: Arvind, V., et al.
Pubblicazione: (2022)
di: Arvind, V., et al.
Pubblicazione: (2022)
Beyond Aesthetics: Cultural Competence in Text-to-Image Models
di: Kannen, Nithish, et al.
Pubblicazione: (2024)
di: Kannen, Nithish, et al.
Pubblicazione: (2024)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
di: Jin, Jiho, et al.
Pubblicazione: (2026)
di: Jin, Jiho, et al.
Pubblicazione: (2026)
High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models
di: Lorandi, Michela, et al.
Pubblicazione: (2024)
di: Lorandi, Michela, et al.
Pubblicazione: (2024)
Community Notes are Vulnerable to Rater Bias and Manipulation
di: Truong, Bao Tran, et al.
Pubblicazione: (2025)
di: Truong, Bao Tran, et al.
Pubblicazione: (2025)
LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop
di: Zhao, Runcong, et al.
Pubblicazione: (2025)
di: Zhao, Runcong, et al.
Pubblicazione: (2025)
Malaysian English News Decoded: A Linguistic Resource for Named Entity and Relation Extraction
di: Chanthran, Mohan Raj, et al.
Pubblicazione: (2024)
di: Chanthran, Mohan Raj, et al.
Pubblicazione: (2024)
Scaling Cultural Resources for Improving Generative Models
di: Stepanyan, Hayk, et al.
Pubblicazione: (2025)
di: Stepanyan, Hayk, et al.
Pubblicazione: (2025)
MetaScoreLens: Evaluating User Feedback Across Digital Entertainment Systems
di: Ellington, Christian, et al.
Pubblicazione: (2025)
di: Ellington, Christian, et al.
Pubblicazione: (2025)
Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
di: Zhang, Xingjian, et al.
Pubblicazione: (2025)
di: Zhang, Xingjian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives
di: Wang, Ding, et al.
Pubblicazione: (2025) -
Insights on Disagreement Patterns in Multimodal Safety Perception across Diverse Rater Groups
di: Rastogi, Charvi, et al.
Pubblicazione: (2024) -
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
di: Rastogi, Charvi, et al.
Pubblicazione: (2025) -
GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives
di: Prabhakaran, Vinodkumar, et al.
Pubblicazione: (2023) -
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
di: Quaye, Jessica, et al.
Pubblicazione: (2025)