The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Kirk, Hannah Rose, Whitefield, Alexander, Röttger, Paul, Bean, Andrew, Margatina, Katerina, Ciro, Juan, Mosquera, Rafael, Bartolo, Max, Williams, Adina, He, He, Vidgen, Bertie, Hale, Scott A. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
por: Kirk, Hannah Rose, et al.
Publicado: (2026)
por: Kirk, Hannah Rose, et al.
Publicado: (2026)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
por: Vidgen, Bertie, et al.
Publicado: (2023)
por: Vidgen, Bertie, et al.
Publicado: (2023)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
por: Röttger, Paul, et al.
Publicado: (2023)
por: Röttger, Paul, et al.
Publicado: (2023)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
por: Rystrøm, Jonathan, et al.
Publicado: (2025)
por: Rystrøm, Jonathan, et al.
Publicado: (2025)
Why human-AI relationships need socioaffective alignment
por: Kirk, Hannah Rose, et al.
Publicado: (2025)
por: Kirk, Hannah Rose, et al.
Publicado: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
por: Röttger, Paul, et al.
Publicado: (2024)
por: Röttger, Paul, et al.
Publicado: (2024)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
por: Kirk, Hannah Rose, et al.
Publicado: (2025)
por: Kirk, Hannah Rose, et al.
Publicado: (2025)
Classification is a RAG problem: A case study on hate speech detection
por: Willats, Richard, et al.
Publicado: (2025)
por: Willats, Richard, et al.
Publicado: (2025)
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
por: Khandelwal, Khyati, et al.
Publicado: (2023)
por: Khandelwal, Khyati, et al.
Publicado: (2023)
Measuring and Mitigating Persona Distortions from AI Writing Assistance
por: Röttger, Paul, et al.
Publicado: (2026)
por: Röttger, Paul, et al.
Publicado: (2026)
PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment
por: Xu, Chenning, et al.
Publicado: (2026)
por: Xu, Chenning, et al.
Publicado: (2026)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
por: Styles, Olly, et al.
Publicado: (2024)
por: Styles, Olly, et al.
Publicado: (2024)
A Study on Leveraging Search and Self-Feedback for Agent Reasoning
por: K, Karthikeyan, et al.
Publicado: (2025)
por: K, Karthikeyan, et al.
Publicado: (2025)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
por: Li, Nanxi, et al.
Publicado: (2025)
por: Li, Nanxi, et al.
Publicado: (2025)
Explicit Trait Inference for Multi-Agent Coordination
por: Abdurahman, Suhaib, et al.
Publicado: (2026)
por: Abdurahman, Suhaib, et al.
Publicado: (2026)
HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
por: Tonneau, Manuel, et al.
Publicado: (2024)
por: Tonneau, Manuel, et al.
Publicado: (2024)
Religious Individualisation
Publicado: (2020)
Publicado: (2020)
Biological Plausibility and Representational Alignment of Feedback Alignment in Convolutional Networks
por: Lance, Jake, et al.
Publicado: (2026)
por: Lance, Jake, et al.
Publicado: (2026)
Federated Learning with Feedback Alignment
por: Baek, Incheol, et al.
Publicado: (2025)
por: Baek, Incheol, et al.
Publicado: (2025)
The AI Consumer Index (ACE)
por: Benchek, Julien, et al.
Publicado: (2025)
por: Benchek, Julien, et al.
Publicado: (2025)
Co-Constructing Alignment: A Participatory Approach to Situate AI Values
por: Arzberger, Anne, et al.
Publicado: (2026)
por: Arzberger, Anne, et al.
Publicado: (2026)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
por: Shi, Zhengyan, et al.
Publicado: (2024)
por: Shi, Zhengyan, et al.
Publicado: (2024)
Democratizing Reward Design for Personal and Representative Value-Alignment
por: Blair, Carter, et al.
Publicado: (2024)
por: Blair, Carter, et al.
Publicado: (2024)
NPO: Learning Alignment and Meta-Alignment through Structured Human Feedback
por: Gaikwad, Madhava, et al.
Publicado: (2025)
por: Gaikwad, Madhava, et al.
Publicado: (2025)
PRISM: Perspective Reasoning for Integrated Synthesis and Mediation as a Multi-Perspective Framework for AI Alignment
por: Diamond, Anthony
Publicado: (2025)
por: Diamond, Anthony
Publicado: (2025)
CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions
por: Alkhouli, Tamer, et al.
Publicado: (2025)
por: Alkhouli, Tamer, et al.
Publicado: (2025)
Clinical knowledge in LLMs does not translate to human interactions
por: Bean, Andrew M., et al.
Publicado: (2025)
por: Bean, Andrew M., et al.
Publicado: (2025)
On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation
por: Wen, Xueru, et al.
Publicado: (2024)
por: Wen, Xueru, et al.
Publicado: (2024)
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
por: Padmakumar, Vishakh, et al.
Publicado: (2024)
por: Padmakumar, Vishakh, et al.
Publicado: (2024)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
por: Chen, Tiejin, et al.
Publicado: (2026)
por: Chen, Tiejin, et al.
Publicado: (2026)
Representative Social Choice: From Learning Theory to AI Alignment
por: Qiu, Tianyi
Publicado: (2024)
por: Qiu, Tianyi
Publicado: (2024)
AnimateZoo: Zero-shot Video Generation of Cross-Species Animation via Subject Alignment
por: Xu, Yuanfeng, et al.
Publicado: (2024)
por: Xu, Yuanfeng, et al.
Publicado: (2024)
Truth-Revealing Participatory Budgeting
por: Han, Qishen, et al.
Publicado: (2026)
por: Han, Qishen, et al.
Publicado: (2026)
Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM
por: Hu, Zizhao, et al.
Publicado: (2026)
por: Hu, Zizhao, et al.
Publicado: (2026)
An Aligned Sub-Neptune Revealed with MAROON-X and a Tendency Towards Alignment for Small Planets
por: Polanski, Alex S., et al.
Publicado: (2025)
por: Polanski, Alex S., et al.
Publicado: (2025)
Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory Management
por: Ma, Weitao, et al.
Publicado: (2026)
por: Ma, Weitao, et al.
Publicado: (2026)
SAIL: Self-Amplified Iterative Learning for Diffusion Model Alignment with Minimal Human Feedback
por: He, Xiaoxuan, et al.
Publicado: (2026)
por: He, Xiaoxuan, et al.
Publicado: (2026)
Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment
por: Tian, Ran, et al.
Publicado: (2024)
por: Tian, Ran, et al.
Publicado: (2024)
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
por: Bean, Andrew M., et al.
Publicado: (2024)
por: Bean, Andrew M., et al.
Publicado: (2024)
The Polytechnic Library and Individualised Learning.
por: Revill, D. H.
Publicado: (1981)
por: Revill, D. H.
Publicado: (1981)
Ejemplares similares
-
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
por: Kirk, Hannah Rose, et al.
Publicado: (2026) -
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
por: Vidgen, Bertie, et al.
Publicado: (2023) -
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
por: Röttger, Paul, et al.
Publicado: (2023) -
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
por: Rystrøm, Jonathan, et al.
Publicado: (2025) -
Why human-AI relationships need socioaffective alignment
por: Kirk, Hannah Rose, et al.
Publicado: (2025)