EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
Fuente:
arXiv
Guardado en:
| Autores principales: | Ghate, Kshitish, Liu, Andy, Jain, Devansh, Sorensen, Taylor, Kasirzadeh, Atoosa, Caliskan, Aylin, Diab, Mona T., Sap, Maarten |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes
por: Ghate, Kshitish, et al.
Publicado: (2025)
por: Ghate, Kshitish, et al.
Publicado: (2025)
Generative Value Conflicts Reveal LLM Priorities
por: Liu, Andy, et al.
Publicado: (2025)
por: Liu, Andy, et al.
Publicado: (2025)
Personal Information Parroting in Language Models
por: Subramani, Nishant, et al.
Publicado: (2026)
por: Subramani, Nishant, et al.
Publicado: (2026)
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
por: Ghate, Kshitish, et al.
Publicado: (2025)
por: Ghate, Kshitish, et al.
Publicado: (2025)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
por: Chopra, Harshita, et al.
Publicado: (2026)
por: Chopra, Harshita, et al.
Publicado: (2026)
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
por: Li, Wenkai, et al.
Publicado: (2024)
por: Li, Wenkai, et al.
Publicado: (2024)
Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models
por: Veerendranath, Vishruth, et al.
Publicado: (2024)
por: Veerendranath, Vishruth, et al.
Publicado: (2024)
Measurement challenges in AI catastrophic risk governance and safety frameworks
por: Kasirzadeh, Atoosa
Publicado: (2024)
por: Kasirzadeh, Atoosa
Publicado: (2024)
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
por: Jain, Devansh, et al.
Publicado: (2024)
por: Jain, Devansh, et al.
Publicado: (2024)
Evaluating Large Language Model Biases in Persona-Steered Generation
por: Liu, Andy, et al.
Publicado: (2024)
por: Liu, Andy, et al.
Publicado: (2024)
Two Types of AI Existential Risk: Decisive and Accumulative
por: Kasirzadeh, Atoosa
Publicado: (2024)
por: Kasirzadeh, Atoosa
Publicado: (2024)
Bayesian Preference Learning for Test-Time Steerable Reward Models
por: Hong, Jiwoo, et al.
Publicado: (2026)
por: Hong, Jiwoo, et al.
Publicado: (2026)
Deep Reasoning in General Purpose Agents via Structured Meta-Cognition
por: Light, Dean, et al.
Publicado: (2026)
por: Light, Dean, et al.
Publicado: (2026)
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
por: Pistilli, Giada, et al.
Publicado: (2024)
por: Pistilli, Giada, et al.
Publicado: (2024)
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
por: Kumar, Priyanshu, et al.
Publicado: (2025)
por: Kumar, Priyanshu, et al.
Publicado: (2025)
A Taxonomy of Stereotype Content in Large Language Models
por: Nicolas, Gandalf, et al.
Publicado: (2024)
por: Nicolas, Gandalf, et al.
Publicado: (2024)
LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
por: Liu, Jiarui, et al.
Publicado: (2025)
por: Liu, Jiarui, et al.
Publicado: (2025)
Can Language Models Reason about Individualistic Human Values and Preferences?
por: Jiang, Liwei, et al.
Publicado: (2024)
por: Jiang, Liwei, et al.
Publicado: (2024)
Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
por: Gueorguieva, Anna-Maria, et al.
Publicado: (2025)
por: Gueorguieva, Anna-Maria, et al.
Publicado: (2025)
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
por: Liu, Jiarui, et al.
Publicado: (2025)
por: Liu, Jiarui, et al.
Publicado: (2025)
Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
por: Wilson, Kyra, et al.
Publicado: (2024)
por: Wilson, Kyra, et al.
Publicado: (2024)
ChatGPT Perpetuates Gender Bias in Machine Translation and Ignores Non-Gendered Pronouns: Findings across Bengali and Five other Low-Resource Languages
por: Ghosh, Sourojit, et al.
Publicado: (2023)
por: Ghosh, Sourojit, et al.
Publicado: (2023)
Beyond Model Interpretability: Socio-Structural Explanations in Machine Learning
por: Smart, Andrew, et al.
Publicado: (2024)
por: Smart, Andrew, et al.
Publicado: (2024)
Automatic Generation of Model and Data Cards: A Step Towards Responsible AI
por: Liu, Jiarui, et al.
Publicado: (2024)
por: Liu, Jiarui, et al.
Publicado: (2024)
AI Safety for Everyone
por: Gyevnar, Balint, et al.
Publicado: (2025)
por: Gyevnar, Balint, et al.
Publicado: (2025)
Explanation Hacking: The perils of algorithmic recourse
por: Sullivan, Emily, et al.
Publicado: (2024)
por: Sullivan, Emily, et al.
Publicado: (2024)
Characterizing AI Agents for Alignment and Governance
por: Kasirzadeh, Atoosa, et al.
Publicado: (2025)
por: Kasirzadeh, Atoosa, et al.
Publicado: (2025)
Bridging the Gap in the Responsible AI Divides
por: Gyevnár, Bálint, et al.
Publicado: (2026)
por: Gyevnár, Bálint, et al.
Publicado: (2026)
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
por: Rao, Abhinav, et al.
Publicado: (2024)
por: Rao, Abhinav, et al.
Publicado: (2024)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
por: Wedgwood, James, et al.
Publicado: (2026)
por: Wedgwood, James, et al.
Publicado: (2026)
A Note on Bias to Complete
por: Xu, Jia, et al.
Publicado: (2024)
por: Xu, Jia, et al.
Publicado: (2024)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
por: Sorensen, Taylor, et al.
Publicado: (2023)
por: Sorensen, Taylor, et al.
Publicado: (2023)
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
por: Dash, Saloni, et al.
Publicado: (2025)
por: Dash, Saloni, et al.
Publicado: (2025)
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
por: Kim, Woojin, et al.
Publicado: (2026)
por: Kim, Woojin, et al.
Publicado: (2026)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
por: Li, Jing-Jing, et al.
Publicado: (2024)
por: Li, Jing-Jing, et al.
Publicado: (2024)
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
por: Sorensen, Taylor, et al.
Publicado: (2025)
por: Sorensen, Taylor, et al.
Publicado: (2025)
VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models
por: Raj, Chahat, et al.
Publicado: (2025)
por: Raj, Chahat, et al.
Publicado: (2025)
BiasDora: Exploring Hidden Biased Associations in Vision-Language Models
por: Raj, Chahat, et al.
Publicado: (2024)
por: Raj, Chahat, et al.
Publicado: (2024)
REALM: A Dataset of Real-World LLM Use Cases
por: Cheng, Jingwen, et al.
Publicado: (2025)
por: Cheng, Jingwen, et al.
Publicado: (2025)
Rejected Dialects: Biases Against African American Language in Reward Models
por: Mire, Joel, et al.
Publicado: (2025)
por: Mire, Joel, et al.
Publicado: (2025)
Ejemplares similares
-
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes
por: Ghate, Kshitish, et al.
Publicado: (2025) -
Generative Value Conflicts Reveal LLM Priorities
por: Liu, Andy, et al.
Publicado: (2025) -
Personal Information Parroting in Language Models
por: Subramani, Nishant, et al.
Publicado: (2026) -
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
por: Ghate, Kshitish, et al.
Publicado: (2025) -
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
por: Chopra, Harshita, et al.
Publicado: (2026)