Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Nie, Shangrui, Mai, Florian, Kaczér, David, Welch, Charles, Zhao, Zhixue, Flek, Lucie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PERSPECTRA: A Scalable and Configurable Pluralist Benchmark of Perspectives from Arguments
by: Nie, Shangrui, et al.
Published: (2026)
by: Nie, Shangrui, et al.
Published: (2026)
Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
by: Jørgenvåg, Magnus, et al.
Published: (2026)
by: Jørgenvåg, Magnus, et al.
Published: (2026)
Superalignment with Dynamic Human Values
by: Mai, Florian, et al.
Published: (2025)
by: Mai, Florian, et al.
Published: (2025)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
by: Schlicht, Ipek Baris, et al.
Published: (2025)
by: Schlicht, Ipek Baris, et al.
Published: (2025)
The Muddy Waters of Modeling Empathy in Language: The Practical Impacts of Theoretical Constructs
by: Lahnala, Allison, et al.
Published: (2025)
by: Lahnala, Allison, et al.
Published: (2025)
A Critical Reflection and Forward Perspective on Empathy and Natural Language Processing
by: Lahnala, Allison, et al.
Published: (2022)
by: Lahnala, Allison, et al.
Published: (2022)
IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall?
by: Rawat, Shivam, et al.
Published: (2026)
by: Rawat, Shivam, et al.
Published: (2026)
Funzac at CoMeDi Shared Task: Modeling Annotator Disagreement from Word-In-Context Perspectives
by: Sarumi, Olufunke O., et al.
Published: (2025)
by: Sarumi, Olufunke O., et al.
Published: (2025)
Do Multilingual Large Language Models Mitigate Stereotype Bias?
by: Nie, Shangrui, et al.
Published: (2024)
by: Nie, Shangrui, et al.
Published: (2024)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
by: Kurz, Simon, et al.
Published: (2024)
by: Kurz, Simon, et al.
Published: (2024)
Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models
by: Nickel, Christian, et al.
Published: (2026)
by: Nickel, Christian, et al.
Published: (2026)
Multi-Hop Reasoning for Question Answering with Hyperbolic Representations
by: Welz, Simon, et al.
Published: (2025)
by: Welz, Simon, et al.
Published: (2025)
Exploring Robustness of LLMs to Paraphrasing Based on Sociodemographic Factors
by: Arora, Pulkit, et al.
Published: (2025)
by: Arora, Pulkit, et al.
Published: (2025)
Exploring Robustness of Multilingual LLMs on Real-World Noisy Data
by: Aliakbarzadeh, Amirhossein, et al.
Published: (2025)
by: Aliakbarzadeh, Amirhossein, et al.
Published: (2025)
ISCA: A Framework for Interview-Style Conversational Agents
by: Welch, Charles, et al.
Published: (2025)
by: Welch, Charles, et al.
Published: (2025)
ArithmAttack: Evaluating Robustness of LLMs to Noisy Context in Math Problem Solving
by: Abedin, Zain Ul, et al.
Published: (2025)
by: Abedin, Zain Ul, et al.
Published: (2025)
Pitfalls of Conversational LLMs on News Debiasing
by: Schlicht, Ipek Baris, et al.
Published: (2024)
by: Schlicht, Ipek Baris, et al.
Published: (2024)
Encoder Fine-tuning with Stochastic Sampling Outperforms Open-weight GPT in Astronomy Knowledge Extraction
by: Rawat, Shivam, et al.
Published: (2025)
by: Rawat, Shivam, et al.
Published: (2025)
RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
by: Chan, Jason, et al.
Published: (2024)
by: Chan, Jason, et al.
Published: (2024)
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
by: Fatimah, Shiza, et al.
Published: (2026)
by: Fatimah, Shiza, et al.
Published: (2026)
Corpus Considerations for Annotator Modeling and Scaling
by: Sarumi, Olufunke O., et al.
Published: (2024)
by: Sarumi, Olufunke O., et al.
Published: (2024)
Disparities in Multilingual LLM-Based Healthcare Q&A
by: Schlicht, Ipek Baris, et al.
Published: (2025)
by: Schlicht, Ipek Baris, et al.
Published: (2025)
Probing the Robustness of Theory of Mind in Large Language Models
by: Nickel, Christian, et al.
Published: (2024)
by: Nickel, Christian, et al.
Published: (2024)
Label-Consistent Data Generation for Aspect-Based Sentiment Analysis Using LLM Agents
by: Monfared, Mohammad H. A., et al.
Published: (2026)
by: Monfared, Mohammad H. A., et al.
Published: (2026)
On the Credibility of Evaluating LLMs using Survey Questions
by: Libovický, Jindřich
Published: (2026)
by: Libovický, Jindřich
Published: (2026)
Explanation Generation for Contradiction Reconciliation with LLMs
by: Chan, Jason, et al.
Published: (2026)
by: Chan, Jason, et al.
Published: (2026)
AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys
by: Lin, Chenxi, et al.
Published: (2025)
by: Lin, Chenxi, et al.
Published: (2025)
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
by: Chan, Jason, et al.
Published: (2025)
by: Chan, Jason, et al.
Published: (2025)
Can Stories Help LLMs Reason? Curating Information Space Through Narrative
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs
by: Chan, Jason, et al.
Published: (2026)
by: Chan, Jason, et al.
Published: (2026)
More Agents Improve Math Problem Solving but Adversarial Robustness Gap Persists
by: Alavi, Khashayar, et al.
Published: (2025)
by: Alavi, Khashayar, et al.
Published: (2025)
Disambiguation in Conversational Question Answering in the Era of LLMs and Agents: A Survey
by: Tanjim, Md Mehrab, et al.
Published: (2025)
by: Tanjim, Md Mehrab, et al.
Published: (2025)
From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs
by: Adilazuarda, Muhammad Farid, et al.
Published: (2025)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2025)
In-Training Defenses against Emergent Misalignment in Language Models
by: Kaczér, David, et al.
Published: (2025)
by: Kaczér, David, et al.
Published: (2025)
Tracing and Reversing Edits in LLMs
by: Youssef, Paul, et al.
Published: (2025)
by: Youssef, Paul, et al.
Published: (2025)
A Survey on Human-Centric LLMs
by: Wang, Jing Yi, et al.
Published: (2024)
by: Wang, Jing Yi, et al.
Published: (2024)
Tucano 2 Cool: Better Open Source LLMs for Portuguese
by: Corrêa, Nicholas Kluge, et al.
Published: (2026)
by: Corrêa, Nicholas Kluge, et al.
Published: (2026)
Self-Alignment: Improving Alignment of Cultural Values in LLMs via In-Context Learning
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
Similar Items
-
PERSPECTRA: A Scalable and Configurable Pluralist Benchmark of Perspectives from Arguments
by: Nie, Shangrui, et al.
Published: (2026) -
Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
by: Jørgenvåg, Magnus, et al.
Published: (2026) -
Superalignment with Dynamic Human Values
by: Mai, Florian, et al.
Published: (2025) -
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
by: Schlicht, Ipek Baris, et al.
Published: (2025) -
The Muddy Waters of Modeling Empathy in Language: The Practical Impacts of Theoretical Constructs
by: Lahnala, Allison, et al.
Published: (2025)