Are Language Models Sensitive to Morally Irrelevant Distractors?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shaw, Andrew, Hahn, Christina, Rasgaitis, Catherine, Mishra, Yash, Liu, Alisa, Jaques, Natasha, Tsvetkov, Yulia, Zhang, Amy X. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
von: Wu, Addison J., et al.
Veröffentlicht: (2026)
von: Wu, Addison J., et al.
Veröffentlicht: (2026)
Tuning Language Models by Proxy
von: Liu, Alisa, et al.
Veröffentlicht: (2024)
von: Liu, Alisa, et al.
Veröffentlicht: (2024)
DiVERT: Distractor Generation with Variational Errors Represented as Text for Math Multiple-choice Questions
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions
von: Preniqi, Vjosa, et al.
Veröffentlicht: (2024)
von: Preniqi, Vjosa, et al.
Veröffentlicht: (2024)
Language Models as Critical Thinking Tools: A Case Study of Philosophers
von: Ye, Andre, et al.
Veröffentlicht: (2024)
von: Ye, Andre, et al.
Veröffentlicht: (2024)
Finding Flawed Fictions: Evaluating Complex Reasoning in Language Models via Plot Hole Detection
von: Ahuja, Kabir, et al.
Veröffentlicht: (2025)
von: Ahuja, Kabir, et al.
Veröffentlicht: (2025)
The Moral Machine Experiment on Large Language Models
von: Takemoto, Kazuhiro
Veröffentlicht: (2023)
von: Takemoto, Kazuhiro
Veröffentlicht: (2023)
Investigating Political and Demographic Associations in Large Language Models Through Moral Foundations Theory
von: Smith-Vaniz, Nicole, et al.
Veröffentlicht: (2025)
von: Smith-Vaniz, Nicole, et al.
Veröffentlicht: (2025)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
von: Walsh, Cole, et al.
Veröffentlicht: (2026)
von: Walsh, Cole, et al.
Veröffentlicht: (2026)
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
von: Wang, Yike, et al.
Veröffentlicht: (2025)
von: Wang, Yike, et al.
Veröffentlicht: (2025)
The Moral Gap of Large Language Models
von: Skorski, Maciej, et al.
Veröffentlicht: (2025)
von: Skorski, Maciej, et al.
Veröffentlicht: (2025)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
von: Sclar, Melanie, et al.
Veröffentlicht: (2023)
von: Sclar, Melanie, et al.
Veröffentlicht: (2023)
Fine-grained Hallucination Detection and Editing for Language Models
von: Mishra, Abhika, et al.
Veröffentlicht: (2024)
von: Mishra, Abhika, et al.
Veröffentlicht: (2024)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
von: Sauter, Adrian, et al.
Veröffentlicht: (2026)
von: Sauter, Adrian, et al.
Veröffentlicht: (2026)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
von: Costa, Davi Bastos, et al.
Veröffentlicht: (2025)
von: Costa, Davi Bastos, et al.
Veröffentlicht: (2025)
Whose Emotions and Moral Sentiments Do Language Models Reflect?
von: He, Zihao, et al.
Veröffentlicht: (2024)
von: He, Zihao, et al.
Veröffentlicht: (2024)
That is Unacceptable: the Moral Foundations of Canceling
von: Lo, Soda Marem, et al.
Veröffentlicht: (2025)
von: Lo, Soda Marem, et al.
Veröffentlicht: (2025)
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
von: Morabito, Robert, et al.
Veröffentlicht: (2024)
von: Morabito, Robert, et al.
Veröffentlicht: (2024)
Generative Debunking of Climate Misinformation
von: Zanartu, Francisco, et al.
Veröffentlicht: (2024)
von: Zanartu, Francisco, et al.
Veröffentlicht: (2024)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
von: Ding, Junchen, et al.
Veröffentlicht: (2025)
von: Ding, Junchen, et al.
Veröffentlicht: (2025)
The Moral Foundations Reddit Corpus
von: Trager, Jackson, et al.
Veröffentlicht: (2022)
von: Trager, Jackson, et al.
Veröffentlicht: (2022)
SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
von: Jiang, Yuru, et al.
Veröffentlicht: (2025)
von: Jiang, Yuru, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
von: Patil, Parth, et al.
Veröffentlicht: (2026)
von: Patil, Parth, et al.
Veröffentlicht: (2026)
Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models
von: Banerjee, Somnath, et al.
Veröffentlicht: (2024)
von: Banerjee, Somnath, et al.
Veröffentlicht: (2024)
Harnessing Structured Knowledge: A Concept Map-Based Approach for High-Quality Multiple Choice Question Generation with Effective Distractors
von: Scaria, Nicy, et al.
Veröffentlicht: (2025)
von: Scaria, Nicy, et al.
Veröffentlicht: (2025)
Polarization and Morality: Lexical Analysis of Abortion Discourse on Reddit
von: Stanier, Tessa, et al.
Veröffentlicht: (2024)
von: Stanier, Tessa, et al.
Veröffentlicht: (2024)
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
von: Liu, Zhining, et al.
Veröffentlicht: (2026)
von: Liu, Zhining, et al.
Veröffentlicht: (2026)
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
A Moral Imperative: The Need for Continual Superalignment of Large Language Models
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
von: Puthumanaillam, Gokul, et al.
Veröffentlicht: (2024)
Does Cross-Cultural Alignment Change the Commonsense Morality of Language Models?
von: Jinnai, Yuu
Veröffentlicht: (2024)
von: Jinnai, Yuu
Veröffentlicht: (2024)
Recognition Without Authorization: LLMs and the Moral Order of Online Advice
von: van Nuenen, Tom
Veröffentlicht: (2026)
von: van Nuenen, Tom
Veröffentlicht: (2026)
MoVa: Towards Generalizable Classification of Human Morals and Values
von: Chen, Ziyu, et al.
Veröffentlicht: (2025)
von: Chen, Ziyu, et al.
Veröffentlicht: (2025)
Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories
von: Bilalpur, Maneesh, et al.
Veröffentlicht: (2025)
von: Bilalpur, Maneesh, et al.
Veröffentlicht: (2025)
Patterns in the Transition From Founder-Leadership to Community Governance of Open Source
von: Noori, Mobina, et al.
Veröffentlicht: (2025)
von: Noori, Mobina, et al.
Veröffentlicht: (2025)
The Single-Multi Evolution Loop for Self-Improving Model Collaboration Systems
von: Feng, Shangbin, et al.
Veröffentlicht: (2026)
von: Feng, Shangbin, et al.
Veröffentlicht: (2026)
Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems
von: Yang, Ziyuan, et al.
Veröffentlicht: (2026)
von: Yang, Ziyuan, et al.
Veröffentlicht: (2026)
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
von: Khandelwal, Khyati, et al.
Veröffentlicht: (2023)
von: Khandelwal, Khyati, et al.
Veröffentlicht: (2023)
ComPO: Community Preferences for Language Model Personalization
von: Kumar, Sachin, et al.
Veröffentlicht: (2024)
von: Kumar, Sachin, et al.
Veröffentlicht: (2024)
Irrelevant Alternatives Bias Large Language Model Hiring Decisions
von: Valkanova, Kremena, et al.
Veröffentlicht: (2024)
von: Valkanova, Kremena, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
von: Wu, Addison J., et al.
Veröffentlicht: (2026) -
Tuning Language Models by Proxy
von: Liu, Alisa, et al.
Veröffentlicht: (2024) -
DiVERT: Distractor Generation with Variational Errors Represented as Text for Math Multiple-choice Questions
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024) -
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions
von: Preniqi, Vjosa, et al.
Veröffentlicht: (2024) -
Language Models as Critical Thinking Tools: A Case Study of Philosophers
von: Ye, Andre, et al.
Veröffentlicht: (2024)