How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
Fuente:
arXiv
Salvato in:
| Autore principale: | Fukui, Hiroki |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A molecular clock for writing systems reveals the quantitative impact of imperial power on cultural evolution
di: Fukui, Hiroki
Pubblicazione: (2026)
di: Fukui, Hiroki
Pubblicazione: (2026)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
di: Ding, Junchen, et al.
Pubblicazione: (2025)
di: Ding, Junchen, et al.
Pubblicazione: (2025)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
di: Duan, Shitong, et al.
Pubblicazione: (2023)
di: Duan, Shitong, et al.
Pubblicazione: (2023)
Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems
di: Fukui, Hiroki
Pubblicazione: (2026)
di: Fukui, Hiroki
Pubblicazione: (2026)
Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
Semantic Consistency for Assuring Reliability of Large Language Models
di: Raj, Harsh, et al.
Pubblicazione: (2023)
di: Raj, Harsh, et al.
Pubblicazione: (2023)
Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
Alignment as Iatrogenesis: Pastoral Power, Collective Pathology, and the Structural Limits of Monolingual Safety Evaluation
di: Fukui, Hiroki
Pubblicazione: (2026)
di: Fukui, Hiroki
Pubblicazione: (2026)
TaCIE: Enhancing Instruction Comprehension in Large Language Models through Task-Centred Instruction Evolution
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
How Large Language Models are Designed to Hallucinate
di: Ackermann, Richard, et al.
Pubblicazione: (2025)
di: Ackermann, Richard, et al.
Pubblicazione: (2025)
Do Large Language Models Get Caught in Hofstadter-Mobius Loops?
di: Hryszko, Jaroslaw
Pubblicazione: (2026)
di: Hryszko, Jaroslaw
Pubblicazione: (2026)
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
di: Wang, Xing, et al.
Pubblicazione: (2025)
di: Wang, Xing, et al.
Pubblicazione: (2025)
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
di: Meadows, Gwenyth Isobel, et al.
Pubblicazione: (2024)
di: Meadows, Gwenyth Isobel, et al.
Pubblicazione: (2024)
From Argumentation to Deliberation: Perspectivized Stance Vectors for Fine-grained (Dis)agreement Analysis
di: Plenz, Moritz, et al.
Pubblicazione: (2025)
di: Plenz, Moritz, et al.
Pubblicazione: (2025)
MedSimAI: Simulation and Formative Feedback Generation to Enhance Deliberate Practice in Medical Education
di: Hicke, Yann, et al.
Pubblicazione: (2025)
di: Hicke, Yann, et al.
Pubblicazione: (2025)
Cancer Vaccine Adjuvant Name Recognition from Biomedical Literature using Large Language Models
di: Rehana, Hasin, et al.
Pubblicazione: (2025)
di: Rehana, Hasin, et al.
Pubblicazione: (2025)
Do GPT Language Models Suffer From Split Personality Disorder? The Advent Of Substrate-Free Psychometrics
di: Romero, Peter, et al.
Pubblicazione: (2024)
di: Romero, Peter, et al.
Pubblicazione: (2024)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
di: Wu, Addison J., et al.
Pubblicazione: (2026)
di: Wu, Addison J., et al.
Pubblicazione: (2026)
Teaching Language Models How to Code Like Learners: Conversational Serialization for Student Simulation
di: Koutcheme, Charles, et al.
Pubblicazione: (2026)
di: Koutcheme, Charles, et al.
Pubblicazione: (2026)
Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems
di: Fukui, Hiroki
Pubblicazione: (2026)
di: Fukui, Hiroki
Pubblicazione: (2026)
Anecdoctoring: Automated Red-Teaming Across Language and Place
di: Cuevas, Alejandro, et al.
Pubblicazione: (2025)
di: Cuevas, Alejandro, et al.
Pubblicazione: (2025)
Interactive DualChecker for Mitigating Hallucinations in Distilling Large Language Models
di: Wang, Meiyun, et al.
Pubblicazione: (2024)
di: Wang, Meiyun, et al.
Pubblicazione: (2024)
DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
di: Fu, Jiachen, et al.
Pubblicazione: (2025)
di: Fu, Jiachen, et al.
Pubblicazione: (2025)
On the Creativity of Large Language Models
di: Franceschelli, Giorgio, et al.
Pubblicazione: (2023)
di: Franceschelli, Giorgio, et al.
Pubblicazione: (2023)
Do Language Models Reason Across Languages?
di: Meng, Yan, et al.
Pubblicazione: (2026)
di: Meng, Yan, et al.
Pubblicazione: (2026)
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
di: Gallegos, Isabel O., et al.
Pubblicazione: (2024)
di: Gallegos, Isabel O., et al.
Pubblicazione: (2024)
Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
di: Sun, Yuxi, et al.
Pubblicazione: (2025)
di: Sun, Yuxi, et al.
Pubblicazione: (2025)
Cross-Language Bias Examination in Large Language Models
di: Liang, Yuxuan, et al.
Pubblicazione: (2025)
di: Liang, Yuxuan, et al.
Pubblicazione: (2025)
The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models
di: Jamshidi, Saeid, et al.
Pubblicazione: (2025)
di: Jamshidi, Saeid, et al.
Pubblicazione: (2025)
How Do Vision-Language Models Process Conflicting Information Across Modalities?
di: Hua, Tianze, et al.
Pubblicazione: (2025)
di: Hua, Tianze, et al.
Pubblicazione: (2025)
EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI
di: Kasu, Sai Kartheek Reddy
Pubblicazione: (2025)
di: Kasu, Sai Kartheek Reddy
Pubblicazione: (2025)
Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics Statements
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2024)
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2024)
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
di: Venkit, Pranav Narayanan, et al.
Pubblicazione: (2025)
di: Venkit, Pranav Narayanan, et al.
Pubblicazione: (2025)
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
di: Ahmed, Ahmed Haj, et al.
Pubblicazione: (2024)
di: Ahmed, Ahmed Haj, et al.
Pubblicazione: (2024)
LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models
di: Frisch, Ivar, et al.
Pubblicazione: (2024)
di: Frisch, Ivar, et al.
Pubblicazione: (2024)
Do Language Models Think Consistently? A Study of Value Preferences Across Varying Response Lengths
di: Nair, Inderjeet, et al.
Pubblicazione: (2025)
di: Nair, Inderjeet, et al.
Pubblicazione: (2025)
Talking the Talk Does Not Entail Walking the Walk: On the Limits of Large Language Models in Lexical Entailment Recognition
di: Greco, Candida M., et al.
Pubblicazione: (2024)
di: Greco, Candida M., et al.
Pubblicazione: (2024)
Anticipating Innovation Using Large Language Models
di: Fenoaltea, Enrico Maria, et al.
Pubblicazione: (2026)
di: Fenoaltea, Enrico Maria, et al.
Pubblicazione: (2026)
Evaluating Large Language Models for Detecting Antisemitism
di: Patel, Jay, et al.
Pubblicazione: (2025)
di: Patel, Jay, et al.
Pubblicazione: (2025)
Open-Ended Wargames with Large Language Models
di: Hogan, Daniel P., et al.
Pubblicazione: (2024)
di: Hogan, Daniel P., et al.
Pubblicazione: (2024)
Documenti analoghi
-
A molecular clock for writing systems reveals the quantitative impact of imperial power on cultural evolution
di: Fukui, Hiroki
Pubblicazione: (2026) -
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
di: Ding, Junchen, et al.
Pubblicazione: (2025) -
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
di: Duan, Shitong, et al.
Pubblicazione: (2023) -
Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems
di: Fukui, Hiroki
Pubblicazione: (2026) -
Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
di: Xu, Rongwu, et al.
Pubblicazione: (2024)