Language Model Alignment in Multilingual Trolley Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Zhijing, Kleiman-Weiner, Max, Piatti, Giorgio, Levine, Sydney, Liu, Jiarui, Gonzalez, Fernando, Ortu, Francesco, Strausz, András, Sachan, Mrinmaya, Mihalcea, Rada, Choi, Yejin, Schölkopf, Bernhard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
by: Piatti, Giorgio, et al.
Published: (2024)
by: Piatti, Giorgio, et al.
Published: (2024)
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis
by: Lyu, Zhiheng, et al.
Published: (2024)
by: Lyu, Zhiheng, et al.
Published: (2024)
Are Language Models Consequentialist or Deontological Moral Reasoners?
by: Samway, Keenan, et al.
Published: (2025)
by: Samway, Keenan, et al.
Published: (2025)
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
by: Ortu, Francesco, et al.
Published: (2024)
by: Ortu, Francesco, et al.
Published: (2024)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Implicit Personalization in Language Models: A Systematic Study
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Are LLMs Good Safety Agents or a Propaganda Engine?
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025)
by: Simko, Samuel, et al.
Published: (2025)
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
by: Ceraolo, Roberto, et al.
Published: (2024)
by: Ceraolo, Roberto, et al.
Published: (2024)
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
by: Ortu, Francesco, et al.
Published: (2026)
by: Ortu, Francesco, et al.
Published: (2026)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
by: Jenny, David F., et al.
Published: (2023)
by: Jenny, David F., et al.
Published: (2023)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
CLadder: Assessing Causal Reasoning in Language Models
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
by: Backmann, Steffen, et al.
Published: (2025)
by: Backmann, Steffen, et al.
Published: (2025)
CausalCite: A Causal Formulation of Paper Citations
by: Kumar, Ishan, et al.
Published: (2023)
by: Kumar, Ishan, et al.
Published: (2023)
The Curious Case of Curiosity across Human Cultures and LLMs
by: Borah, Angana, et al.
Published: (2025)
by: Borah, Angana, et al.
Published: (2025)
When Do Language Models Endorse Limitations on Human Rights Principles?
by: Samway, Keenan, et al.
Published: (2026)
by: Samway, Keenan, et al.
Published: (2026)
Learning to Reason Efficiently with A* Post-Training
by: Opedal, Andreas, et al.
Published: (2026)
by: Opedal, Andreas, et al.
Published: (2026)
Voices of Her: Analyzing Gender Differences in the AI Publication World
by: Ding, Yiwen, et al.
Published: (2023)
by: Ding, Yiwen, et al.
Published: (2023)
Uncovering Hidden Correctness in LLM Causal Reasoning via Symbolic Verification
by: He, Paul, et al.
Published: (2026)
by: He, Paul, et al.
Published: (2026)
Value Internalization: Learning and Generalizing from Social Reward
by: Rong, Frieda, et al.
Published: (2024)
by: Rong, Frieda, et al.
Published: (2024)
Evaluating LLMs in Open-Source Games
by: Sistla, Swadesh, et al.
Published: (2025)
by: Sistla, Swadesh, et al.
Published: (2025)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
by: Arif, Samee, et al.
Published: (2026)
by: Arif, Samee, et al.
Published: (2026)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Intuitions of Compromise: Utilitarianism vs. Contractualism
by: Moore, Jared, et al.
Published: (2024)
by: Moore, Jared, et al.
Published: (2024)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
by: Ignat, Oana, et al.
Published: (2024)
by: Ignat, Oana, et al.
Published: (2024)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
by: Harrasse, Abir, et al.
Published: (2025)
by: Harrasse, Abir, et al.
Published: (2025)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
by: Do, Heejin, et al.
Published: (2026)
by: Do, Heejin, et al.
Published: (2026)
Fluid Representations in Reasoning Models
by: Kharlapenko, Dmitrii, et al.
Published: (2026)
by: Kharlapenko, Dmitrii, et al.
Published: (2026)
Causality can systematically address the monsters under the bench(marks)
by: Leeb, Felix, et al.
Published: (2025)
by: Leeb, Felix, et al.
Published: (2025)
Causal Responsibility Attribution for Human-AI Collaboration
by: Qi, Yahang, et al.
Published: (2024)
by: Qi, Yahang, et al.
Published: (2024)
Test of Time: Rethinking Temporal Signal of Benchmark Contamination
by: Zhang, Terry Jingchen, et al.
Published: (2025)
by: Zhang, Terry Jingchen, et al.
Published: (2025)
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
by: Cui, Peng, et al.
Published: (2025)
by: Cui, Peng, et al.
Published: (2025)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
by: Li, Jing-Jing, et al.
Published: (2024)
by: Li, Jing-Jing, et al.
Published: (2024)
A Learned Cost Model-based Cross-engine Optimizer for SQL Workloads
by: Strausz, András, et al.
Published: (2025)
by: Strausz, András, et al.
Published: (2025)
World Models for Math Story Problems
by: Opedal, Andreas, et al.
Published: (2023)
by: Opedal, Andreas, et al.
Published: (2023)
Boundedly Rational Meta-Learning in Sequential Consumer Choice
by: Khosravi, Mehrzad, et al.
Published: (2026)
by: Khosravi, Mehrzad, et al.
Published: (2026)
Similar Items
-
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
by: Piatti, Giorgio, et al.
Published: (2024) -
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis
by: Lyu, Zhiheng, et al.
Published: (2024) -
Are Language Models Consequentialist or Deontological Moral Reasoners?
by: Samway, Keenan, et al.
Published: (2025) -
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
by: Ortu, Francesco, et al.
Published: (2024) -
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)