Resource Rational Contractualism Should Guide AI Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Levine, Sydney, Franklin, Matija, Zhi-Xuan, Tan, Guyot, Secil Yanik, Wong, Lionel, Kilov, Daniel, Choi, Yejin, Tenenbaum, Joshua B., Goodman, Noah, Lazar, Seth, Gabriel, Iason |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in LLMs
von: Kilov, Daniel, et al.
Veröffentlicht: (2025)
von: Kilov, Daniel, et al.
Veröffentlicht: (2025)
Intuitions of Compromise: Utilitarianism vs. Contractualism
von: Moore, Jared, et al.
Veröffentlicht: (2024)
von: Moore, Jared, et al.
Veröffentlicht: (2024)
Resource‐Rational Virtual Bargaining for Moral Judgment: Toward a Probabilistic Cognitive Model
von: Diego Trujillo, et al.
Veröffentlicht: (2025)
von: Diego Trujillo, et al.
Veröffentlicht: (2025)
Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2025)
Emergence is Overrated: AGI as an Archipelago of Experts
von: Kilov, Daniel
Veröffentlicht: (2026)
von: Kilov, Daniel
Veröffentlicht: (2026)
Can Language Models Reason about Individualistic Human Values and Preferences?
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
Beyond Preferences in AI Alignment
von: Zhi-Xuan, Tan, et al.
Veröffentlicht: (2024)
von: Zhi-Xuan, Tan, et al.
Veröffentlicht: (2024)
Grounding Language about Belief in a Bayesian Theory-of-Mind
von: Ying, Lance, et al.
Veröffentlicht: (2024)
von: Ying, Lance, et al.
Veröffentlicht: (2024)
Understanding Epistemic Language with a Language-augmented Bayesian Theory of Mind
von: Ying, Lance, et al.
Veröffentlicht: (2024)
von: Ying, Lance, et al.
Veröffentlicht: (2024)
Domed Monuments Cluster at Babylonian Beru Harmonics: A Longitude Enrichment Test on the UNESCO Cultural and Mixed Heritage Corpus
von: Tenenbaum, Seth
Veröffentlicht: (2026)
von: Tenenbaum, Seth
Veröffentlicht: (2026)
Longitude Quantization in the UNESCO World Heritage Corpus - Domes, Stupas, and the Babylonian Beru
von: Tenenbaum, Seth
Veröffentlicht: (2026)
von: Tenenbaum, Seth
Veröffentlicht: (2026)
Longitude Quantization in the UNESCO World Heritage Corpus: Domes, Stupas, and the Babylonian Beru
von: Tenenbaum, Seth
Veröffentlicht: (2026)
von: Tenenbaum, Seth
Veröffentlicht: (2026)
What Makes a Maze Look Like a Maze?
von: Hsu, Joy, et al.
Veröffentlicht: (2024)
von: Hsu, Joy, et al.
Veröffentlicht: (2024)
Language-Informed Synthesis of Rational Agent Models for Grounded Theory-of-Mind Reasoning On-The-Fly
von: Ying, Lance, et al.
Veröffentlicht: (2025)
von: Ying, Lance, et al.
Veröffentlicht: (2025)
Large Language Models, and LLM-Based Agents, Should Be Used to Enhance the Digital Public Sphere
von: Lazar, Seth, et al.
Veröffentlicht: (2024)
von: Lazar, Seth, et al.
Veröffentlicht: (2024)
Rational $D(q)$-quadruples
von: Dražić, Goran, et al.
Veröffentlicht: (2020)
von: Dražić, Goran, et al.
Veröffentlicht: (2020)
Lecture II: Communicative Justice and the Distribution of Attention
von: Lazar, Seth
Veröffentlicht: (2024)
von: Lazar, Seth
Veröffentlicht: (2024)
Automatic Authorities: Power and AI
von: Lazar, Seth
Veröffentlicht: (2024)
von: Lazar, Seth
Veröffentlicht: (2024)
Frontier AI Ethics: Anticipating and Evaluating the Societal Impacts of Language Model Agents
von: Lazar, Seth
Veröffentlicht: (2024)
von: Lazar, Seth
Veröffentlicht: (2024)
Lecture I: Governing the Algorithmic City
von: Lazar, Seth
Veröffentlicht: (2024)
von: Lazar, Seth
Veröffentlicht: (2024)
Model-Free RL Agents Demonstrate System 1-Like Intentionality
von: Ashton, Hal, et al.
Veröffentlicht: (2025)
von: Ashton, Hal, et al.
Veröffentlicht: (2025)
Language and Experience: A Computational Model of Social Learning in Complex Tasks
von: Colas, Cédric, et al.
Veröffentlicht: (2025)
von: Colas, Cédric, et al.
Veröffentlicht: (2025)
Scaling up the think-aloud method
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
Should Scientific Journals Be Printed? A Personal View.
von: Goodman, David
Veröffentlicht: (2000)
von: Goodman, David
Veröffentlicht: (2000)
Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
von: Grand, Gabriel, et al.
Veröffentlicht: (2025)
von: Grand, Gabriel, et al.
Veröffentlicht: (2025)
Language Model Alignment in Multilingual Trolley Problems
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
Contractualism and Risk
von: Korbinian Rüger
Veröffentlicht: (2025)
von: Korbinian Rüger
Veröffentlicht: (2025)
Rational Diophantine sextuples with strong pair
von: Dujella, Andrej, et al.
Veröffentlicht: (2024)
von: Dujella, Andrej, et al.
Veröffentlicht: (2024)
Finding structure in logographic writing with library learning
von: Jiang, Guangyuan, et al.
Veröffentlicht: (2024)
von: Jiang, Guangyuan, et al.
Veröffentlicht: (2024)
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
Quartic Rational Diophantine Quadruples and the Euler Surface
von: Andrašek, Alen, et al.
Veröffentlicht: (2026)
von: Andrašek, Alen, et al.
Veröffentlicht: (2026)
Characterizing AI Agents for Alignment and Governance
von: Kasirzadeh, Atoosa, et al.
Veröffentlicht: (2025)
von: Kasirzadeh, Atoosa, et al.
Veröffentlicht: (2025)
People use fast, goal-directed simulation to reason about novel games
von: Zhang, Cedegao E., et al.
Veröffentlicht: (2024)
von: Zhang, Cedegao E., et al.
Veröffentlicht: (2024)
In-Context Learning Strategies Emerge Rationally
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
Systematic Identification of g-Mode Pulsations in Subdwarf B Stars for Kepler Data
von: Guyot, Nathan
Veröffentlicht: (2025)
von: Guyot, Nathan
Veröffentlicht: (2025)
The stable rank of $\mathbb{Z}[x]$ is $3$
von: Guyot, Luc
Veröffentlicht: (2021)
von: Guyot, Luc
Veröffentlicht: (2021)
La construcción territorial de cabezas de puente antárticas rivales: Ushuaia (Argentina) y Punta Arenas (Chile)
von: Sylvain Guyot
Veröffentlicht: (2013)
von: Sylvain Guyot
Veröffentlicht: (2013)
On state complexity for subword-closed languages
von: Guyot, Jérôme
Veröffentlicht: (2024)
von: Guyot, Jérôme
Veröffentlicht: (2024)
De la magia de las preguntas de la infancia a la lucidez de la interrogación filosófica Atestación de una experiencia de enseñanza de la filosofía
von: Violeta Guyot
Veröffentlicht: (2007)
von: Violeta Guyot
Veröffentlicht: (2007)
Diversidad lingüística, comunicación y espacio público
von: Jacques Guyot
Veröffentlicht: (2006)
von: Jacques Guyot
Veröffentlicht: (2006)
Ähnliche Einträge
-
Discerning What Matters: A Multi-Dimensional Assessment of Moral Competence in LLMs
von: Kilov, Daniel, et al.
Veröffentlicht: (2025) -
Intuitions of Compromise: Utilitarianism vs. Contractualism
von: Moore, Jared, et al.
Veröffentlicht: (2024) -
Resource‐Rational Virtual Bargaining for Moral Judgment: Toward a Probabilistic Cognitive Model
von: Diego Trujillo, et al.
Veröffentlicht: (2025) -
Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2025) -
Emergence is Overrated: AGI as an Archipelago of Experts
von: Kilov, Daniel
Veröffentlicht: (2026)