Combining Theory of Mind and Kindness for Self-Supervised Human-AI Alignment
Fuente:
arXiv
Saved in:
| Main Author: | Hewson, Joshua T. S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
We Urgently Need Intrinsically Kind Machines
by: Hewson, Joshua T. S.
Published: (2024)
by: Hewson, Joshua T. S.
Published: (2024)
Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind
by: Tong, Haibo, et al.
Published: (2024)
by: Tong, Haibo, et al.
Published: (2024)
One system for learning and remembering episodes and rules
by: Hewson, Joshua T. S., et al.
Published: (2024)
by: Hewson, Joshua T. S., et al.
Published: (2024)
Mutual Theory of Mind for Human-AI Communication
by: Wang, Qiaosi, et al.
Published: (2022)
by: Wang, Qiaosi, et al.
Published: (2022)
Machine Theory of Mind and the Structure of Human Values
by: de Font-Reaulx, Paul
Published: (2025)
by: de Font-Reaulx, Paul
Published: (2025)
Three Kinds of AI Ethics
by: Ratti, Emanuele
Published: (2025)
by: Ratti, Emanuele
Published: (2025)
Theory of Mind for Explainable Human-Robot Interaction
by: Bauer, Marie S., et al.
Published: (2025)
by: Bauer, Marie S., et al.
Published: (2025)
HCC Is All You Need: Alignment-The Sensible Kind Anyway-Is Just Human-Centered Computing
by: Gilbert, Eric
Published: (2024)
by: Gilbert, Eric
Published: (2024)
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations
by: Gong, Nanxu, et al.
Published: (2026)
by: Gong, Nanxu, et al.
Published: (2026)
Automated Meta Prompt Engineering for Alignment with the Theory of Mind
by: Baughman, Aaron, et al.
Published: (2025)
by: Baughman, Aaron, et al.
Published: (2025)
MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
by: Chen, Zehui, et al.
Published: (2024)
by: Chen, Zehui, et al.
Published: (2024)
"My Kind of Woman": Analysing Gender Stereotypes in AI through The Averageness Theory and EU Law
by: Doh, Miriam, et al.
Published: (2024)
by: Doh, Miriam, et al.
Published: (2024)
Measuring AI Alignment with Human Flourishing
by: Hilliard, Elizabeth, et al.
Published: (2025)
by: Hilliard, Elizabeth, et al.
Published: (2025)
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
by: Kim, Junsol, et al.
Published: (2026)
by: Kim, Junsol, et al.
Published: (2026)
Grounding Language about Belief in a Bayesian Theory-of-Mind
by: Ying, Lance, et al.
Published: (2024)
by: Ying, Lance, et al.
Published: (2024)
Kantian Deontology Meets AI Alignment: Towards Morally Grounded Fairness Metrics
by: Mougan, Carlos, et al.
Published: (2023)
by: Mougan, Carlos, et al.
Published: (2023)
Silico-centric Theory of Mind
by: Mukherjee, Anirban, et al.
Published: (2024)
by: Mukherjee, Anirban, et al.
Published: (2024)
LLM Theory of Mind and Alignment: Opportunities and Risks
by: Street, Winnie
Published: (2024)
by: Street, Winnie
Published: (2024)
Combining Cognitive and Generative AI for Self-explanation in Interactive AI Agents
by: Sushri, Shalini, et al.
Published: (2024)
by: Sushri, Shalini, et al.
Published: (2024)
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
Pragmatic Embodied Spoken Instruction Following in Human-Robot Collaboration with Theory of Mind
by: Ying, Lance, et al.
Published: (2024)
by: Ying, Lance, et al.
Published: (2024)
MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Understanding Epistemic Language with a Language-augmented Bayesian Theory of Mind
by: Ying, Lance, et al.
Published: (2024)
by: Ying, Lance, et al.
Published: (2024)
AI Alignment via Incentives and Correction
by: Agarwal, Rohit, et al.
Published: (2026)
by: Agarwal, Rohit, et al.
Published: (2026)
Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models
by: Kim, Hyunwoo, et al.
Published: (2025)
by: Kim, Hyunwoo, et al.
Published: (2025)
Readable Minds: Emergent Theory-of-Mind-Like Behavior in LLM Poker Agents
by: Lin, Hsieh-Ting, et al.
Published: (2026)
by: Lin, Hsieh-Ting, et al.
Published: (2026)
MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
Learning to See the Elephant in the Room: Self-Supervised Context Reasoning in Humans and AI
by: Liu, Xiao, et al.
Published: (2022)
by: Liu, Xiao, et al.
Published: (2022)
Towards Dialogues for Joint Human-AI Reasoning and Value Alignment
by: Bezou-Vrakatseli, Elfia, et al.
Published: (2024)
by: Bezou-Vrakatseli, Elfia, et al.
Published: (2024)
Human-Alignment Influences the Utility of AI-assisted Decision Making
by: Benz, Nina L. Corvelo, et al.
Published: (2025)
by: Benz, Nina L. Corvelo, et al.
Published: (2025)
Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
by: Li, Yubo, et al.
Published: (2025)
by: Li, Yubo, et al.
Published: (2025)
CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models
by: Tong, Haibo, et al.
Published: (2026)
by: Tong, Haibo, et al.
Published: (2026)
CoSToM:Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models
by: Li, Mengfan, et al.
Published: (2026)
by: Li, Mengfan, et al.
Published: (2026)
Understanding the Process of Human-AI Value Alignment
by: McKinlay, Jack, et al.
Published: (2025)
by: McKinlay, Jack, et al.
Published: (2025)
Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models
by: Cross, Logan, et al.
Published: (2024)
by: Cross, Logan, et al.
Published: (2024)
Self-Supervised Visual Preference Alignment
by: Zhu, Ke, et al.
Published: (2024)
by: Zhu, Ke, et al.
Published: (2024)
Maia-2: A Unified Model for Human-AI Alignment in Chess
by: Tang, Zhenwei, et al.
Published: (2024)
by: Tang, Zhenwei, et al.
Published: (2024)
Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?
by: Lu, Yi-Long, et al.
Published: (2025)
by: Lu, Yi-Long, et al.
Published: (2025)
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
by: Chuang, Yung-Sung, et al.
Published: (2025)
by: Chuang, Yung-Sung, et al.
Published: (2025)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
by: Sun, Zhiqing, et al.
Published: (2024)
by: Sun, Zhiqing, et al.
Published: (2024)
Similar Items
-
We Urgently Need Intrinsically Kind Machines
by: Hewson, Joshua T. S.
Published: (2024) -
Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind
by: Tong, Haibo, et al.
Published: (2024) -
One system for learning and remembering episodes and rules
by: Hewson, Joshua T. S., et al.
Published: (2024) -
Mutual Theory of Mind for Human-AI Communication
by: Wang, Qiaosi, et al.
Published: (2022) -
Machine Theory of Mind and the Structure of Human Values
by: de Font-Reaulx, Paul
Published: (2025)