Evaluating LLM Alignment With Human Trust Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Debnath, Anushka, Cranefield, Stephen, Savarimuthu, Bastin Tony Roy, Lorini, Emiliano
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914374542688256
author Debnath, Anushka
Cranefield, Stephen
Savarimuthu, Bastin Tony Roy
Lorini, Emiliano
author_facet Debnath, Anushka
Cranefield, Stephen
Savarimuthu, Bastin Tony Roy
Lorini, Emiliano
contents Trust plays a pivotal role in enabling effective cooperation, reducing uncertainty, and guiding decision-making in both human interactions and multi-agent systems. Although it is significant, there is limited understanding of how large language models (LLMs) internally conceptualize and reason about trust. This work presents a white-box analysis of trust representation in EleutherAI/gpt-j-6B, using contrastive prompting to generate embedding vectors within the activation space of the LLM for diadic trust and related interpersonal relationship attributes. We first identified trust-related concepts from five established human trust models. We then determined a threshold for significant conceptual alignment by computing pairwise cosine similarities across 60 general emotional concepts. Then we measured the cosine similarities between the LLM's internal representation of trust and the derived trust-related concepts. Our results show that the internal trust representation of EleutherAI/gpt-j-6B aligns most closely with the Castelfranchi socio-cognitive model, followed by the Marsh Model. These findings indicate that LLMs encode socio-cognitive constructs in their activation space in ways that support meaningful comparative analyses, inform theories of social cognition, and support the design of human-AI collaborative systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05839
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating LLM Alignment With Human Trust Models
Debnath, Anushka
Cranefield, Stephen
Savarimuthu, Bastin Tony Roy
Lorini, Emiliano
Multiagent Systems
Artificial Intelligence
Trust plays a pivotal role in enabling effective cooperation, reducing uncertainty, and guiding decision-making in both human interactions and multi-agent systems. Although it is significant, there is limited understanding of how large language models (LLMs) internally conceptualize and reason about trust. This work presents a white-box analysis of trust representation in EleutherAI/gpt-j-6B, using contrastive prompting to generate embedding vectors within the activation space of the LLM for diadic trust and related interpersonal relationship attributes. We first identified trust-related concepts from five established human trust models. We then determined a threshold for significant conceptual alignment by computing pairwise cosine similarities across 60 general emotional concepts. Then we measured the cosine similarities between the LLM's internal representation of trust and the derived trust-related concepts. Our results show that the internal trust representation of EleutherAI/gpt-j-6B aligns most closely with the Castelfranchi socio-cognitive model, followed by the Marsh Model. These findings indicate that LLMs encode socio-cognitive constructs in their activation space in ways that support meaningful comparative analyses, inform theories of social cognition, and support the design of human-AI collaborative systems.
title Evaluating LLM Alignment With Human Trust Models
topic Multiagent Systems
Artificial Intelligence
url https://arxiv.org/abs/2603.05839