TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ziyi, Zhang, Chen, Peng, Wenjun, Wu, Qi, Wang, Xinyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913053757407232
author Wang, Ziyi
Zhang, Chen
Peng, Wenjun
Wu, Qi
Wang, Xinyu
author_facet Wang, Ziyi
Zhang, Chen
Peng, Wenjun
Wu, Qi
Wang, Xinyu
contents Explainability for Large Language Model (LLM) agents is especially challenging in interactive, partially observable settings, where decisions depend on evolving beliefs and other agents. We present \textbf{TriEx}, a tri-view explainability framework that instruments sequential decision making with aligned artifacts: (i) structured first-person self-reasoning bound to an action, (ii) explicit second-person belief states about opponents updated over time, and (iii) third-person oracle audits grounded in environment-derived reference signals. This design turns explanations from free-form narratives into evidence-anchored objects that can be compared and checked across time and perspectives. Using imperfect-information strategic games as a controlled testbed, we show that TriEx enables scalable analysis of explanation faithfulness, belief dynamics, and evaluator reliability, revealing systematic mismatches between what agents say, what they believe, and what they do. Our results highlight explainability as an interaction-dependent property and motivate multi-view, evidence-grounded evaluation for LLM agents. Code is available at https://github.com/Einsam1819/TriEx.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20043
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs
Wang, Ziyi
Zhang, Chen
Peng, Wenjun
Wu, Qi
Wang, Xinyu
Computation and Language
Artificial Intelligence
Explainability for Large Language Model (LLM) agents is especially challenging in interactive, partially observable settings, where decisions depend on evolving beliefs and other agents. We present \textbf{TriEx}, a tri-view explainability framework that instruments sequential decision making with aligned artifacts: (i) structured first-person self-reasoning bound to an action, (ii) explicit second-person belief states about opponents updated over time, and (iii) third-person oracle audits grounded in environment-derived reference signals. This design turns explanations from free-form narratives into evidence-anchored objects that can be compared and checked across time and perspectives. Using imperfect-information strategic games as a controlled testbed, we show that TriEx enables scalable analysis of explanation faithfulness, belief dynamics, and evaluator reliability, revealing systematic mismatches between what agents say, what they believe, and what they do. Our results highlight explainability as an interaction-dependent property and motivate multi-view, evidence-grounded evaluation for LLM agents. Code is available at https://github.com/Einsam1819/TriEx.
title TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.20043