Interpreting Agentic Systems: Beyond Model Explanations to System-Level Accountability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Judy, Gandhi, Dhari, Joshi, Himanshu, Mianroodi, Ahmad Rezaie, Kocak, Sedef Akinli, Ramachandran, Dhanesh
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912845202980864
author Zhu, Judy
Gandhi, Dhari
Joshi, Himanshu
Mianroodi, Ahmad Rezaie
Kocak, Sedef Akinli
Ramachandran, Dhanesh
author_facet Zhu, Judy
Gandhi, Dhari
Joshi, Himanshu
Mianroodi, Ahmad Rezaie
Kocak, Sedef Akinli
Ramachandran, Dhanesh
contents Agentic systems have transformed how Large Language Models (LLMs) can be leveraged to create autonomous systems with goal-directed behaviors, consisting of multi-step planning and the ability to interact with different environments. These systems differ fundamentally from traditional machine learning models, both in architecture and deployment, introducing unique AI safety challenges, including goal misalignment, compounding decision errors, and coordination risks among interacting agents, that necessitate embedding interpretability and explainability by design to ensure traceability and accountability across their autonomous behaviors. Current interpretability techniques, developed primarily for static models, show limitations when applied to agentic systems. The temporal dynamics, compounding decisions, and context-dependent behaviors of agentic systems demand new analytical approaches. This paper assesses the suitability and limitations of existing interpretability methods in the context of agentic systems, identifying gaps in their capacity to provide meaningful insight into agent decision-making. We propose future directions for developing interpretability techniques specifically designed for agentic systems, pinpointing where interpretability is required to embed oversight mechanisms across the agent lifecycle from goal formation, through environmental interaction, to outcome evaluation. These advances are essential to ensure the safe and accountable deployment of agentic AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17168
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Interpreting Agentic Systems: Beyond Model Explanations to System-Level Accountability
Zhu, Judy
Gandhi, Dhari
Joshi, Himanshu
Mianroodi, Ahmad Rezaie
Kocak, Sedef Akinli
Ramachandran, Dhanesh
Artificial Intelligence
Multiagent Systems
Agentic systems have transformed how Large Language Models (LLMs) can be leveraged to create autonomous systems with goal-directed behaviors, consisting of multi-step planning and the ability to interact with different environments. These systems differ fundamentally from traditional machine learning models, both in architecture and deployment, introducing unique AI safety challenges, including goal misalignment, compounding decision errors, and coordination risks among interacting agents, that necessitate embedding interpretability and explainability by design to ensure traceability and accountability across their autonomous behaviors. Current interpretability techniques, developed primarily for static models, show limitations when applied to agentic systems. The temporal dynamics, compounding decisions, and context-dependent behaviors of agentic systems demand new analytical approaches. This paper assesses the suitability and limitations of existing interpretability methods in the context of agentic systems, identifying gaps in their capacity to provide meaningful insight into agent decision-making. We propose future directions for developing interpretability techniques specifically designed for agentic systems, pinpointing where interpretability is required to embed oversight mechanisms across the agent lifecycle from goal formation, through environmental interaction, to outcome evaluation. These advances are essential to ensure the safe and accountable deployment of agentic AI systems.
title Interpreting Agentic Systems: Beyond Model Explanations to System-Level Accountability
topic Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2601.17168