How Good Are LLMs at Processing Tool Outputs?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kate, Kiran, Rizk, Yara, Ghosh, Poulami, Gulati, Ashu, Chakraborti, Tathagata, Wright, Zidane, Agarwal, Mayank
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918303388139520
author Kate, Kiran
Rizk, Yara
Ghosh, Poulami
Gulati, Ashu
Chakraborti, Tathagata
Wright, Zidane
Agarwal, Mayank
author_facet Kate, Kiran
Rizk, Yara
Ghosh, Poulami
Gulati, Ashu
Chakraborti, Tathagata
Wright, Zidane
Agarwal, Mayank
contents Most realistic task automation problems require large language models (LLMs) to call tools, which often return complex JSON responses. These responses must be further processed to derive the information necessary for task completion. The ability of LLMs to do so is under-studied. In this paper, we study the tool response processing task and LLMs' abilities to process structured (JSON) responses. We created a dataset for this task, and evaluated 15 open and closed weight models using multiple prompting approaches. Our results show that JSON processing remains a difficult task even for frontier models across multiple prompting strategies. The optimal response processing strategy depends on both the nature and size of the tool outputs, as well as the complexity of the required reasoning. Variations in processing approaches can lead to performance differences ranging from 3\% to 50\%.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15955
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Good Are LLMs at Processing Tool Outputs?
Kate, Kiran
Rizk, Yara
Ghosh, Poulami
Gulati, Ashu
Chakraborti, Tathagata
Wright, Zidane
Agarwal, Mayank
Machine Learning
Artificial Intelligence
Most realistic task automation problems require large language models (LLMs) to call tools, which often return complex JSON responses. These responses must be further processed to derive the information necessary for task completion. The ability of LLMs to do so is under-studied. In this paper, we study the tool response processing task and LLMs' abilities to process structured (JSON) responses. We created a dataset for this task, and evaluated 15 open and closed weight models using multiple prompting approaches. Our results show that JSON processing remains a difficult task even for frontier models across multiple prompting strategies. The optimal response processing strategy depends on both the nature and size of the tool outputs, as well as the complexity of the required reasoning. Variations in processing approaches can lead to performance differences ranging from 3\% to 50\%.
title How Good Are LLMs at Processing Tool Outputs?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.15955