PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Vuddanti, Sri Vatsa, Shah, Aarav, Chittiprolu, Satwik Kumar, Song, Tony, Dev, Sunishchal, Zhu, Kevin, Chaudhary, Maheep
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912615409647616
author Vuddanti, Sri Vatsa
Shah, Aarav
Chittiprolu, Satwik Kumar
Song, Tony
Dev, Sunishchal
Zhu, Kevin
Chaudhary, Maheep
author_facet Vuddanti, Sri Vatsa
Shah, Aarav
Chittiprolu, Satwik Kumar
Song, Tony
Dev, Sunishchal
Zhu, Kevin
Chaudhary, Maheep
contents Tool-augmented language agents frequently fail in real-world deployment due to tool malfunctions--timeouts, API exceptions, or inconsistent outputs--triggering cascading reasoning errors and task abandonment. Existing agent training pipelines optimize only for success trajectories, failing to expose models to the tool failures that dominate real-world usage. We propose \textbf{PALADIN}, a generalizable framework for equipping language agents with robust failure recovery capabilities. PALADIN trains on 50,000+ recovery-annotated trajectories constructed via systematic failure injection and expert demonstrations on an enhanced ToolBench dataset. Training uses LoRA-based fine-tuning to retain base capabilities while injecting recovery competence. At inference, PALADIN detects execution-time errors and retrieves the most similar case from a curated bank of 55+ failure exemplars aligned with ToolScan's taxonomy, then executes the corresponding recovery action. This approach generalizes to novel failures beyond the training distribution, retaining 95.2\% recovery performance on unseen tool APIs. Evaluation across PaladinEval and ToolReflectEval demonstrates consistent improvements in Recovery Rate (RR), Task Success Rate (TSR), Catastrophic Success Rate (CSR), and Efficiency Score (ES). PALADIN improves RR from 32.76% to 89.68% (+57% relative) over ToolBench and outperforms the strongest baseline CRITIC (76.34%) by +13.3%. Against vanilla agents, PALADIN achieves 89.86\% RR (+66% relative improvement from 23.75%). These results establish PALADIN as an effective method for building fault-tolerant agents capable of robust recovery in real-world tool environments.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25238
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
Vuddanti, Sri Vatsa
Shah, Aarav
Chittiprolu, Satwik Kumar
Song, Tony
Dev, Sunishchal
Zhu, Kevin
Chaudhary, Maheep
Machine Learning
Artificial Intelligence
Tool-augmented language agents frequently fail in real-world deployment due to tool malfunctions--timeouts, API exceptions, or inconsistent outputs--triggering cascading reasoning errors and task abandonment. Existing agent training pipelines optimize only for success trajectories, failing to expose models to the tool failures that dominate real-world usage. We propose \textbf{PALADIN}, a generalizable framework for equipping language agents with robust failure recovery capabilities. PALADIN trains on 50,000+ recovery-annotated trajectories constructed via systematic failure injection and expert demonstrations on an enhanced ToolBench dataset. Training uses LoRA-based fine-tuning to retain base capabilities while injecting recovery competence. At inference, PALADIN detects execution-time errors and retrieves the most similar case from a curated bank of 55+ failure exemplars aligned with ToolScan's taxonomy, then executes the corresponding recovery action. This approach generalizes to novel failures beyond the training distribution, retaining 95.2\% recovery performance on unseen tool APIs. Evaluation across PaladinEval and ToolReflectEval demonstrates consistent improvements in Recovery Rate (RR), Task Success Rate (TSR), Catastrophic Success Rate (CSR), and Efficiency Score (ES). PALADIN improves RR from 32.76% to 89.68% (+57% relative) over ToolBench and outperforms the strongest baseline CRITIC (76.34%) by +13.3%. Against vanilla agents, PALADIN achieves 89.86\% RR (+66% relative improvement from 23.75%). These results establish PALADIN as an effective method for building fault-tolerant agents capable of robust recovery in real-world tool environments.
title PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.25238