ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zeng, Xiyin, Sun, Yuyu, Li, Haoyang, Liu, Shouqiang, Wang, Hao
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910208670826496
author Zeng, Xiyin
Sun, Yuyu
Li, Haoyang
Liu, Shouqiang
Wang, Hao
author_facet Zeng, Xiyin
Sun, Yuyu
Li, Haoyang
Liu, Shouqiang
Wang, Hao
contents Vision-Language-Action systems follow instructions to execute multi-step tasks in multimodal environments. Recent VLA approaches typically rely on post-hoc correction mechanisms or operate under fixed task decompositions and alignment schemes. However, once an intermediate step is mis-specified, local errors propagate through subsequent steps and eventually accumulate into cascading failures. To mitigate this compounding effect, we propose Predictive Alignment and Planning Architecture, a framework that uses prediction and contrast to adjust deviations across three levels: actions, subgoals, and trajectories. Semantic alignment is enforced at all levels using a Sinkhorn-based module and a Score-field module. The predictive correction and alignment jointly update the action generator during training, enabling it to adjust fine-grained steps to remain aligned with the overall intent. We further introduce two new metrics to quantify error propagation and recovery processes in tasks, capturing how mistakes spread and fade over long-horizon execution. Experiments show that ReCAPA achieves competitive results on embodied agent benchmarks such as VisualAgentBench, MineDojo, and AI2-THOR, outperforming strong proprietary and open-source Large Language Model baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21232
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures
Zeng, Xiyin
Sun, Yuyu
Li, Haoyang
Liu, Shouqiang
Wang, Hao
Artificial Intelligence
Vision-Language-Action systems follow instructions to execute multi-step tasks in multimodal environments. Recent VLA approaches typically rely on post-hoc correction mechanisms or operate under fixed task decompositions and alignment schemes. However, once an intermediate step is mis-specified, local errors propagate through subsequent steps and eventually accumulate into cascading failures. To mitigate this compounding effect, we propose Predictive Alignment and Planning Architecture, a framework that uses prediction and contrast to adjust deviations across three levels: actions, subgoals, and trajectories. Semantic alignment is enforced at all levels using a Sinkhorn-based module and a Score-field module. The predictive correction and alignment jointly update the action generator during training, enabling it to adjust fine-grained steps to remain aligned with the overall intent. We further introduce two new metrics to quantify error propagation and recovery processes in tasks, capturing how mistakes spread and fade over long-horizon execution. Experiments show that ReCAPA achieves competitive results on embodied agent benchmarks such as VisualAgentBench, MineDojo, and AI2-THOR, outperforming strong proprietary and open-source Large Language Model baselines.
title ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading Failures
topic Artificial Intelligence
url https://arxiv.org/abs/2604.21232