Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuxin, Hu, Mengxue, Lin, Zheng, Fan, Xiaoyi, Xie, Fan, Fang, Zihan, Yang, Jing, Zhu, Wenjun, Chen, Zhiwen, Lv, Chengfei, Chen, Zhe
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910252230770688
author Zhang, Yuxin
Hu, Mengxue
Lin, Zheng
Fan, Xiaoyi
Xie, Fan
Fang, Zihan
Yang, Jing
Zhu, Wenjun
Chen, Zhiwen
Lv, Chengfei
Chen, Zhe
author_facet Zhang, Yuxin
Hu, Mengxue
Lin, Zheng
Fan, Xiaoyi
Xie, Fan
Fang, Zihan
Yang, Jing
Zhu, Wenjun
Chen, Zhiwen
Lv, Chengfei
Chen, Zhe
contents Large language model (LLM) agents excel at solving complex long-horizon tasks through autonomous interaction with environments. However, their real-world deployment faces a fundamental device--cloud dilemma: on-device models are efficient but often brittle, while cloud models are stronger but costly in computation. State-of-the-art LLM device--cloud routers usually make coarse task-level decisions, which cannot adapt to the changing difficulty of multi-step agent interactions. To address this issue, we present Hera, a step-level device--cloud LLM agent coordinator for long-horizon tasks achieving a strong performance--cost Pareto frontier. Hera adopts a novel two-stage training paradigm: (1) imitation learning for cold-start, followed by (2) reinforcement learning that jointly optimizes task success and cloud usage efficiency. The first stage casts step-level routing as a supervised classification problem: the device agent is replayed on cloud trajectories, with each state labeled by the agreement between device and cloud actions. In the second stage, we perform cost-aware reinforcement learning by grouping identical states across trajectories and updating Hera with labels favoring higher expected return and fewer future cloud calls. We evaluate Hera on ALFWorld, WebShop, and AppWorld, where it consistently outperforms prior methods, achieving 92.5% of the cloud-only success rate with cloud use in only 46.3% of steps.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24598
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents
Zhang, Yuxin
Hu, Mengxue
Lin, Zheng
Fan, Xiaoyi
Xie, Fan
Fang, Zihan
Yang, Jing
Zhu, Wenjun
Chen, Zhiwen
Lv, Chengfei
Chen, Zhe
Artificial Intelligence
Multiagent Systems
Large language model (LLM) agents excel at solving complex long-horizon tasks through autonomous interaction with environments. However, their real-world deployment faces a fundamental device--cloud dilemma: on-device models are efficient but often brittle, while cloud models are stronger but costly in computation. State-of-the-art LLM device--cloud routers usually make coarse task-level decisions, which cannot adapt to the changing difficulty of multi-step agent interactions. To address this issue, we present Hera, a step-level device--cloud LLM agent coordinator for long-horizon tasks achieving a strong performance--cost Pareto frontier. Hera adopts a novel two-stage training paradigm: (1) imitation learning for cold-start, followed by (2) reinforcement learning that jointly optimizes task success and cloud usage efficiency. The first stage casts step-level routing as a supervised classification problem: the device agent is replayed on cloud trajectories, with each state labeled by the agreement between device and cloud actions. In the second stage, we perform cost-aware reinforcement learning by grouping identical states across trajectories and updating Hera with labels favoring higher expected return and fewer future cloud calls. We evaluate Hera on ALFWorld, WebShop, and AppWorld, where it consistently outperforms prior methods, achieving 92.5% of the cloud-only success rate with cloud use in only 46.3% of steps.
title Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents
topic Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2605.24598