UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Zhengyang, Zhang, Daoan, Zhou, Huichi, Huang, Rui, Li, Bobo, Zhang, Yuechen, Wu, Shengqiong, Wang, Xiaohan, Luo, Jiebo, Liao, Lizi, Fei, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912702286266368
author Liang, Zhengyang
Zhang, Daoan
Zhou, Huichi
Huang, Rui
Li, Bobo
Zhang, Yuechen
Wu, Shengqiong
Wang, Xiaohan
Luo, Jiebo
Liao, Lizi
Fei, Hao
author_facet Liang, Zhengyang
Zhang, Daoan
Zhou, Huichi
Huang, Rui
Li, Bobo
Zhang, Yuechen
Wu, Shengqiong
Wang, Xiaohan
Luo, Jiebo
Liao, Lizi
Fei, Hao
contents While specialized AI models excel at isolated video tasks like generation or understanding, real-world applications demand complex, iterative workflows that combine these capabilities. To bridge this gap, we introduce UniVA, an open-source, omni-capable multi-agent framework for next-generation video generalists that unifies video understanding, segmentation, editing, and generation into cohesive workflows. UniVA employs a Plan-and-Act dual-agent architecture that drives a highly automated and proactive workflow: a planner agent interprets user intentions and decomposes them into structured video-processing steps, while executor agents execute these through modular, MCP-based tool servers (for analysis, generation, editing, tracking, etc.). Through a hierarchical multi-level memory (global knowledge, task context, and user-specific preferences), UniVA sustains long-horizon reasoning, contextual continuity, and inter-agent communication, enabling interactive and self-reflective video creation with full traceability. This design enables iterative and any-conditioned video workflows (e.g., text/image/video-conditioned generation $\rightarrow$ multi-round editing $\rightarrow$ object segmentation $\rightarrow$ compositional synthesis) that were previously cumbersome to achieve with single-purpose models or monolithic video-language models. We also introduce UniVA-Bench, a benchmark suite of multi-step video tasks spanning understanding, editing, segmentation, and generation, to rigorously evaluate such agentic video systems. Both UniVA and UniVA-Bench are fully open-sourced, aiming to catalyze research on interactive, agentic, and general-purpose video intelligence for the next generation of multimodal AI systems. (https://univa.online/)
format Preprint
id arxiv_https___arxiv_org_abs_2511_08521
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
Liang, Zhengyang
Zhang, Daoan
Zhou, Huichi
Huang, Rui
Li, Bobo
Zhang, Yuechen
Wu, Shengqiong
Wang, Xiaohan
Luo, Jiebo
Liao, Lizi
Fei, Hao
Computer Vision and Pattern Recognition
While specialized AI models excel at isolated video tasks like generation or understanding, real-world applications demand complex, iterative workflows that combine these capabilities. To bridge this gap, we introduce UniVA, an open-source, omni-capable multi-agent framework for next-generation video generalists that unifies video understanding, segmentation, editing, and generation into cohesive workflows. UniVA employs a Plan-and-Act dual-agent architecture that drives a highly automated and proactive workflow: a planner agent interprets user intentions and decomposes them into structured video-processing steps, while executor agents execute these through modular, MCP-based tool servers (for analysis, generation, editing, tracking, etc.). Through a hierarchical multi-level memory (global knowledge, task context, and user-specific preferences), UniVA sustains long-horizon reasoning, contextual continuity, and inter-agent communication, enabling interactive and self-reflective video creation with full traceability. This design enables iterative and any-conditioned video workflows (e.g., text/image/video-conditioned generation $\rightarrow$ multi-round editing $\rightarrow$ object segmentation $\rightarrow$ compositional synthesis) that were previously cumbersome to achieve with single-purpose models or monolithic video-language models. We also introduce UniVA-Bench, a benchmark suite of multi-step video tasks spanning understanding, editing, segmentation, and generation, to rigorously evaluate such agentic video systems. Both UniVA and UniVA-Bench are fully open-sourced, aiming to catalyze research on interactive, agentic, and general-purpose video intelligence for the next generation of multimodal AI systems. (https://univa.online/)
title UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.08521