UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Zhengyang, Zhang, Daoan, Zhou, Huichi, Huang, Rui, Li, Bobo, Zhang, Yuechen, Wu, Shengqiong, Wang, Xiaohan, Luo, Jiebo, Liao, Lizi, Fei, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Video-Browser: Towards Agentic Open-web Video Browsing
by: Liang, Zhengyang, et al.
Published: (2025)
by: Liang, Zhengyang, et al.
Published: (2025)
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025)
by: Luo, Jiabin, et al.
Published: (2025)
A Versatile Multimodal Agent for Multimedia Content Generation
by: Zhang, Daoan, et al.
Published: (2026)
by: Zhang, Daoan, et al.
Published: (2026)
Aligned Multi-View Scripts for Universal Chart-to-Code Generation
by: Zhang, Zhihan, et al.
Published: (2026)
by: Zhang, Zhihan, et al.
Published: (2026)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
by: Liu, Kai, et al.
Published: (2026)
by: Liu, Kai, et al.
Published: (2026)
NUS-Emo at SemEval-2024 Task 3: Instruction-Tuning LLM for Multimodal Emotion-Cause Analysis in Conversations
by: Luo, Meng, et al.
Published: (2024)
by: Luo, Meng, et al.
Published: (2024)
On Path to Multimodal Generalist: General-Level and General-Bench
by: Fei, Hao, et al.
Published: (2025)
by: Fei, Hao, et al.
Published: (2025)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Cognitive Kernel: An Open-source Agent System towards Generalist Autopilots
by: Zhang, Hongming, et al.
Published: (2024)
by: Zhang, Hongming, et al.
Published: (2024)
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
by: Huang, Haojian, et al.
Published: (2025)
by: Huang, Haojian, et al.
Published: (2025)
Sphinx: Benchmarking and Modeling for LLM-Driven Pull Request Review
by: Zhang, Daoan, et al.
Published: (2026)
by: Zhang, Daoan, et al.
Published: (2026)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
by: Wang, Yaoting, et al.
Published: (2025)
by: Wang, Yaoting, et al.
Published: (2025)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
by: Fei, Hao, et al.
Published: (2023)
by: Fei, Hao, et al.
Published: (2023)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
by: Wang, Yongqi, et al.
Published: (2024)
by: Wang, Yongqi, et al.
Published: (2024)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
UniVS: Unified and Universal Video Segmentation with Prompts as Queries
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
by: Zhou, Zhenghong, et al.
Published: (2024)
by: Zhou, Zhenghong, et al.
Published: (2024)
EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
by: Lin, Jingyang, et al.
Published: (2026)
by: Lin, Jingyang, et al.
Published: (2026)
Octo: An Open-Source Generalist Robot Policy
by: Octo Model Team, et al.
Published: (2024)
by: Octo Model Team, et al.
Published: (2024)
Conversation Disentanglement with Bi-Level Contrastive Learning
by: Huang, Chengyu, et al.
Published: (2022)
by: Huang, Chengyu, et al.
Published: (2022)
Learning Brain Tumor Representation in 3D High-Resolution MR Images via Interpretable State Space Models
by: Hu, Qingqiao, et al.
Published: (2024)
by: Hu, Qingqiao, et al.
Published: (2024)
Universal Scene Graph Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
by: Zhang, Daoan, et al.
Published: (2024)
by: Zhang, Daoan, et al.
Published: (2024)
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
Global Commander and Local Operative: A Dual-Agent Framework for Scene Navigation
by: Jin, Kaiming, et al.
Published: (2026)
by: Jin, Kaiming, et al.
Published: (2026)
Aurora: Unified Video Editing with a Tool-Using Agent
by: Yu, Yongsheng, et al.
Published: (2026)
by: Yu, Yongsheng, et al.
Published: (2026)
What Happens Next? Next Scene Prediction with a Unified Video Model
by: Li, Xinjie, et al.
Published: (2025)
by: Li, Xinjie, et al.
Published: (2025)
SpeechEE: A Novel Benchmark for Speech Event Extraction
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
by: Qin, You, et al.
Published: (2024)
by: Qin, You, et al.
Published: (2024)
InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction
by: Lei, Bin, et al.
Published: (2025)
by: Lei, Bin, et al.
Published: (2025)
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
by: Ding, Xiaohan, et al.
Published: (2023)
by: Ding, Xiaohan, et al.
Published: (2023)
Video Understanding with Large Language Models: A Survey
by: Tang, Yolo Y., et al.
Published: (2023)
by: Tang, Yolo Y., et al.
Published: (2023)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
by: Zhou, Xiaofeng, et al.
Published: (2025)
by: Zhou, Xiaofeng, et al.
Published: (2025)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
by: Luo, Yongdong, et al.
Published: (2024)
by: Luo, Yongdong, et al.
Published: (2024)
Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided Refinement
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
Similar Items
-
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
by: Wang, Shuai, et al.
Published: (2025) -
Video-Browser: Towards Agentic Open-web Video Browsing
by: Liang, Zhengyang, et al.
Published: (2025) -
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025) -
A Versatile Multimodal Agent for Multimedia Content Generation
by: Zhang, Daoan, et al.
Published: (2026) -
Aligned Multi-View Scripts for Universal Chart-to-Code Generation
by: Zhang, Zhihan, et al.
Published: (2026)