Saved in:
| Main Authors: | Wong, Tsz-To, Huang, Ching-Chun, Shuai, Hong-Han |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.01853 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
by: Huang, Kaiyi, et al.
Published: (2024)
by: Huang, Kaiyi, et al.
Published: (2024)
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
by: Zhou, Yiyang, et al.
Published: (2025)
by: Zhou, Yiyang, et al.
Published: (2025)
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
by: Liu, Haowei, et al.
Published: (2025)
by: Liu, Haowei, et al.
Published: (2025)
Long-Video Audio Synthesis with Multi-Agent Collaboration
by: Zhang, Yehang, et al.
Published: (2025)
by: Zhang, Yehang, et al.
Published: (2025)
Sports-Traj: A Unified Trajectory Generation Model for Multi-Agent Movement in Sports
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning
by: Yang, Jingru, et al.
Published: (2024)
by: Yang, Jingru, et al.
Published: (2024)
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
by: Kugo, Noriyuki, et al.
Published: (2025)
by: Kugo, Noriyuki, et al.
Published: (2025)
Arbitrary-Resolution and Arbitrary-Scale Face Super-Resolution with Implicit Representation Networks
by: Tsai, Yi Ting, et al.
Published: (2025)
by: Tsai, Yi Ting, et al.
Published: (2025)
Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
by: Jiang, Haichao, et al.
Published: (2026)
by: Jiang, Haichao, et al.
Published: (2026)
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
by: Hu, Panwen, et al.
Published: (2024)
by: Hu, Panwen, et al.
Published: (2024)
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
by: Chen, Haodong, et al.
Published: (2025)
by: Chen, Haodong, et al.
Published: (2025)
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
by: Chen, Kerui, et al.
Published: (2026)
by: Chen, Kerui, et al.
Published: (2026)
AdaSports-Traj: Role- and Domain-Aware Adaptation for Multi-Agent Trajectory Modeling in Sports
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
FetalAgents: A Multi-Agent System for Fetal Ultrasound Image and Video Analysis
by: Hu, Xiaotian, et al.
Published: (2026)
by: Hu, Xiaotian, et al.
Published: (2026)
CREA: A Collaborative Multi-Agent Framework for Creative Image Editing and Generation
by: Venkatesh, Kavana, et al.
Published: (2025)
by: Venkatesh, Kavana, et al.
Published: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Automated Detection of Sport Highlights from Audio and Video Sources
by: Della Santa, Francesco, et al.
Published: (2025)
by: Della Santa, Francesco, et al.
Published: (2025)
Single Document Image Highlight Removal via A Large-Scale Real-World Dataset and A Location-Aware Network
by: Pan, Lu, et al.
Published: (2025)
by: Pan, Lu, et al.
Published: (2025)
Aerial View River Landform Video segmentation: A Weakly Supervised Context-aware Temporal Consistency Distillation Approach
by: Chen, Chi-Han, et al.
Published: (2025)
by: Chen, Chi-Han, et al.
Published: (2025)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
by: He, Liu, et al.
Published: (2024)
by: He, Liu, et al.
Published: (2024)
Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework
by: Han, Xiao, et al.
Published: (2024)
by: Han, Xiao, et al.
Published: (2024)
Pragmatic Communication in Multi-Agent Collaborative Perception
by: Hu, Yue, et al.
Published: (2024)
by: Hu, Yue, et al.
Published: (2024)
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
by: Zhang, Lu, et al.
Published: (2024)
by: Zhang, Lu, et al.
Published: (2024)
DNA: Dual-branch Network with Adaptation for Open-Set Online Handwriting Generation
by: Huang, Tsai-Ling, et al.
Published: (2025)
by: Huang, Tsai-Ling, et al.
Published: (2025)
Mora: Enabling Generalist Video Generation via A Multi-Agent Framework
by: Yuan, Zhengqing, et al.
Published: (2024)
by: Yuan, Zhengqing, et al.
Published: (2024)
Towards Reliable Fetal Ultrasound Interpretation with Multi-Agent Collaboration
by: Hu, Xiaotian, et al.
Published: (2026)
by: Hu, Xiaotian, et al.
Published: (2026)
MA-CBP: A Criminal Behavior Prediction Framework Based on Multi-Agent Asynchronous Collaboration
by: Liu, Cheng, et al.
Published: (2025)
by: Liu, Cheng, et al.
Published: (2025)
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
by: Yang, Zhongyu, et al.
Published: (2026)
by: Yang, Zhongyu, et al.
Published: (2026)
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
by: Zeng, Qinglin, et al.
Published: (2025)
by: Zeng, Qinglin, et al.
Published: (2025)
LungNoduleAgent: A Collaborative Multi-Agent System for Precision Diagnosis of Lung Nodules
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition
by: Wang, Zeheng, et al.
Published: (2026)
by: Wang, Zeheng, et al.
Published: (2026)
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
by: Wang, Zheng, et al.
Published: (2026)
by: Wang, Zheng, et al.
Published: (2026)
Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition
by: Huang, Guanjie, et al.
Published: (2025)
by: Huang, Guanjie, et al.
Published: (2025)
A General Framework for Jersey Number Recognition in Sports Video
by: Koshkina, Maria, et al.
Published: (2024)
by: Koshkina, Maria, et al.
Published: (2024)
CompAgent: An Agentic Framework for Visual Compliance Verification
by: Ghosh, Rahul, et al.
Published: (2025)
by: Ghosh, Rahul, et al.
Published: (2025)
TranSPORTmer: A Holistic Approach to Trajectory Understanding in Multi-Agent Sports
by: Capellera, Guillem, et al.
Published: (2024)
by: Capellera, Guillem, et al.
Published: (2024)
MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
HighlightMe: Detecting Highlights from Human-Centric Videos
by: Bhattacharya, Uttaran, et al.
Published: (2021)
by: Bhattacharya, Uttaran, et al.
Published: (2021)
Similar Items
-
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
by: Huang, Kaiyi, et al.
Published: (2024) -
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
by: Zhou, Yiyang, et al.
Published: (2025) -
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
by: Liu, Haowei, et al.
Published: (2025) -
Long-Video Audio Synthesis with Multi-Agent Collaboration
by: Zhang, Yehang, et al.
Published: (2025) -
Sports-Traj: A Unified Trajectory Generation Model for Multi-Agent Movement in Sports
by: Xu, Yi, et al.
Published: (2024)