Saved in:
| Main Authors: | Yang, Zhongyu, Yang, Zuhao, Zhan, Shuo, Yue, Tan, Pang, Wei, Yuan, Yingfang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.05079 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration
by: Yang, Zhongyu, et al.
Published: (2025)
by: Yang, Zhongyu, et al.
Published: (2025)
Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
by: Yang, Zhongyu, et al.
Published: (2025)
by: Yang, Zhongyu, et al.
Published: (2025)
XR: Cross-Modal Agents for Composed Image Retrieval
by: Yang, Zhongyu, et al.
Published: (2026)
by: Yang, Zhongyu, et al.
Published: (2026)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
by: Zhao, Zixu, et al.
Published: (2025)
by: Zhao, Zixu, et al.
Published: (2025)
Event-based Video Person Re-identification via Cross-Modality and Temporal Collaboration
by: Li, Renkai, et al.
Published: (2025)
by: Li, Renkai, et al.
Published: (2025)
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
by: Xu, Dannong, et al.
Published: (2026)
by: Xu, Dannong, et al.
Published: (2026)
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
by: Li, Jiahua, et al.
Published: (2025)
by: Li, Jiahua, et al.
Published: (2025)
MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights
by: Yang, Jasmine, et al.
Published: (2026)
by: Yang, Jasmine, et al.
Published: (2026)
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
by: Chen, Kerui, et al.
Published: (2026)
by: Chen, Kerui, et al.
Published: (2026)
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026)
by: Yang, Zuhao, et al.
Published: (2026)
Long-Video Audio Synthesis with Multi-Agent Collaboration
by: Zhang, Yehang, et al.
Published: (2025)
by: Zhang, Yehang, et al.
Published: (2025)
DATE: Dynamic Absolute Time Enhancement for Long Video Understanding
by: Yuan, Chao, et al.
Published: (2025)
by: Yuan, Chao, et al.
Published: (2025)
VCA: Video Curious Agent for Long Video Understanding
by: Yang, Zeyuan, et al.
Published: (2024)
by: Yang, Zeyuan, et al.
Published: (2024)
MR. Video: "MapReduce" is the Principle for Long Video Understanding
by: Pang, Ziqi, et al.
Published: (2025)
by: Pang, Ziqi, et al.
Published: (2025)
EEA: Exploration-Exploitation Agent for Long Video Understanding
by: Yang, Te, et al.
Published: (2025)
by: Yang, Te, et al.
Published: (2025)
Communication-Efficient Multi-Agent 3D Detection via Hybrid Collaboration
by: Hu, Yue, et al.
Published: (2025)
by: Hu, Yue, et al.
Published: (2025)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
by: Yin, Yufei, et al.
Published: (2026)
by: Yin, Yufei, et al.
Published: (2026)
Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
Versatile Transition Generation with Image-to-Video Diffusion
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Pragmatic Communication in Multi-Agent Collaborative Perception
by: Hu, Yue, et al.
Published: (2024)
by: Hu, Yue, et al.
Published: (2024)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
by: Ding, Yanbo, et al.
Published: (2024)
by: Ding, Yanbo, et al.
Published: (2024)
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction
by: Man, Yuanbin, et al.
Published: (2024)
by: Man, Yuanbin, et al.
Published: (2024)
Facilitating Video Story Interaction with Multi-Agent Collaborative System
by: Zhang, Yiwen, et al.
Published: (2025)
by: Zhang, Yiwen, et al.
Published: (2025)
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
by: Feng, Yue, et al.
Published: (2025)
by: Feng, Yue, et al.
Published: (2025)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
by: Xie, Yuan, et al.
Published: (2025)
by: Xie, Yuan, et al.
Published: (2025)
BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
by: Wang, Zheng, et al.
Published: (2026)
by: Wang, Zheng, et al.
Published: (2026)
MambaSOD: Dual Mamba-Driven Cross-Modal Fusion Network for RGB-D Salient Object Detection
by: Zhan, Yue, et al.
Published: (2024)
by: Zhan, Yue, et al.
Published: (2024)
SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
by: Sheng, Yuan, et al.
Published: (2025)
by: Sheng, Yuan, et al.
Published: (2025)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
by: Wang, Mengyue, et al.
Published: (2025)
by: Wang, Mengyue, et al.
Published: (2025)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
by: Ma, David, et al.
Published: (2025)
by: Ma, David, et al.
Published: (2025)
LongVideoAgent: Multi-Agent Reasoning with Long Videos
by: Liu, Runtao, et al.
Published: (2025)
by: Liu, Runtao, et al.
Published: (2025)
Similar Items
-
InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration
by: Yang, Zhongyu, et al.
Published: (2025) -
Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
by: Yang, Zhongyu, et al.
Published: (2025) -
XR: Cross-Modal Agents for Composed Image Retrieval
by: Yang, Zhongyu, et al.
Published: (2026) -
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025) -
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
by: Wei, Zheng, et al.
Published: (2025)