EEA: Exploration-Exploitation Agent for Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Te, Zhu, Xiangyu, Wang, Bo, Chen, Quan, Jiang, Peng, Lei, Zhen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
VCA: Video Curious Agent for Long Video Understanding
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
von: Yang, Te, et al.
Veröffentlicht: (2024)
von: Yang, Te, et al.
Veröffentlicht: (2024)
Text-Video Multi-Grained Integration for Video Moment Montage
von: Yin, Zhihui, et al.
Veröffentlicht: (2024)
von: Yin, Zhihui, et al.
Veröffentlicht: (2024)
Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
AffectSeek: Agentic Affective Understanding in Long Videos under Vague User Queries
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
S2TD-Face: Reconstruct a Detailed 3D Face with Controllable Texture from a Single Sketch
von: Wang, Zidu, et al.
Veröffentlicht: (2024)
von: Wang, Zidu, et al.
Veröffentlicht: (2024)
LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
von: Li, Handong, et al.
Veröffentlicht: (2026)
von: Li, Handong, et al.
Veröffentlicht: (2026)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
von: Yang, Zhongyu, et al.
Veröffentlicht: (2026)
von: Yang, Zhongyu, et al.
Veröffentlicht: (2026)
Training-free Subject-Enhanced Attention Guidance for Compositional Text-to-image Generation
von: Liu, Shengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Shengyuan, et al.
Veröffentlicht: (2024)
STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstruction
von: Zhao, Jiankuo, et al.
Veröffentlicht: (2025)
von: Zhao, Jiankuo, et al.
Veröffentlicht: (2025)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
DevFD: Developmental Face Forgery Detection by Learning Shared and Orthogonal LoRA Subspaces
von: Zhang, Tianshuo, et al.
Veröffentlicht: (2025)
von: Zhang, Tianshuo, et al.
Veröffentlicht: (2025)
Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
von: Yang, Pinci, et al.
Veröffentlicht: (2025)
von: Yang, Pinci, et al.
Veröffentlicht: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching
von: Liu, Jingyu, et al.
Veröffentlicht: (2024)
von: Liu, Jingyu, et al.
Veröffentlicht: (2024)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
von: Ren, Weiming, et al.
Veröffentlicht: (2024)
von: Ren, Weiming, et al.
Veröffentlicht: (2024)
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
von: Yan, Haiyang, et al.
Veröffentlicht: (2026)
von: Yan, Haiyang, et al.
Veröffentlicht: (2026)
MR. Video: "MapReduce" is the Principle for Long Video Understanding
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
Improving Large Vision-Language Models' Understanding for Flow Field Data
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
DATE: Dynamic Absolute Time Enhancement for Long Video Understanding
von: Yuan, Chao, et al.
Veröffentlicht: (2025)
von: Yuan, Chao, et al.
Veröffentlicht: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning
von: Chen, Cheng, et al.
Veröffentlicht: (2025)
von: Chen, Cheng, et al.
Veröffentlicht: (2025)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
von: Li, Guangyuan, et al.
Veröffentlicht: (2025) -
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
von: Chen, Boyu, et al.
Veröffentlicht: (2025) -
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026) -
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026) -
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)