Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yun, Zhang, Long, Liu, Jingren, Yan, Jiaqi, Zhang, Zhanjie, Zheng, Jiahao, Ma, Ao, Ling, Run, Yang, Xun, Wu, Dapeng, Chen, Xiangyu, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Infinite Video Understanding
von: Zhang, Dell, et al.
Veröffentlicht: (2025)
von: Zhang, Dell, et al.
Veröffentlicht: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Generative Video Compression: Towards 0.01% Compression Rate for Video Transmission
von: Chen, Xiangyu, et al.
Veröffentlicht: (2025)
von: Chen, Xiangyu, et al.
Veröffentlicht: (2025)
Towards Event-oriented Long Video Understanding
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
EV-NVC: Efficient Variable bitrate Neural Video Compression
von: Hu, Yongcun, et al.
Veröffentlicht: (2025)
von: Hu, Yongcun, et al.
Veröffentlicht: (2025)
HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
von: Zhang, Long, et al.
Veröffentlicht: (2025)
von: Zhang, Long, et al.
Veröffentlicht: (2025)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models
von: Lin, Zixing, et al.
Veröffentlicht: (2026)
von: Lin, Zixing, et al.
Veröffentlicht: (2026)
Multimodal Chaptering for Long-Form TV Newscast Video
von: Guetari, Khalil, et al.
Veröffentlicht: (2024)
von: Guetari, Khalil, et al.
Veröffentlicht: (2024)
Shorts on the Rise: Assessing the Effects of YouTube Shorts on Long-Form Video Content
von: Rajendran, Prajit T., et al.
Veröffentlicht: (2024)
von: Rajendran, Prajit T., et al.
Veröffentlicht: (2024)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
von: You, Qijie, et al.
Veröffentlicht: (2026)
von: You, Qijie, et al.
Veröffentlicht: (2026)
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
Short-Form Video Viewing Behavior Analysis and Multi-Step Viewing Time Prediction
von: Yen, Vu Thi Hai, et al.
Veröffentlicht: (2026)
von: Yen, Vu Thi Hai, et al.
Veröffentlicht: (2026)
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
von: Tian, Chong, et al.
Veröffentlicht: (2026)
von: Tian, Chong, et al.
Veröffentlicht: (2026)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
Predicting Outcomes in Video Games with Long Short Term Memory Networks
von: Chulajata, Kittimate, et al.
Veröffentlicht: (2024)
von: Chulajata, Kittimate, et al.
Veröffentlicht: (2024)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
FedCVU: Federated Learning for Cross-View Video Understanding
von: Zhang, Shenghan, et al.
Veröffentlicht: (2026)
von: Zhang, Shenghan, et al.
Veröffentlicht: (2026)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
Towards Open-Vocabulary Video Semantic Segmentation
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
von: Li, Jie, et al.
Veröffentlicht: (2025)
von: Li, Jie, et al.
Veröffentlicht: (2025)
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
von: Hao, Haoran, et al.
Veröffentlicht: (2025)
von: Hao, Haoran, et al.
Veröffentlicht: (2025)
PRVR: Partially Relevant Video Retrieval
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
von: He, Xusheng, et al.
Veröffentlicht: (2025)
von: He, Xusheng, et al.
Veröffentlicht: (2025)
Anchorage: Visual Analysis of Satisfaction in Customer Service Videos via Anchor Events
von: Wong, Kam Kwai, et al.
Veröffentlicht: (2023)
von: Wong, Kam Kwai, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Infinite Video Understanding
von: Zhang, Dell, et al.
Veröffentlicht: (2025) -
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024) -
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025) -
Generative Video Compression: Towards 0.01% Compression Rate for Video Transmission
von: Chen, Xiangyu, et al.
Veröffentlicht: (2025) -
Towards Event-oriented Long Video Understanding
von: Du, Yifan, et al.
Veröffentlicht: (2024)