VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Baoyu, Su, Qile, Zhu, Shoutai, Liang, Yuchen, Tong, Chao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
von: Su, Qile, et al.
Veröffentlicht: (2025)
von: Su, Qile, et al.
Veröffentlicht: (2025)
A Survey of Video Datasets for Grounded Event Understanding
von: Sanders, Kate, et al.
Veröffentlicht: (2024)
von: Sanders, Kate, et al.
Veröffentlicht: (2024)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
von: Li, Qi, et al.
Veröffentlicht: (2025)
von: Li, Qi, et al.
Veröffentlicht: (2025)
Semantic Event Graphs for Long-Form Video Question Answering
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026)
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026)
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026)
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026)
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
von: He, Zhihao, et al.
Veröffentlicht: (2026)
von: He, Zhihao, et al.
Veröffentlicht: (2026)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
von: Tang, Yolo Y., et al.
Veröffentlicht: (2024)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2024)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
EventVL: Understand Event Streams via Multimodal Large Language Model
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
von: Feng, Weixi, et al.
Veröffentlicht: (2025)
von: Feng, Weixi, et al.
Veröffentlicht: (2025)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
von: Fang, Ye, et al.
Veröffentlicht: (2025)
von: Fang, Ye, et al.
Veröffentlicht: (2025)
SafeVid: Toward Safety Aligned Video Large Multimodal Models
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
Enhancing Long Video Understanding via Hierarchical Event-Based Memory
von: Cheng, Dingxin, et al.
Veröffentlicht: (2024)
von: Cheng, Dingxin, et al.
Veröffentlicht: (2024)
VERHallu: Evaluating and Mitigating Event Relation Hallucination in Video Large Language Models
von: Zhang, Zefan, et al.
Veröffentlicht: (2026)
von: Zhang, Zefan, et al.
Veröffentlicht: (2026)
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
von: Zhang, Runze, et al.
Veröffentlicht: (2025)
von: Zhang, Runze, et al.
Veröffentlicht: (2025)
AdaVid: Adaptive Video-Language Pretraining
von: Patel, Chaitanya, et al.
Veröffentlicht: (2025)
von: Patel, Chaitanya, et al.
Veröffentlicht: (2025)
When and Where do Events Switch in Multi-Event Video Generation?
von: Liao, Ruotong, et al.
Veröffentlicht: (2025)
von: Liao, Ruotong, et al.
Veröffentlicht: (2025)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
von: Cheng, Sijie, et al.
Veröffentlicht: (2024)
von: Cheng, Sijie, et al.
Veröffentlicht: (2024)
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios
von: Yan, Peizheng, et al.
Veröffentlicht: (2026)
von: Yan, Peizheng, et al.
Veröffentlicht: (2026)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
VidTwin: Video VAE with Decoupled Structure and Dynamics
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
von: Kang, Minseok, et al.
Veröffentlicht: (2026)
von: Kang, Minseok, et al.
Veröffentlicht: (2026)
VidDoS: Universal Denial-of-Service Attack on Video-based Large Language Models
von: Tang, Duoxun, et al.
Veröffentlicht: (2026)
von: Tang, Duoxun, et al.
Veröffentlicht: (2026)
Uneven Event Modeling for Partially Relevant Video Retrieval
von: Zhu, Sa, et al.
Veröffentlicht: (2025)
von: Zhu, Sa, et al.
Veröffentlicht: (2025)
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2022)
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2022)
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
Localizing Events in Videos with Multimodal Queries
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2024)
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline
von: Jia, Qi, et al.
Veröffentlicht: (2024)
von: Jia, Qi, et al.
Veröffentlicht: (2024)
MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction
von: Tang, Shuo, et al.
Veröffentlicht: (2025)
von: Tang, Shuo, et al.
Veröffentlicht: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
Fostering Video Reasoning via Next-Event Prediction
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
von: Lin, Rui, et al.
Veröffentlicht: (2026)
von: Lin, Rui, et al.
Veröffentlicht: (2026)
Event-Enhanced Blurry Video Super-Resolution
von: Kai, Dachun, et al.
Veröffentlicht: (2025)
von: Kai, Dachun, et al.
Veröffentlicht: (2025)
EdgeVidSum: Real-Time Personalized Video Summarization at the Edge
von: Mujtaba, Ghulam, et al.
Veröffentlicht: (2025)
von: Mujtaba, Ghulam, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
von: Su, Qile, et al.
Veröffentlicht: (2025) -
A Survey of Video Datasets for Grounded Event Understanding
von: Sanders, Kate, et al.
Veröffentlicht: (2024) -
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
von: Li, Qi, et al.
Veröffentlicht: (2025) -
Semantic Event Graphs for Long-Form Video Question Answering
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026) -
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026)