LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Hongjie, Dong, Lu, Liu, Yi, Huang, Yifei, Wang, Yali, Wang, Limin, Qiao, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
VideoMamba: State Space Model for Efficient Video Understanding
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining
von: Dong, Lu, et al.
Veröffentlicht: (2025)
von: Dong, Lu, et al.
Veröffentlicht: (2025)
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
von: Dong, Lu, et al.
Veröffentlicht: (2025)
von: Dong, Lu, et al.
Veröffentlicht: (2025)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
VideoChat: Chat-Centric Video Understanding
von: Li, KunChang, et al.
Veröffentlicht: (2023)
von: Li, KunChang, et al.
Veröffentlicht: (2023)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
von: Wang, Yi, et al.
Veröffentlicht: (2024)
von: Wang, Yi, et al.
Veröffentlicht: (2024)
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
von: Wang, Zining, et al.
Veröffentlicht: (2025)
von: Wang, Zining, et al.
Veröffentlicht: (2025)
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
von: Huang, Ziqi, et al.
Veröffentlicht: (2024)
von: Huang, Ziqi, et al.
Veröffentlicht: (2024)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
A Simple LLM Framework for Long-Range Video Question-Answering
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
Encoding and Controlling Global Semantics for Long-form Video Question Answering
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
von: He, Yuping, et al.
Veröffentlicht: (2025)
von: He, Yuping, et al.
Veröffentlicht: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
von: Xu, Yicheng, et al.
Veröffentlicht: (2025)
von: Xu, Yicheng, et al.
Veröffentlicht: (2025)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
von: Li, Xinhao, et al.
Veröffentlicht: (2025)
von: Li, Xinhao, et al.
Veröffentlicht: (2025)
Harvest Video Foundation Models via Efficient Post-Pretraining
von: Li, Yizhuo, et al.
Veröffentlicht: (2023)
von: Li, Yizhuo, et al.
Veröffentlicht: (2023)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving
von: Chen, Siran, et al.
Veröffentlicht: (2025)
von: Chen, Siran, et al.
Veröffentlicht: (2025)
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
von: Qian, Tianwen, et al.
Veröffentlicht: (2023)
von: Qian, Tianwen, et al.
Veröffentlicht: (2023)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
CinePile: A Long Video Question Answering Dataset and Benchmark
von: Rawal, Ruchit, et al.
Veröffentlicht: (2024)
von: Rawal, Ruchit, et al.
Veröffentlicht: (2024)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
von: Song, Enxin, et al.
Veröffentlicht: (2024)
von: Song, Enxin, et al.
Veröffentlicht: (2024)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024) -
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
von: Li, Kunchang, et al.
Veröffentlicht: (2023) -
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025) -
VideoMamba: State Space Model for Efficient Video Understanding
von: Li, Kunchang, et al.
Veröffentlicht: (2024) -
Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining
von: Dong, Lu, et al.
Veröffentlicht: (2025)