ALLVB: All-in-One Long Video Understanding Benchmark
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tan, Xichen, Luo, Yuanjing, Ye, Yunfan, Liu, Fang, Cai, Zhiping |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding
par: Tan, Xichen, et autres
Publié: (2025)
par: Tan, Xichen, et autres
Publié: (2025)
HOCA-Bench: Beyond Semantic Perception to Predictive World Modeling via Hegelian Ontological-Causal Anomalies
par: Liu, Chang, et autres
Publié: (2026)
par: Liu, Chang, et autres
Publié: (2026)
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
par: Liu, Chang, et autres
Publié: (2025)
par: Liu, Chang, et autres
Publié: (2025)
StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References
par: He, Boyu, et autres
Publié: (2026)
par: He, Boyu, et autres
Publié: (2026)
DiffusionEdge: Diffusion Probabilistic Model for Crisp Edge Detection
par: Ye, Yunfan, et autres
Publié: (2024)
par: Ye, Yunfan, et autres
Publié: (2024)
From Physical Degradation Models to Task-Aware All-in-One Image Restoration
par: Gao, Hu, et autres
Publié: (2026)
par: Gao, Hu, et autres
Publié: (2026)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
par: Zhong, Yangyang, et autres
Publié: (2025)
par: Zhong, Yangyang, et autres
Publié: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
par: Fang, Xinyu, et autres
Publié: (2024)
par: Fang, Xinyu, et autres
Publié: (2024)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
par: Ma, David, et autres
Publié: (2025)
par: Ma, David, et autres
Publié: (2025)
Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation
par: Liang, Tianming, et autres
Publié: (2025)
par: Liang, Tianming, et autres
Publié: (2025)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
par: Zhou, Wenqi, et autres
Publié: (2025)
par: Zhou, Wenqi, et autres
Publié: (2025)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
par: Qiu, Jihao, et autres
Publié: (2026)
par: Qiu, Jihao, et autres
Publié: (2026)
VACE: All-in-One Video Creation and Editing
par: Jiang, Zeyinzi, et autres
Publié: (2025)
par: Jiang, Zeyinzi, et autres
Publié: (2025)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
par: Chen, Guo, et autres
Publié: (2024)
par: Chen, Guo, et autres
Publié: (2024)
Cambrian-P: Pose-Grounded Video Understanding
par: Yang, Jihan, et autres
Publié: (2026)
par: Yang, Jihan, et autres
Publié: (2026)
MoA-VR: A Mixture-of-Agents System Towards All-in-One Video Restoration
par: Liu, Lu, et autres
Publié: (2025)
par: Liu, Lu, et autres
Publié: (2025)
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
par: He, Haichen, et autres
Publié: (2026)
par: He, Haichen, et autres
Publié: (2026)
DrVideo: Document Retrieval Based Long Video Understanding
par: Ma, Ziyu, et autres
Publié: (2024)
par: Ma, Ziyu, et autres
Publié: (2024)
LVBench: An Extreme Long Video Understanding Benchmark
par: Wang, Weihan, et autres
Publié: (2024)
par: Wang, Weihan, et autres
Publié: (2024)
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
par: Chen, Tao, et autres
Publié: (2025)
par: Chen, Tao, et autres
Publié: (2025)
OneThinker: All-in-one Reasoning Model for Image and Video
par: Feng, Kaituo, et autres
Publié: (2025)
par: Feng, Kaituo, et autres
Publié: (2025)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
par: Fang, Pengcheng, et autres
Publié: (2025)
par: Fang, Pengcheng, et autres
Publié: (2025)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
par: Xue, Zhucun, et autres
Publié: (2025)
par: Xue, Zhucun, et autres
Publié: (2025)
MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
par: Tan, Wenhui, et autres
Publié: (2026)
par: Tan, Wenhui, et autres
Publié: (2026)
Towards Long Video Understanding via Fine-detailed Video Story Generation
par: You, Zeng, et autres
Publié: (2024)
par: You, Zeng, et autres
Publié: (2024)
MLVU: Benchmarking Multi-task Long Video Understanding
par: Zhou, Junjie, et autres
Publié: (2024)
par: Zhou, Junjie, et autres
Publié: (2024)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
par: Lin, Jingyang, et autres
Publié: (2025)
par: Lin, Jingyang, et autres
Publié: (2025)
LongDiff: Training-Free Long Video Generation in One Go
par: Li, Zhuoling, et autres
Publié: (2025)
par: Li, Zhuoling, et autres
Publié: (2025)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
par: Feng, Xiang, et autres
Publié: (2026)
par: Feng, Xiang, et autres
Publié: (2026)
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
par: Zou, Heqing, et autres
Publié: (2025)
par: Zou, Heqing, et autres
Publié: (2025)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
par: Zhang, Hongjie, et autres
Publié: (2023)
par: Zhang, Hongjie, et autres
Publié: (2023)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
par: Li, Kunchang, et autres
Publié: (2023)
par: Li, Kunchang, et autres
Publié: (2023)
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
par: Zhang, Zheyu, et autres
Publié: (2026)
par: Zhang, Zheyu, et autres
Publié: (2026)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
par: Rahman, Tanzila, et autres
Publié: (2026)
par: Rahman, Tanzila, et autres
Publié: (2026)
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
par: Chen, Seng Nam, et autres
Publié: (2026)
par: Chen, Seng Nam, et autres
Publié: (2026)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
par: Wu, Haoning, et autres
Publié: (2024)
par: Wu, Haoning, et autres
Publié: (2024)
Hallucination Mitigation Prompts Long-term Video Understanding
par: Sun, Yiwei, et autres
Publié: (2024)
par: Sun, Yiwei, et autres
Publié: (2024)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
par: Zhang, Zicheng, et autres
Publié: (2024)
par: Zhang, Zicheng, et autres
Publié: (2024)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
par: Zuo, Jialong, et autres
Publié: (2025)
par: Zuo, Jialong, et autres
Publié: (2025)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
par: Nagrani, Arsha, et autres
Publié: (2024)
par: Nagrani, Arsha, et autres
Publié: (2024)
Documents similaires
-
RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding
par: Tan, Xichen, et autres
Publié: (2025) -
HOCA-Bench: Beyond Semantic Perception to Predictive World Modeling via Hegelian Ontological-Causal Anomalies
par: Liu, Chang, et autres
Publié: (2026) -
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
par: Liu, Chang, et autres
Publié: (2025) -
StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References
par: He, Boyu, et autres
Publié: (2026) -
DiffusionEdge: Diffusion Probabilistic Model for Crisp Edge Detection
par: Ye, Yunfan, et autres
Publié: (2024)