LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Xiaodong, Huang, Langling, Wu, Zhirong, Zhao, Xu, Xu, Teng, Xia, Xuhong, Peng, Peixi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model
par: Wang, Xiaodong, et autres
Publié: (2025)
par: Wang, Xiaodong, et autres
Publié: (2025)
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
par: Tao, Keda, et autres
Publié: (2026)
par: Tao, Keda, et autres
Publié: (2026)
ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
par: Wang, Xiaodong, et autres
Publié: (2025)
par: Wang, Xiaodong, et autres
Publié: (2025)
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
par: Wang, Xiaodong, et autres
Publié: (2025)
par: Wang, Xiaodong, et autres
Publié: (2025)
Active Perception Agent for Omnimodal Audio-Video Understanding
par: Tao, Keda, et autres
Publié: (2025)
par: Tao, Keda, et autres
Publié: (2025)
ViStoryBench: Comprehensive Benchmark Suite for Story Visualization
par: Zhuang, Cailin, et autres
Publié: (2025)
par: Zhuang, Cailin, et autres
Publié: (2025)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
par: Chen, Harold Haodong, et autres
Publié: (2025)
par: Chen, Harold Haodong, et autres
Publié: (2025)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
par: Zhang, Zicheng, et autres
Publié: (2024)
par: Zhang, Zicheng, et autres
Publié: (2024)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
par: Chen, Guo, et autres
Publié: (2024)
par: Chen, Guo, et autres
Publié: (2024)
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
par: Liang, Yongyuan, et autres
Publié: (2025)
par: Liang, Yongyuan, et autres
Publié: (2025)
ViMU: Benchmarking Video Metaphorical Understanding
par: Li, Qi, et autres
Publié: (2026)
par: Li, Qi, et autres
Publié: (2026)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
par: He, Yuping, et autres
Publié: (2025)
par: He, Yuping, et autres
Publié: (2025)
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
par: Oh, Yeongtak, et autres
Publié: (2026)
par: Oh, Yeongtak, et autres
Publié: (2026)
InstructionBench: An Instructional Video Understanding Benchmark
par: Wei, Haiwan, et autres
Publié: (2025)
par: Wei, Haiwan, et autres
Publié: (2025)
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
par: Hu, Lanxiang, et autres
Publié: (2025)
par: Hu, Lanxiang, et autres
Publié: (2025)
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
par: Kong, Fanqi, et autres
Publié: (2025)
par: Kong, Fanqi, et autres
Publié: (2025)
Prototype Embedding Optimization for Human-Object Interaction Detection in Livestreaming
par: Zhang, Menghui, et autres
Publié: (2025)
par: Zhang, Menghui, et autres
Publié: (2025)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
par: Wang, Xijun, et autres
Publié: (2025)
par: Wang, Xijun, et autres
Publié: (2025)
FreeGen: Feed-Forward Reconstruction-Generation Co-Training for Free-Viewpoint Driving Scene Synthesis
par: Chen, Shijie, et autres
Publié: (2025)
par: Chen, Shijie, et autres
Publié: (2025)
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
par: Wang, Xinran, et autres
Publié: (2025)
par: Wang, Xinran, et autres
Publié: (2025)
ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
par: Liberatori, Benedetta, et autres
Publié: (2025)
par: Liberatori, Benedetta, et autres
Publié: (2025)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
par: Hong, Wenyi, et autres
Publié: (2025)
par: Hong, Wenyi, et autres
Publié: (2025)
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
par: Tu, Chongjun, et autres
Publié: (2025)
par: Tu, Chongjun, et autres
Publié: (2025)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
par: Zhong, Yangyang, et autres
Publié: (2025)
par: Zhong, Yangyang, et autres
Publié: (2025)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
par: Zhou, Wenqi, et autres
Publié: (2025)
par: Zhou, Wenqi, et autres
Publié: (2025)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
par: Zhang, Zhihong, et autres
Publié: (2025)
par: Zhang, Zhihong, et autres
Publié: (2025)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
par: Zhang, Hongjie, et autres
Publié: (2023)
par: Zhang, Hongjie, et autres
Publié: (2023)
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
par: Hong, Jack, et autres
Publié: (2025)
par: Hong, Jack, et autres
Publié: (2025)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
par: Liu, Xuannan, et autres
Publié: (2025)
par: Liu, Xuannan, et autres
Publié: (2025)
ViLCo-Bench: VIdeo Language COntinual learning Benchmark
par: Tang, Tianqi, et autres
Publié: (2024)
par: Tang, Tianqi, et autres
Publié: (2024)
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
par: Zhang, Shi-Xue, et autres
Publié: (2025)
par: Zhang, Shi-Xue, et autres
Publié: (2025)
ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
par: Wu, Xuecheng, et autres
Publié: (2025)
par: Wu, Xuecheng, et autres
Publié: (2025)
LongViTU: Instruction Tuning for Long-Form Video Understanding
par: Wu, Rujie, et autres
Publié: (2025)
par: Wu, Rujie, et autres
Publié: (2025)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
par: Pan, Yaning, et autres
Publié: (2025)
par: Pan, Yaning, et autres
Publié: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
par: Xu, Yicheng, et autres
Publié: (2025)
par: Xu, Yicheng, et autres
Publié: (2025)
U-Bench: A Comprehensive Understanding of U-Net through 100-Variant Benchmarking
par: Tang, Fenghe, et autres
Publié: (2025)
par: Tang, Fenghe, et autres
Publié: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
par: Wu, Haoning, et autres
Publié: (2024)
par: Wu, Haoning, et autres
Publié: (2024)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
par: Zhong, Hao, et autres
Publié: (2025)
par: Zhong, Hao, et autres
Publié: (2025)
PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
par: Lu, Xudong, et autres
Publié: (2026)
par: Lu, Xudong, et autres
Publié: (2026)
Scaling Language-Centric Omnimodal Representation Learning
par: Xiao, Chenghao, et autres
Publié: (2025)
par: Xiao, Chenghao, et autres
Publié: (2025)
Documents similaires
-
LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model
par: Wang, Xiaodong, et autres
Publié: (2025) -
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
par: Tao, Keda, et autres
Publié: (2026) -
ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
par: Wang, Xiaodong, et autres
Publié: (2025) -
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
par: Wang, Xiaodong, et autres
Publié: (2025) -
Active Perception Agent for Omnimodal Audio-Video Understanding
par: Tao, Keda, et autres
Publié: (2025)