SCBench: A Sports Commentary Benchmark for Video LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ge, Kuangzhi, Chen, Lingjun, Zhang, Kevin, Luo, Yulin, Shi, Tianyu, Fan, Liaoyuan, Li, Xiang, Wang, Guanqun, Zhang, Shanghang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Can World Models Benefit VLMs for World Dynamics?
par: Zhang, Kevin, et autres
Publié: (2025)
par: Zhang, Kevin, et autres
Publié: (2025)
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
par: Mak, Chak-Wing, et autres
Publié: (2026)
par: Mak, Chak-Wing, et autres
Publié: (2026)
Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test
par: Fan, Chun-Kai, et autres
Publié: (2026)
par: Fan, Chun-Kai, et autres
Publié: (2026)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
par: Zhang, Yulin, et autres
Publié: (2026)
par: Zhang, Yulin, et autres
Publié: (2026)
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
par: Yang, Junqi, et autres
Publié: (2026)
par: Yang, Junqi, et autres
Publié: (2026)
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
par: Kou, Qian, et autres
Publié: (2026)
par: Kou, Qian, et autres
Publié: (2026)
MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception
par: Wang, Guanqun, et autres
Publié: (2024)
par: Wang, Guanqun, et autres
Publié: (2024)
DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning
par: Zou, Junbo, et autres
Publié: (2025)
par: Zou, Junbo, et autres
Publié: (2025)
A Physical Coherence Benchmark for Evaluating Video Generation Models via Optical Flow-guided Frame Prediction
par: Chen, Yongfan, et autres
Publié: (2025)
par: Chen, Yongfan, et autres
Publié: (2025)
AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
par: Zha, Jirong, et autres
Publié: (2025)
par: Zha, Jirong, et autres
Publié: (2025)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
par: An, Ruichuan, et autres
Publié: (2025)
par: An, Ruichuan, et autres
Publié: (2025)
TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
par: Li, Baiqi, et autres
Publié: (2026)
par: Li, Baiqi, et autres
Publié: (2026)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
par: Gong, Zhantao, et autres
Publié: (2025)
par: Gong, Zhantao, et autres
Publié: (2025)
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
par: Huang, Yiyang, et autres
Publié: (2026)
par: Huang, Yiyang, et autres
Publié: (2026)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
par: Pan, Yaning, et autres
Publié: (2025)
par: Pan, Yaning, et autres
Publié: (2025)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
par: An, Ruichuan, et autres
Publié: (2024)
par: An, Ruichuan, et autres
Publié: (2024)
FastInit: Fast Noise Initialization for Temporally Consistent Video Generation
par: Bai, Chengyu, et autres
Publié: (2025)
par: Bai, Chengyu, et autres
Publié: (2025)
VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
par: Jin, Hongbo, et autres
Publié: (2025)
par: Jin, Hongbo, et autres
Publié: (2025)
Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation
par: Zhang, Rongyu, et autres
Publié: (2024)
par: Zhang, Rongyu, et autres
Publié: (2024)
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
par: Zou, Heqing, et autres
Publié: (2025)
par: Zou, Heqing, et autres
Publié: (2025)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
par: Zhou, Wenrui, et autres
Publié: (2025)
par: Zhou, Wenrui, et autres
Publié: (2025)
High-Quality 3D Creation from A Single Image Using Subject-Specific Knowledge Prior
par: Huang, Nan, et autres
Publié: (2023)
par: Huang, Nan, et autres
Publié: (2023)
SportsGPT: An LLM-driven Framework for Interpretable Sports Motion Assessment and Training Guidance
par: Tian, Wenbo, et autres
Publié: (2025)
par: Tian, Wenbo, et autres
Publié: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
par: Li, Yunxin, et autres
Publié: (2024)
par: Li, Yunxin, et autres
Publié: (2024)
Video-Bench: Human-Aligned Video Generation Benchmark
par: Han, Hui, et autres
Publié: (2025)
par: Han, Hui, et autres
Publié: (2025)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
par: Lin, Kuanwei, et autres
Publié: (2026)
par: Lin, Kuanwei, et autres
Publié: (2026)
PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm
par: Chang, Tianyu, et autres
Publié: (2024)
par: Chang, Tianyu, et autres
Publié: (2024)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
par: Cai, Yuxuan, et autres
Publié: (2025)
par: Cai, Yuxuan, et autres
Publié: (2025)
SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
par: Xia, Haotian, et autres
Publié: (2025)
par: Xia, Haotian, et autres
Publié: (2025)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
par: Li, Chenglin, et autres
Publié: (2024)
par: Li, Chenglin, et autres
Publié: (2024)
Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly
par: Du, Hang, et autres
Publié: (2024)
par: Du, Hang, et autres
Publié: (2024)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
par: Chen, Houlun, et autres
Publié: (2024)
par: Chen, Houlun, et autres
Publié: (2024)
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
par: Wu, Qi, et autres
Publié: (2025)
par: Wu, Qi, et autres
Publié: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
par: Zhang, Jun, et autres
Publié: (2025)
par: Zhang, Jun, et autres
Publié: (2025)
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
par: Chen, Siqi, et autres
Publié: (2025)
par: Chen, Siqi, et autres
Publié: (2025)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
par: Li, Chenglin, et autres
Publié: (2026)
par: Li, Chenglin, et autres
Publié: (2026)
WM-MoE: Weather-aware Multi-scale Mixture-of-Experts for Blind Adverse Weather Removal
par: Luo, Yulin, et autres
Publié: (2023)
par: Luo, Yulin, et autres
Publié: (2023)
MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs
par: Ma, Junpeng, et autres
Publié: (2025)
par: Ma, Junpeng, et autres
Publié: (2025)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
par: Li, Chenxuan, et autres
Publié: (2024)
par: Li, Chenxuan, et autres
Publié: (2024)
Documents similaires
-
Can World Models Benefit VLMs for World Dynamics?
par: Zhang, Kevin, et autres
Publié: (2025) -
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
par: Mak, Chak-Wing, et autres
Publié: (2026) -
Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test
par: Fan, Chun-Kai, et autres
Publié: (2026) -
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
par: Zhang, Yulin, et autres
Publié: (2026) -
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
par: Yang, Junqi, et autres
Publié: (2026)