TempCompass: Do Video LLMs Really Understand Videos?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Yuanxin, Li, Shicheng, Liu, Yi, Wang, Yuxiang, Ren, Shuhuai, Li, Lei, Chen, Sishuo, Sun, Xu, Hou, Lu |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
par: Li, Shicheng, et autres
Publié: (2023)
par: Li, Shicheng, et autres
Publié: (2023)
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
par: Chen, Sishuo, et autres
Publié: (2024)
par: Chen, Sishuo, et autres
Publié: (2024)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
par: Ren, Shuhuai, et autres
Publié: (2023)
par: Ren, Shuhuai, et autres
Publié: (2023)
TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
par: Li, Shicheng, et autres
Publié: (2025)
par: Li, Shicheng, et autres
Publié: (2025)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
par: Liu, Yuanxin, et autres
Publié: (2025)
par: Liu, Yuanxin, et autres
Publié: (2025)
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
par: Yao, Linli, et autres
Publié: (2024)
par: Yao, Linli, et autres
Publié: (2024)
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
par: Yao, Linli, et autres
Publié: (2025)
par: Yao, Linli, et autres
Publié: (2025)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
par: Liu, Wenqi, et autres
Publié: (2026)
par: Liu, Wenqi, et autres
Publié: (2026)
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
par: Li, Yunheng, et autres
Publié: (2025)
par: Li, Yunheng, et autres
Publié: (2025)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
par: Ren, Shuhuai, et autres
Publié: (2025)
par: Ren, Shuhuai, et autres
Publié: (2025)
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
par: Shen, Leqi, et autres
Publié: (2024)
par: Shen, Leqi, et autres
Publié: (2024)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
par: Liu, Yuanxin, et autres
Publié: (2025)
par: Liu, Yuanxin, et autres
Publié: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
par: Ouyang, Kun, et autres
Publié: (2025)
par: Ouyang, Kun, et autres
Publié: (2025)
Temporal Reasoning Transfer from Text to Video
par: Li, Lei, et autres
Publié: (2024)
par: Li, Lei, et autres
Publié: (2024)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
par: Jiang, Jindong, et autres
Publié: (2025)
par: Jiang, Jindong, et autres
Publié: (2025)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
par: Liu, Dongyang, et autres
Publié: (2025)
par: Liu, Dongyang, et autres
Publié: (2025)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
par: Fu, Chaoyou, et autres
Publié: (2024)
par: Fu, Chaoyou, et autres
Publié: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
par: Ouyang, Kun, et autres
Publié: (2024)
par: Ouyang, Kun, et autres
Publié: (2024)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
par: Liu, Yunze, et autres
Publié: (2024)
par: Liu, Yunze, et autres
Publié: (2024)
AURA: Always-On Understanding and Real-Time Assistance via Video Streams
par: Lu, Xudong, et autres
Publié: (2026)
par: Lu, Xudong, et autres
Publié: (2026)
Do Vision-Language Models Really Understand Visual Language?
par: Hou, Yifan, et autres
Publié: (2024)
par: Hou, Yifan, et autres
Publié: (2024)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
par: Sun, Yiming, et autres
Publié: (2025)
par: Sun, Yiming, et autres
Publié: (2025)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
par: Wang, Youze, et autres
Publié: (2025)
par: Wang, Youze, et autres
Publié: (2025)
NEWTON: Agentic Planning for Physically Grounded Video Generation
par: Feng, Yuxiang, et autres
Publié: (2026)
par: Feng, Yuxiang, et autres
Publié: (2026)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
par: Li, Chenglin, et autres
Publié: (2025)
par: Li, Chenglin, et autres
Publié: (2025)
VideoMamba: State Space Model for Efficient Video Understanding
par: Li, Kunchang, et autres
Publié: (2024)
par: Li, Kunchang, et autres
Publié: (2024)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
par: Liu, Xiangrui, et autres
Publié: (2025)
par: Liu, Xiangrui, et autres
Publié: (2025)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
par: Sun, Shangkun, et autres
Publié: (2024)
par: Sun, Shangkun, et autres
Publié: (2024)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
par: Li, Kunchang, et autres
Publié: (2023)
par: Li, Kunchang, et autres
Publié: (2023)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
par: Zhang, Zicheng, et autres
Publié: (2024)
par: Zhang, Zicheng, et autres
Publié: (2024)
EMCompress: Video-LLMs with Endomorphic Multimodal Compression
par: Fan, Zheyu, et autres
Publié: (2025)
par: Fan, Zheyu, et autres
Publié: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
par: Wang, Shihao, et autres
Publié: (2025)
par: Wang, Shihao, et autres
Publié: (2025)
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
par: Wang, Yuchi, et autres
Publié: (2024)
par: Wang, Yuchi, et autres
Publié: (2024)
VUDG: A Dataset for Video Understanding Domain Generalization
par: Wang, Ziyi, et autres
Publié: (2025)
par: Wang, Ziyi, et autres
Publié: (2025)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
par: Liu, Qing'an, et autres
Publié: (2026)
par: Liu, Qing'an, et autres
Publié: (2026)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
par: Ma, Wentao, et autres
Publié: (2025)
par: Ma, Wentao, et autres
Publié: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
par: Xu, Jiaqi, et autres
Publié: (2023)
par: Xu, Jiaqi, et autres
Publié: (2023)
Video Understanding: From Geometry and Semantics to Unified Models
par: An, Zhaochong, et autres
Publié: (2026)
par: An, Zhaochong, et autres
Publié: (2026)
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
par: Yu, Zhuang, et autres
Publié: (2026)
par: Yu, Zhuang, et autres
Publié: (2026)
SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
par: Wu, Jinlin, et autres
Publié: (2026)
par: Wu, Jinlin, et autres
Publié: (2026)
Documents similaires
-
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
par: Li, Shicheng, et autres
Publié: (2023) -
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
par: Chen, Sishuo, et autres
Publié: (2024) -
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
par: Ren, Shuhuai, et autres
Publié: (2023) -
TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
par: Li, Shicheng, et autres
Publié: (2025) -
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
par: Liu, Yuanxin, et autres
Publié: (2025)