VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sarkar, Pritam, Etemad, Ali |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
von: Sarkar, Pritam, et al.
Veröffentlicht: (2025)
von: Sarkar, Pritam, et al.
Veröffentlicht: (2025)
Consistency-guided Prompt Learning for Vision-Language Models
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
von: Sarkar, Pritam, et al.
Veröffentlicht: (2024)
von: Sarkar, Pritam, et al.
Veröffentlicht: (2024)
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
von: Roy, Shuvendu, et al.
Veröffentlicht: (2025)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2025)
Exploring the Boundaries of Semi-Supervised Facial Expression Recognition using In-Distribution, Out-of-Distribution, and Unconstrained Data
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
CycleCrash: A Dataset of Bicycle Collision Videos for Collision Prediction and Analysis
von: Desai, Nishq Poorav, et al.
Veröffentlicht: (2024)
von: Desai, Nishq Poorav, et al.
Veröffentlicht: (2024)
CollideNet: Hierarchical Multi-scale Video Representation Learning with Disentanglement for Time-To-Collision Forecasting
von: Desai, Nishq Poorav, et al.
Veröffentlicht: (2026)
von: Desai, Nishq Poorav, et al.
Veröffentlicht: (2026)
Impact of Strategic Sampling and Supervision Policies on Semi-supervised Learning
von: Roy, Shuvendu, et al.
Veröffentlicht: (2022)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2022)
Partial Label Learning for Emotion Recognition from EEG
von: Zhang, Guangyi, et al.
Veröffentlicht: (2023)
von: Zhang, Guangyi, et al.
Veröffentlicht: (2023)
Investigating Video Reasoning Capability of Large Language Models with Tropes in Movies
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
On The Relationship Between Continual Learning and Long-Tailed Recognition
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2023)
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2023)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
Diffusion Models with Deterministic Normalizing Flow Priors
von: Zand, Mohsen, et al.
Veröffentlicht: (2023)
von: Zand, Mohsen, et al.
Veröffentlicht: (2023)
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
von: Dong, Yuhao, et al.
Veröffentlicht: (2024)
von: Dong, Yuhao, et al.
Veröffentlicht: (2024)
Unmasking Deepfakes: Masked Autoencoding Spatiotemporal Transformers for Enhanced Video Forgery Detection
von: Das, Sayantan, et al.
Veröffentlicht: (2023)
von: Das, Sayantan, et al.
Veröffentlicht: (2023)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
Scaling Up Semi-supervised Learning with Unconstrained Unlabelled Data
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
Human Pose Estimation from Ambiguous Pressure Recordings with Spatio-temporal Masked Transformers
von: Davoodnia, Vandad, et al.
Veröffentlicht: (2023)
von: Davoodnia, Vandad, et al.
Veröffentlicht: (2023)
GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
Consistency-Guided Asynchronous Contrastive Tuning for Few-Shot Class-Incremental Tuning of Foundation Models
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
LongVLM: Efficient Long Video Understanding via Large Language Models
von: Weng, Yuetian, et al.
Veröffentlicht: (2024)
von: Weng, Yuetian, et al.
Veröffentlicht: (2024)
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
von: He, Zhihao, et al.
Veröffentlicht: (2025)
von: He, Zhihao, et al.
Veröffentlicht: (2025)
ViLLa: Video Reasoning Segmentation with Large Language Model
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
von: Tang, Liyan, et al.
Veröffentlicht: (2025)
von: Tang, Liyan, et al.
Veröffentlicht: (2025)
Language Model Guided Interpretable Video Action Reasoning
von: Wang, Ning, et al.
Veröffentlicht: (2024)
von: Wang, Ning, et al.
Veröffentlicht: (2024)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
Multi-Frame Vision-Language Model for Long-form Reasoning in Driver Behavior Analysis
von: Takato, Hiroshi, et al.
Veröffentlicht: (2024)
von: Takato, Hiroshi, et al.
Veröffentlicht: (2024)
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
von: Cheng, Ying, et al.
Veröffentlicht: (2025)
von: Cheng, Ying, et al.
Veröffentlicht: (2025)
VISA: Reasoning Video Object Segmentation via Large Language Models
von: Yan, Cilin, et al.
Veröffentlicht: (2024)
von: Yan, Cilin, et al.
Veröffentlicht: (2024)
Follow the Rules: Reasoning for Video Anomaly Detection with Large Language Models
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
von: Salamatian, Ali, et al.
Veröffentlicht: (2026)
von: Salamatian, Ali, et al.
Veröffentlicht: (2026)
VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
von: Li, Yunhao, et al.
Veröffentlicht: (2026)
von: Li, Yunhao, et al.
Veröffentlicht: (2026)
Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
von: Shen, Yiqing, et al.
Veröffentlicht: (2025)
von: Shen, Yiqing, et al.
Veröffentlicht: (2025)
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
von: Xu, Shilin, et al.
Veröffentlicht: (2025)
von: Xu, Shilin, et al.
Veröffentlicht: (2025)
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models
von: Souza, Rafael, et al.
Veröffentlicht: (2024)
von: Souza, Rafael, et al.
Veröffentlicht: (2024)
Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation
von: Luo, Jingnan, et al.
Veröffentlicht: (2026)
von: Luo, Jingnan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
von: Sarkar, Pritam, et al.
Veröffentlicht: (2025) -
Consistency-guided Prompt Learning for Vision-Language Models
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023) -
Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
von: Sarkar, Pritam, et al.
Veröffentlicht: (2024) -
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
von: Roy, Shuvendu, et al.
Veröffentlicht: (2025) -
Exploring the Boundaries of Semi-Supervised Facial Expression Recognition using In-Distribution, Out-of-Distribution, and Unconstrained Data
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)