Video Understanding: Through A Temporal Lens
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Nguyen, Thong Thanh |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
par: Nguyen, Thong, et autres
Publié: (2025)
par: Nguyen, Thong, et autres
Publié: (2025)
Multi-Scale Contrastive Learning for Video Temporal Grounding
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
par: Nguyen, Hong, et autres
Publié: (2025)
par: Nguyen, Hong, et autres
Publié: (2025)
Lightweight Models for Emotional Analysis in Video
par: Nguyen, Quoc-Tien, et autres
Publié: (2025)
par: Nguyen, Quoc-Tien, et autres
Publié: (2025)
Encoding and Controlling Global Semantics for Long-form Video Question Answering
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
par: Yang, Min, et autres
Publié: (2024)
par: Yang, Min, et autres
Publié: (2024)
V-CORE: Temporally Consistent Video Understanding for Video-LLM
par: Kang, Zhengjian, et autres
Publié: (2026)
par: Kang, Zhengjian, et autres
Publié: (2026)
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
par: Zhao, Henghao, et autres
Publié: (2025)
par: Zhao, Henghao, et autres
Publié: (2025)
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
par: Nguyen, Trung Thanh, et autres
Publié: (2024)
par: Nguyen, Trung Thanh, et autres
Publié: (2024)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
par: Cao, Tri, et autres
Publié: (2026)
par: Cao, Tri, et autres
Publié: (2026)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
par: Shi, Jiapeng, et autres
Publié: (2026)
par: Shi, Jiapeng, et autres
Publié: (2026)
LensWalk: Agentic Video Understanding by Planning How You See in Videos
par: Li, Keliang, et autres
Publié: (2026)
par: Li, Keliang, et autres
Publié: (2026)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
par: Xu, Zhiyang, et autres
Publié: (2026)
par: Xu, Zhiyang, et autres
Publié: (2026)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
par: Nguyen, Thong, et autres
Publié: (2023)
par: Nguyen, Thong, et autres
Publié: (2023)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
par: Thanh, Toan Le Ngo, et autres
Publié: (2025)
par: Thanh, Toan Le Ngo, et autres
Publié: (2025)
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
par: Zhao, Xinkui, et autres
Publié: (2025)
par: Zhao, Xinkui, et autres
Publié: (2025)
BTS-rPPG: Orthogonal Butterfly Temporal Shifting for Remote Photoplethysmography
par: Nguyen, Ba-Thinh, et autres
Publié: (2026)
par: Nguyen, Ba-Thinh, et autres
Publié: (2026)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
par: Vu, Huu-An, et autres
Publié: (2025)
par: Vu, Huu-An, et autres
Publié: (2025)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
par: Zhou, Xingcheng, et autres
Publié: (2025)
par: Zhou, Xingcheng, et autres
Publié: (2025)
Seeing Through the Tool: A Controlled Benchmark for Occlusion Robustness in Foundation Segmentation Models
par: Ho, Nhan, et autres
Publié: (2026)
par: Ho, Nhan, et autres
Publié: (2026)
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
par: Rasekh, Ali, et autres
Publié: (2025)
par: Rasekh, Ali, et autres
Publié: (2025)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
par: Sun, Shitong, et autres
Publié: (2026)
par: Sun, Shitong, et autres
Publié: (2026)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
par: Ye, Jinhui, et autres
Publié: (2025)
par: Ye, Jinhui, et autres
Publié: (2025)
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
par: Liu, Zichen, et autres
Publié: (2025)
par: Liu, Zichen, et autres
Publié: (2025)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
par: Wang, Kaibin, et autres
Publié: (2025)
par: Wang, Kaibin, et autres
Publié: (2025)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
par: Guo, Yanan, et autres
Publié: (2025)
par: Guo, Yanan, et autres
Publié: (2025)
HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding
par: Nguyen, Trong-Thuan, et autres
Publié: (2023)
par: Nguyen, Trong-Thuan, et autres
Publié: (2023)
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
par: Yang, Zhenyu, et autres
Publié: (2025)
par: Yang, Zhenyu, et autres
Publié: (2025)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
par: Nguyen, Thong, et autres
Publié: (2023)
par: Nguyen, Thong, et autres
Publié: (2023)
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
par: Nie, Ming, et autres
Publié: (2026)
par: Nie, Ming, et autres
Publié: (2026)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
par: Lee, Yeonkyung, et autres
Publié: (2026)
par: Lee, Yeonkyung, et autres
Publié: (2026)
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
par: Elsharkawi, Ismael, et autres
Publié: (2026)
par: Elsharkawi, Ismael, et autres
Publié: (2026)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
par: Vani, Sameep, et autres
Publié: (2025)
par: Vani, Sameep, et autres
Publié: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
par: Hu, Pengfei, et autres
Publié: (2025)
par: Hu, Pengfei, et autres
Publié: (2025)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
par: Wang, Shaobo, et autres
Publié: (2025)
par: Wang, Shaobo, et autres
Publié: (2025)
Multimodal Contextualized Support for Enhancing Video Retrieval System
par: Nguyen-Le, Quoc-Bao, et autres
Publié: (2024)
par: Nguyen-Le, Quoc-Bao, et autres
Publié: (2024)
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
par: Truong, Thanh-Dat, et autres
Publié: (2025)
par: Truong, Thanh-Dat, et autres
Publié: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
par: Zhang, Jun, et autres
Publié: (2025)
par: Zhang, Jun, et autres
Publié: (2025)
TB-Bench: Training and Testing Multi-Modal AI for Understanding Spatio-Temporal Traffic Behaviors from Dashcam Images/Videos
par: Charoenpitaks, Korawat, et autres
Publié: (2025)
par: Charoenpitaks, Korawat, et autres
Publié: (2025)
Documents similaires
-
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
par: Nguyen, Thong, et autres
Publié: (2025) -
Multi-Scale Contrastive Learning for Video Temporal Grounding
par: Nguyen, Thong Thanh, et autres
Publié: (2024) -
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
par: Nguyen, Thong Thanh, et autres
Publié: (2024) -
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
par: Nguyen, Hong, et autres
Publié: (2025) -
Lightweight Models for Emotional Analysis in Video
par: Nguyen, Quoc-Tien, et autres
Publié: (2025)