HourVideo: 1-Hour Video-Language Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Chandrasegaran, Keshigeyan, Gupta, Agrim, Hadzic, Lea M., Kota, Taran, He, Jimming, Eyzaguirre, Cristóbal, Durante, Zane, Li, Manling, Wu, Jiajun, Fei-Fei, Li |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025)
by: Ye, Jinhui, et al.
Published: (2025)
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
Exploring Diffusion Transformer Designs via Grafting
by: Chandrasegaran, Keshigeyan, et al.
Published: (2025)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2025)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
by: Durante, Zane, et al.
Published: (2026)
by: Durante, Zane, et al.
Published: (2026)
GPIC: A Giant Permissive Image Corpus for Visual Generation
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
by: Liu, Yunong, et al.
Published: (2024)
by: Liu, Yunong, et al.
Published: (2024)
MindCube: Spatial Mental Modeling from Limited Views
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
Understanding Complexity in VideoQA via Visual Program Generation
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
by: Fei, Jiajun, et al.
Published: (2024)
by: Fei, Jiajun, et al.
Published: (2024)
Towards Fine-Grained Video Question Answering
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
by: Lin, Jingyang, et al.
Published: (2025)
by: Lin, Jingyang, et al.
Published: (2025)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
by: An, Joungbin, et al.
Published: (2026)
by: An, Joungbin, et al.
Published: (2026)
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
by: Dharmarajan, Karthik, et al.
Published: (2025)
by: Dharmarajan, Karthik, et al.
Published: (2025)
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
by: Shu, Yan, et al.
Published: (2024)
by: Shu, Yan, et al.
Published: (2024)
Towards Sparse Video Understanding and Reasoning
by: Xu, Chenwei, et al.
Published: (2026)
by: Xu, Chenwei, et al.
Published: (2026)
Universal Visuo-Tactile Video Understanding for Embodied Interaction
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
A Survey on Generative Modeling with Limited Data, Few Shots, and Zero Shot
by: Abdollahzadeh, Milad, et al.
Published: (2023)
by: Abdollahzadeh, Milad, et al.
Published: (2023)
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Model Inversion Robustness: Can Transfer Learning Help?
by: Ho, Sy-Tuyen, et al.
Published: (2024)
by: Ho, Sy-Tuyen, et al.
Published: (2024)
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
by: Zhang, Pingyue, et al.
Published: (2026)
by: Zhang, Pingyue, et al.
Published: (2026)
Autoregressive Flow Matching for Motion Prediction
by: Xie, Johnathan, et al.
Published: (2025)
by: Xie, Johnathan, et al.
Published: (2025)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Streaming Detection of Queried Event Start
by: Eyzaguirre, Cristobal, et al.
Published: (2024)
by: Eyzaguirre, Cristobal, et al.
Published: (2024)
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
by: Ren, Weiming, et al.
Published: (2025)
by: Ren, Weiming, et al.
Published: (2025)
Few-Shot Classification of Interactive Activities of Daily Living (InteractADL)
by: Durante, Zane, et al.
Published: (2024)
by: Durante, Zane, et al.
Published: (2024)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
by: Shi, Jiapeng, et al.
Published: (2026)
by: Shi, Jiapeng, et al.
Published: (2026)
Planning with the Views via Scene Self-Exploration
by: Wang, Kangrui, et al.
Published: (2026)
by: Wang, Kangrui, et al.
Published: (2026)
VideoMamba: State Space Model for Efficient Video Understanding
by: Li, Kunchang, et al.
Published: (2024)
by: Li, Kunchang, et al.
Published: (2024)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
by: Zhang, Yongheng, et al.
Published: (2025)
by: Zhang, Yongheng, et al.
Published: (2025)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
by: Zeng, Nianbo, et al.
Published: (2025)
by: Zeng, Nianbo, et al.
Published: (2025)
VideoChat: Chat-Centric Video Understanding
by: Li, KunChang, et al.
Published: (2023)
by: Li, KunChang, et al.
Published: (2023)
Video ReCap: Recursive Captioning of Hour-Long Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
by: Zou, Heqing, et al.
Published: (2024)
by: Zou, Heqing, et al.
Published: (2024)
Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
by: Lee, Seunghun, et al.
Published: (2025)
by: Lee, Seunghun, et al.
Published: (2025)
Online Video Understanding: OVBench and VideoChat-Online
by: Huang, Zhenpeng, et al.
Published: (2024)
by: Huang, Zhenpeng, et al.
Published: (2024)
Understanding Long Videos with Multimodal Language Models
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
LoL: Longer than Longer, Scaling Video Generation to Hour
by: Cui, Justin, et al.
Published: (2026)
by: Cui, Justin, et al.
Published: (2026)
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Similar Items
-
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025) -
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026) -
Exploring Diffusion Transformer Designs via Grafting
by: Chandrasegaran, Keshigeyan, et al.
Published: (2025) -
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
by: Durante, Zane, et al.
Published: (2026) -
GPIC: A Giant Permissive Image Corpus for Visual Generation
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)