Long Story Short: Story-level Video Understanding from 20K Short Films
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ghermi, Ridouane, Wang, Xi, Kalogeiton, Vicky, Laptev, Ivan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SF20K Competition 2025: Summary and findings
par: Ghermi, Ridouane, et autres
Publié: (2026)
par: Ghermi, Ridouane, et autres
Publié: (2026)
FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild
par: Liu, Zhi-Song, et autres
Publié: (2024)
par: Liu, Zhi-Song, et autres
Publié: (2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
par: Wang, Zun, et autres
Publié: (2024)
par: Wang, Zun, et autres
Publié: (2024)
AKiRa: Augmentation Kit on Rays for optical video generation
par: Wang, Xi, et autres
Publié: (2024)
par: Wang, Xi, et autres
Publié: (2024)
EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation
par: Wu, Xinyi, et autres
Publié: (2026)
par: Wu, Xinyi, et autres
Publié: (2026)
Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
par: Zhu, Yuanzhi, et autres
Publié: (2025)
par: Zhu, Yuanzhi, et autres
Publié: (2025)
Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator
par: Zhu, Yuanzhi, et autres
Publié: (2025)
par: Zhu, Yuanzhi, et autres
Publié: (2025)
Diffusion Reinforcement Learning via Centered Reward Distillation
par: Zhu, Yuanzhi, et autres
Publié: (2026)
par: Zhu, Yuanzhi, et autres
Publié: (2026)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
par: Fei, Jiajun, et autres
Publié: (2024)
par: Fei, Jiajun, et autres
Publié: (2024)
MLVU: Benchmarking Multi-task Long Video Understanding
par: Zhou, Junjie, et autres
Publié: (2024)
par: Zhou, Junjie, et autres
Publié: (2024)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
par: Kawasaki, Haruka, et autres
Publié: (2026)
par: Kawasaki, Haruka, et autres
Publié: (2026)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
par: Sharma, Aditya, et autres
Publié: (2024)
par: Sharma, Aditya, et autres
Publié: (2024)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
par: Wang, Ziyang, et autres
Publié: (2025)
par: Wang, Ziyang, et autres
Publié: (2025)
Adaptive Greedy Frame Selection for Long Video Understanding
par: Huang, Yuning, et autres
Publié: (2026)
par: Huang, Yuning, et autres
Publié: (2026)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
par: Salazar, Israfel, et autres
Publié: (2025)
par: Salazar, Israfel, et autres
Publié: (2025)
Movie2Story: A framework for understanding videos and telling stories in the form of novel text
par: Li, Kangning, et autres
Publié: (2024)
par: Li, Kangning, et autres
Publié: (2024)
USV: Towards Understanding the User-generated Short-form Videos
par: Cheng, Haoyue, et autres
Publié: (2026)
par: Cheng, Haoyue, et autres
Publié: (2026)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
par: Ahn, Jaewoo, et autres
Publié: (2025)
par: Ahn, Jaewoo, et autres
Publié: (2025)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
par: Zhang, Xiaoyi, et autres
Publié: (2025)
par: Zhang, Xiaoyi, et autres
Publié: (2025)
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
par: Chen, Shimin, et autres
Publié: (2024)
par: Chen, Shimin, et autres
Publié: (2024)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
par: Zhang, Jianrui, et autres
Publié: (2024)
par: Zhang, Jianrui, et autres
Publié: (2024)
StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification
par: He, Yichen, et autres
Publié: (2024)
par: He, Yichen, et autres
Publié: (2024)
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System
par: He, Lixuan, et autres
Publié: (2025)
par: He, Lixuan, et autres
Publié: (2025)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
par: Mao, Xiaofeng, et autres
Publié: (2026)
par: Mao, Xiaofeng, et autres
Publié: (2026)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
par: Wang, Ziyang, et autres
Publié: (2026)
par: Wang, Ziyang, et autres
Publié: (2026)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
par: Huang, Yuxiang, et autres
Publié: (2026)
par: Huang, Yuxiang, et autres
Publié: (2026)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
par: Ren, Shuhuai, et autres
Publié: (2023)
par: Ren, Shuhuai, et autres
Publié: (2023)
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
par: Xu, Weihan, et autres
Publié: (2025)
par: Xu, Weihan, et autres
Publié: (2025)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
par: Wang, Xiaohan, et autres
Publié: (2024)
par: Wang, Xiaohan, et autres
Publié: (2024)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
par: Li, Yunxin, et autres
Publié: (2024)
par: Li, Yunxin, et autres
Publié: (2024)
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
par: Xu, Boshen, et autres
Publié: (2025)
par: Xu, Boshen, et autres
Publié: (2025)
LEAD: Latent Realignment for Human Motion Diffusion
par: Andreou, Nefeli, et autres
Publié: (2024)
par: Andreou, Nefeli, et autres
Publié: (2024)
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
par: Chen, Yi-Chun
Publié: (2025)
par: Chen, Yi-Chun
Publié: (2025)
STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives
par: Wang, Bo, et autres
Publié: (2025)
par: Wang, Bo, et autres
Publié: (2025)
One-step Diffusion Models with Bregman Density Ratio Matching
par: Zhu, Yuanzhi, et autres
Publié: (2025)
par: Zhu, Yuanzhi, et autres
Publié: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
par: Wang, Ziyang, et autres
Publié: (2024)
par: Wang, Ziyang, et autres
Publié: (2024)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
par: Lin, Jingyang, et autres
Publié: (2026)
par: Lin, Jingyang, et autres
Publié: (2026)
Scaling RL to Long Videos
par: Chen, Yukang, et autres
Publié: (2025)
par: Chen, Yukang, et autres
Publié: (2025)
Temporal Preference Optimization for Long-Form Video Understanding
par: Li, Rui, et autres
Publié: (2025)
par: Li, Rui, et autres
Publié: (2025)
Documents similaires
-
SF20K Competition 2025: Summary and findings
par: Ghermi, Ridouane, et autres
Publié: (2026) -
FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild
par: Liu, Zhi-Song, et autres
Publié: (2024) -
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
par: Wang, Zun, et autres
Publié: (2024) -
AKiRa: Augmentation Kit on Rays for optical video generation
par: Wang, Xi, et autres
Publié: (2024) -
EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation
par: Wu, Xinyi, et autres
Publié: (2026)