Movie2Story: A framework for understanding videos and telling stories in the form of novel text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Kangning, Jia, Zheyang, Ying, Anyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MovieCORE: COgnitive REasoning in Movies
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2025)
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2025)
Long Story Short: Story-level Video Understanding from 20K Short Films
von: Ghermi, Ridouane, et al.
Veröffentlicht: (2024)
von: Ghermi, Ridouane, et al.
Veröffentlicht: (2024)
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2024)
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2024)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation
von: Wu, Xinyi, et al.
Veröffentlicht: (2026)
von: Wu, Xinyi, et al.
Veröffentlicht: (2026)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
Navigation with VLM framework: Towards Going to Any Language
von: Yin, Zecheng, et al.
Veröffentlicht: (2024)
von: Yin, Zecheng, et al.
Veröffentlicht: (2024)
MovieTeller: Tool-augmented Movie Synopsis with ID Consistent Progressive Abstraction
von: Li, Yizhi, et al.
Veröffentlicht: (2026)
von: Li, Yizhi, et al.
Veröffentlicht: (2026)
GlitchBench: Can large multimodal models detect video game glitches?
von: Taesiri, Mohammad Reza, et al.
Veröffentlicht: (2023)
von: Taesiri, Mohammad Reza, et al.
Veröffentlicht: (2023)
VideoGameBench: Can Vision-Language Models complete popular video games?
von: Zhang, Alex L., et al.
Veröffentlicht: (2025)
von: Zhang, Alex L., et al.
Veröffentlicht: (2025)
RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
von: Yang, Zixi, et al.
Veröffentlicht: (2025)
von: Yang, Zixi, et al.
Veröffentlicht: (2025)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding
von: Zhong, Ziqi, et al.
Veröffentlicht: (2025)
von: Zhong, Ziqi, et al.
Veröffentlicht: (2025)
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
von: Kim, Yoonshik, et al.
Veröffentlicht: (2025)
von: Kim, Yoonshik, et al.
Veröffentlicht: (2025)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
Using Shapley interactions to understand how models use structure
von: Singhvi, Divyansh, et al.
Veröffentlicht: (2024)
von: Singhvi, Divyansh, et al.
Veröffentlicht: (2024)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
von: Xu, Zhe, et al.
Veröffentlicht: (2025)
von: Xu, Zhe, et al.
Veröffentlicht: (2025)
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
von: Jia, Yiming, et al.
Veröffentlicht: (2025)
von: Jia, Yiming, et al.
Veröffentlicht: (2025)
Keyword-Oriented Multimodal Modeling for Euphemism Identification
von: Hu, Yuxue, et al.
Veröffentlicht: (2025)
von: Hu, Yuxue, et al.
Veröffentlicht: (2025)
Spotting tell-tale visual artifacts in face swapping videos: strengths and pitfalls of CNN detectors
von: Ziglio, Riccardo, et al.
Veröffentlicht: (2025)
von: Ziglio, Riccardo, et al.
Veröffentlicht: (2025)
Improving fine-grained understanding in image-text pre-training
von: Bica, Ioana, et al.
Veröffentlicht: (2024)
von: Bica, Ioana, et al.
Veröffentlicht: (2024)
OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2025)
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2025)
LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding
von: Zhao, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2024)
Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
von: Patel, Het, et al.
Veröffentlicht: (2025)
von: Patel, Het, et al.
Veröffentlicht: (2025)
Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline
von: Jia, Qi, et al.
Veröffentlicht: (2024)
von: Jia, Qi, et al.
Veröffentlicht: (2024)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
von: Li, Yanwei, et al.
Veröffentlicht: (2024)
von: Li, Yanwei, et al.
Veröffentlicht: (2024)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
Enhancing Vision-Language Model Pre-training with Image-text Pair Pruning Based on Word Frequency
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
Movie101v2: Improved Movie Narration Benchmark
von: Yue, Zihao, et al.
Veröffentlicht: (2024)
von: Yue, Zihao, et al.
Veröffentlicht: (2024)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
von: Oliveira, Daniel, et al.
Veröffentlicht: (2026)
von: Oliveira, Daniel, et al.
Veröffentlicht: (2026)
ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling
von: Maina, Hernán, et al.
Veröffentlicht: (2025)
von: Maina, Hernán, et al.
Veröffentlicht: (2025)
Speaking images. A novel framework for the automated self-description of artworks
von: Bernasconi, Valentine, et al.
Veröffentlicht: (2025)
von: Bernasconi, Valentine, et al.
Veröffentlicht: (2025)
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
von: Liu, Daixian, et al.
Veröffentlicht: (2026)
Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs
von: Yun, Sukmin, et al.
Veröffentlicht: (2024)
von: Yun, Sukmin, et al.
Veröffentlicht: (2024)
VideoScore2: Think before You Score in Generative Video Evaluation
von: He, Xuan, et al.
Veröffentlicht: (2025)
von: He, Xuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MovieCORE: COgnitive REasoning in Movies
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2025) -
Long Story Short: Story-level Video Understanding from 20K Short Films
von: Ghermi, Ridouane, et al.
Veröffentlicht: (2024) -
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2024) -
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025) -
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)