NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, X., Yu, H., Wu, M., Hu, S., Chen, J., Zhu, C., Wu, J., Chu, X., Huang, K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
von: Feng, X., et al.
Veröffentlicht: (2026)
von: Feng, X., et al.
Veröffentlicht: (2026)
VideoAuteur: Towards Long Narrative Video Generation
von: Xiao, Junfei, et al.
Veröffentlicht: (2025)
von: Xiao, Junfei, et al.
Veröffentlicht: (2025)
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
von: Wu, Meiqi, et al.
Veröffentlicht: (2026)
von: Wu, Meiqi, et al.
Veröffentlicht: (2026)
ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
von: Feng, X., et al.
Veröffentlicht: (2025)
von: Feng, X., et al.
Veröffentlicht: (2025)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
von: Feng, X., et al.
Veröffentlicht: (2024)
von: Feng, X., et al.
Veröffentlicht: (2024)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features
von: Feng, X., et al.
Veröffentlicht: (2025)
von: Feng, X., et al.
Veröffentlicht: (2025)
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
von: Fang, Bo, et al.
Veröffentlicht: (2025)
von: Fang, Bo, et al.
Veröffentlicht: (2025)
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
von: Ha, Hyeonjeong, et al.
Veröffentlicht: (2026)
von: Ha, Hyeonjeong, et al.
Veröffentlicht: (2026)
STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation
von: Wang, Jiamin, et al.
Veröffentlicht: (2025)
von: Wang, Jiamin, et al.
Veröffentlicht: (2025)
Artifact-Aware Evaluation for High-Quality Video Generation
von: Zhu, Chen, et al.
Veröffentlicht: (2026)
von: Zhu, Chen, et al.
Veröffentlicht: (2026)
TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
VEU-Bench: Towards Comprehensive Understanding of Video Editing
von: Li, Bozheng, et al.
Veröffentlicht: (2025)
von: Li, Bozheng, et al.
Veröffentlicht: (2025)
FlowNar: Scalable Streaming Narration for Long-Form Videos
von: Zhong, Zeyun, et al.
Veröffentlicht: (2026)
von: Zhong, Zeyun, et al.
Veröffentlicht: (2026)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
Q-Save: Towards Scoring and Attribution for Generated Video Evaluation
von: Wu, Xiele, et al.
Veröffentlicht: (2025)
von: Wu, Xiele, et al.
Veröffentlicht: (2025)
FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component Analysis
von: Tan, Jiangtong, et al.
Veröffentlicht: (2025)
von: Tan, Jiangtong, et al.
Veröffentlicht: (2025)
ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)
Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
von: Meng, Yihao, et al.
Veröffentlicht: (2025)
von: Meng, Yihao, et al.
Veröffentlicht: (2025)
Multi-sentence Video Grounding for Long Video Generation
von: Feng, Wei, et al.
Veröffentlicht: (2024)
von: Feng, Wei, et al.
Veröffentlicht: (2024)
VMBench: A Benchmark for Perception-Aligned Video Motion Generation
von: Ling, Xinran, et al.
Veröffentlicht: (2025)
von: Ling, Xinran, et al.
Veröffentlicht: (2025)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
Towards One-step Causal Video Generation via Adversarial Self-Distillation
von: Yang, Yongqi, et al.
Veröffentlicht: (2025)
von: Yang, Yongqi, et al.
Veröffentlicht: (2025)
HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
von: Li, Yayuan, et al.
Veröffentlicht: (2024)
von: Li, Yayuan, et al.
Veröffentlicht: (2024)
Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
MorphoHELM: A Comprehensive Benchmark for Evaluating Representations for Microscopy-Based Morphology Assays
von: Hayir, Emre, et al.
Veröffentlicht: (2026)
von: Hayir, Emre, et al.
Veröffentlicht: (2026)
LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders
von: Naiman, Ilan, et al.
Veröffentlicht: (2025)
von: Naiman, Ilan, et al.
Veröffentlicht: (2025)
AnimationBench: Are Video Models Good at Character-Centric Animation?
von: Wu, Leyi, et al.
Veröffentlicht: (2026)
von: Wu, Leyi, et al.
Veröffentlicht: (2026)
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2024)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2025)
von: Bansal, Hritik, et al.
Veröffentlicht: (2025)
VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding
von: Yang, Ruoliu, et al.
Veröffentlicht: (2026)
von: Yang, Ruoliu, et al.
Veröffentlicht: (2026)
VideoMerge: Towards Training-free Long Video Generation
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
Towards Chunk-Wise Generation for Long Videos
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
Narrative Aligned Long Form Video Question Answering
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
LongLive: Real-time Interactive Long Video Generation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
von: Feng, X., et al.
Veröffentlicht: (2026) -
VideoAuteur: Towards Long Narrative Video Generation
von: Xiao, Junfei, et al.
Veröffentlicht: (2025) -
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
von: Wu, Meiqi, et al.
Veröffentlicht: (2026) -
ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
von: Feng, X., et al.
Veröffentlicht: (2025) -
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
von: Feng, X., et al.
Veröffentlicht: (2024)