LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xiong, Tianwei, Wang, Yuqing, Zhou, Daquan, Lin, Zhijie, Feng, Jiashi, Liu, Xihui |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
par: Wang, Yuqing, et autres
Publié: (2024)
par: Wang, Yuqing, et autres
Publié: (2024)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
par: Xu, Lin, et autres
Publié: (2024)
par: Xu, Lin, et autres
Publié: (2024)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
par: Xiong, Tianwei, et autres
Publié: (2026)
par: Xiong, Tianwei, et autres
Publié: (2026)
MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
par: Wang, Weimin, et autres
Publié: (2024)
par: Wang, Weimin, et autres
Publié: (2024)
Dense Video Captioning using Graph-based Sentence Summarization
par: Zhang, Zhiwang, et autres
Publié: (2025)
par: Zhang, Zhiwang, et autres
Publié: (2025)
Magic-Me: Identity-Specific Video Customized Diffusion
par: Ma, Ze, et autres
Publié: (2024)
par: Ma, Ze, et autres
Publié: (2024)
How Far is Video Generation from World Model: A Physical Law Perspective
par: Kang, Bingyi, et autres
Publié: (2024)
par: Kang, Bingyi, et autres
Publié: (2024)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
par: Chen, Sili, et autres
Publié: (2025)
par: Chen, Sili, et autres
Publié: (2025)
TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
par: Cheng, Wei-Yuan, et autres
Publié: (2026)
par: Cheng, Wei-Yuan, et autres
Publié: (2026)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
par: Kim, Ye-Chan, et autres
Publié: (2026)
par: Kim, Ye-Chan, et autres
Publié: (2026)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
par: Zhou, Yupeng, et autres
Publié: (2024)
par: Zhou, Yupeng, et autres
Publié: (2024)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
par: Xiong, Tianwei, et autres
Publié: (2025)
par: Xiong, Tianwei, et autres
Publié: (2025)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
par: Wu, Hao, et autres
Publié: (2024)
par: Wu, Hao, et autres
Publié: (2024)
LVD-GS: Gaussian Splatting SLAM for Dynamic Scenes via Hierarchical Explicit-Implicit Representation Collaboration Rendering
par: Zhu, Wenkai, et autres
Publié: (2025)
par: Zhu, Wenkai, et autres
Publié: (2025)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
par: Wang, Yuqing, et autres
Publié: (2025)
par: Wang, Yuqing, et autres
Publié: (2025)
Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
par: Zhang, Zhiwang, et autres
Publié: (2025)
par: Zhang, Zhiwang, et autres
Publié: (2025)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
par: Sasse, Kuleen, et autres
Publié: (2025)
par: Sasse, Kuleen, et autres
Publié: (2025)
Editing Massive Concepts in Text-to-Image Diffusion Models
par: Xiong, Tianwei, et autres
Publié: (2024)
par: Xiong, Tianwei, et autres
Publié: (2024)
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
par: Awadalla, Anas, et autres
Publié: (2024)
par: Awadalla, Anas, et autres
Publié: (2024)
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
par: Chen, Ruizhe, et autres
Publié: (2025)
par: Chen, Ruizhe, et autres
Publié: (2025)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
par: Ge, Shiping, et autres
Publié: (2024)
par: Ge, Shiping, et autres
Publié: (2024)
DeVAn: Dense Video Annotation for Video-Language Models
par: Liu, Tingkai, et autres
Publié: (2023)
par: Liu, Tingkai, et autres
Publié: (2023)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
par: Wu, Shengqiong, et autres
Publié: (2025)
par: Wu, Shengqiong, et autres
Publié: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
par: Xing, Long, et autres
Publié: (2025)
par: Xing, Long, et autres
Publié: (2025)
Towards Fine-Grained Human Motion Video Captioning
par: Song, Guorui, et autres
Publié: (2025)
par: Song, Guorui, et autres
Publié: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
par: Lian, Long, et autres
Publié: (2025)
par: Lian, Long, et autres
Publié: (2025)
Accurate and Fast Compressed Video Captioning
par: Shen, Yaojie, et autres
Publié: (2023)
par: Shen, Yaojie, et autres
Publié: (2023)
Streaming Dense Video Captioning
par: Zhou, Xingyi, et autres
Publié: (2024)
par: Zhou, Xingyi, et autres
Publié: (2024)
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
par: Liu, Xiangrui, et autres
Publié: (2025)
par: Liu, Xiangrui, et autres
Publié: (2025)
Video Summarization: Towards Entity-Aware Captions
par: Ayyubi, Hammad A., et autres
Publié: (2023)
par: Ayyubi, Hammad A., et autres
Publié: (2023)
Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization
par: Luo, Yongdong, et autres
Publié: (2024)
par: Luo, Yongdong, et autres
Publié: (2024)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
par: Truong, Quang-Trung, et autres
Publié: (2025)
par: Truong, Quang-Trung, et autres
Publié: (2025)
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
par: Chen, Shimin, et autres
Publié: (2024)
par: Chen, Shimin, et autres
Publié: (2024)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
par: Ding, Yang, et autres
Publié: (2025)
par: Ding, Yang, et autres
Publié: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
par: Fu, Honghao, et autres
Publié: (2026)
par: Fu, Honghao, et autres
Publié: (2026)
Image Captioning in news report scenario
par: Liu, Tianrui, et autres
Publié: (2024)
par: Liu, Tianrui, et autres
Publié: (2024)
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
par: Zhang, Runze, et autres
Publié: (2025)
par: Zhang, Runze, et autres
Publié: (2025)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
par: Lim, Junyoung, et autres
Publié: (2025)
par: Lim, Junyoung, et autres
Publié: (2025)
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
par: Yan, Haiyang, et autres
Publié: (2026)
par: Yan, Haiyang, et autres
Publié: (2026)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
par: Pham, Anh-Cuong, et autres
Publié: (2024)
par: Pham, Anh-Cuong, et autres
Publié: (2024)
Documents similaires
-
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
par: Wang, Yuqing, et autres
Publié: (2024) -
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
par: Xu, Lin, et autres
Publié: (2024) -
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
par: Xiong, Tianwei, et autres
Publié: (2026) -
MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
par: Wang, Weimin, et autres
Publié: (2024) -
Dense Video Captioning using Graph-based Sentence Summarization
par: Zhang, Zhiwang, et autres
Publié: (2025)