Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Qiuheng, Shi, Yukai, Ou, Jiarong, Chen, Rui, Lin, Ke, Wang, Jiahao, Jiang, Boyuan, Yang, Haotian, Zheng, Mingwu, Tao, Xin, Yang, Fei, Wan, Pengfei, Zhang, Di |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
von: Hu, Jiahao, et al.
Veröffentlicht: (2024)
von: Hu, Jiahao, et al.
Veröffentlicht: (2024)
Imbalance in Balance: Online Concept Balancing in Generation Models
von: Shi, Yukai, et al.
Veröffentlicht: (2025)
von: Shi, Yukai, et al.
Veröffentlicht: (2025)
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
von: Wang, Ruotong, et al.
Veröffentlicht: (2025)
von: Wang, Ruotong, et al.
Veröffentlicht: (2025)
Towards Precise Scaling Laws for Video Diffusion Transformers
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
Improving Video Generation with Human Feedback
von: Liu, Jie, et al.
Veröffentlicht: (2025)
von: Liu, Jie, et al.
Veröffentlicht: (2025)
Unleashing Video Language Models for Fine-grained HRCT Report Generation
von: Fang, Yingying, et al.
Veröffentlicht: (2026)
von: Fang, Yingying, et al.
Veröffentlicht: (2026)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
von: Hong, Wenyi, et al.
Veröffentlicht: (2025)
von: Hong, Wenyi, et al.
Veröffentlicht: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
VRMM: A Volumetric Relightable Morphable Head Model
von: Yang, Haotian, et al.
Veröffentlicht: (2024)
von: Yang, Haotian, et al.
Veröffentlicht: (2024)
Trajectory Attention for Fine-grained Video Motion Control
von: Xiao, Zeqi, et al.
Veröffentlicht: (2024)
von: Xiao, Zeqi, et al.
Veröffentlicht: (2024)
VideoTetris: Towards Compositional Text-to-Video Generation
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
Towards Fine-grained Interactive Segmentation in Images and Videos
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
Learning Temporally Consistent Video Depth from Video Diffusion Priors
von: Shao, Jiahao, et al.
Veröffentlicht: (2024)
von: Shao, Jiahao, et al.
Veröffentlicht: (2024)
Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control
von: Li, Bingliang, et al.
Veröffentlicht: (2024)
von: Li, Bingliang, et al.
Veröffentlicht: (2024)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
von: Sun, Boyuan, et al.
Veröffentlicht: (2026)
von: Sun, Boyuan, et al.
Veröffentlicht: (2026)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
von: Ji, Sihui, et al.
Veröffentlicht: (2025)
Birding at Koala Park
von: NA
Veröffentlicht: (1989)
von: NA
Veröffentlicht: (1989)
Fisiología del Koala
Veröffentlicht: (1980)
Veröffentlicht: (1980)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
von: Liang, Shuo, et al.
Veröffentlicht: (2025)
von: Liang, Shuo, et al.
Veröffentlicht: (2025)
Gloria: Consistent Character Video Generation via Content Anchors
von: Yang, Yuhang, et al.
Veröffentlicht: (2026)
von: Yang, Yuhang, et al.
Veröffentlicht: (2026)
CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
von: Zi, Bojia, et al.
Veröffentlicht: (2024)
von: Zi, Bojia, et al.
Veröffentlicht: (2024)
FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos
von: Wu, Zhengqian, et al.
Veröffentlicht: (2024)
von: Wu, Zhengqian, et al.
Veröffentlicht: (2024)
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
von: Zhang, Yidan, et al.
Veröffentlicht: (2025)
von: Zhang, Yidan, et al.
Veröffentlicht: (2025)
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
von: Zhong, Tianxiong, et al.
Veröffentlicht: (2025)
von: Zhong, Tianxiong, et al.
Veröffentlicht: (2025)
XS-VID: An Extremely Small Video Object Detection Dataset
von: Guo, Jiahao, et al.
Veröffentlicht: (2024)
von: Guo, Jiahao, et al.
Veröffentlicht: (2024)
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2025)
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2025)
AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
von: Girish, Sharath, et al.
Veröffentlicht: (2025)
von: Girish, Sharath, et al.
Veröffentlicht: (2025)
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
FineVQ: Fine-Grained User Generated Content Video Quality Assessment
von: Duan, Huiyu, et al.
Veröffentlicht: (2024)
von: Duan, Huiyu, et al.
Veröffentlicht: (2024)
ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation
von: Fu, Junhu, et al.
Veröffentlicht: (2026)
von: Fu, Junhu, et al.
Veröffentlicht: (2026)
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos
von: Lin, Jiahao, et al.
Veröffentlicht: (2025)
von: Lin, Jiahao, et al.
Veröffentlicht: (2025)
FingER: Content Aware Fine-grained Evaluation with Reasoning for AI-Generated Videos
von: Chen, Rui, et al.
Veröffentlicht: (2025)
von: Chen, Rui, et al.
Veröffentlicht: (2025)
VideoStudio: Generating Consistent-Content and Multi-Scene Videos
von: Long, Fuchen, et al.
Veröffentlicht: (2024)
von: Long, Fuchen, et al.
Veröffentlicht: (2024)
Koala Tree Planting Day
von: NA
Veröffentlicht: (2023)
von: NA
Veröffentlicht: (2023)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
von: Lu, Ning, et al.
Veröffentlicht: (2025)
von: Lu, Ning, et al.
Veröffentlicht: (2025)
Storyboard guided Alignment for Fine-grained Video Action Recognition
von: Liu, Enqi, et al.
Veröffentlicht: (2024)
von: Liu, Enqi, et al.
Veröffentlicht: (2024)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
von: Cai, Minghong, et al.
Veröffentlicht: (2025)
von: Cai, Minghong, et al.
Veröffentlicht: (2025)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
von: Wu, Jianzong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
von: Hu, Jiahao, et al.
Veröffentlicht: (2024) -
Imbalance in Balance: Online Concept Balancing in Generation Models
von: Shi, Yukai, et al.
Veröffentlicht: (2025) -
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
von: Wang, Ruotong, et al.
Veröffentlicht: (2025) -
Towards Precise Scaling Laws for Video Diffusion Transformers
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024) -
Improving Video Generation with Human Feedback
von: Liu, Jie, et al.
Veröffentlicht: (2025)