VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Wenhao, Yang, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
DeMamba: AI-Generated Video Detection on Million-Scale GenVideo Benchmark
von: Chen, Haoxing, et al.
Veröffentlicht: (2024)
von: Chen, Haoxing, et al.
Veröffentlicht: (2024)
Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation
von: Zhou, Yufan, et al.
Veröffentlicht: (2024)
von: Zhou, Yufan, et al.
Veröffentlicht: (2024)
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
von: Yuan, Shenghai, et al.
Veröffentlicht: (2025)
von: Yuan, Shenghai, et al.
Veröffentlicht: (2025)
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
von: Wang, Jiarui, et al.
Veröffentlicht: (2025)
von: Wang, Jiarui, et al.
Veröffentlicht: (2025)
UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer
von: Liu, Delong, et al.
Veröffentlicht: (2024)
von: Liu, Delong, et al.
Veröffentlicht: (2024)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
von: Wang, Chenting, et al.
Veröffentlicht: (2025)
von: Wang, Chenting, et al.
Veröffentlicht: (2025)
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
von: Wang, Yi, et al.
Veröffentlicht: (2023)
von: Wang, Yi, et al.
Veröffentlicht: (2023)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
VideoTetris: Towards Compositional Text-to-Video Generation
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
A Large-Scale Study on Video Action Dataset Condensation
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
Distilling Vision-Language Models on Millions of Videos
von: Zhao, Yue, et al.
Veröffentlicht: (2024)
von: Zhao, Yue, et al.
Veröffentlicht: (2024)
AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmark
von: Lin, Li, et al.
Veröffentlicht: (2024)
von: Lin, Li, et al.
Veröffentlicht: (2024)
HumanNet: Scaling Human-centric Video Learning to One Million Hours
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
CHUG: Crowdsourced User-Generated HDR Video Quality Dataset
von: Saini, Shreshth, et al.
Veröffentlicht: (2025)
von: Saini, Shreshth, et al.
Veröffentlicht: (2025)
When Words Can't Capture It All: Towards Video-Based User Complaint Text Generation with Multimodal Video Complaint Dataset
von: Das, Sarmistha, et al.
Veröffentlicht: (2025)
von: Das, Sarmistha, et al.
Veröffentlicht: (2025)
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2026)
von: Zheng, Minghang, et al.
Veröffentlicht: (2026)
GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
von: Kang, Jenna, et al.
Veröffentlicht: (2025)
von: Kang, Jenna, et al.
Veröffentlicht: (2025)
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
von: Cheng, Zheng, et al.
Veröffentlicht: (2024)
von: Cheng, Zheng, et al.
Veröffentlicht: (2024)
Exploring AIGC Video Quality: A Focus on Visual Harmony, Video-Text Consistency and Domain Distribution Gap
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
SIGMAN:Scaling 3D Human Gaussian Generation with Millions of Assets
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
CI-VID: A Coherent Interleaved Text-Video Dataset
von: Ju, Yiming, et al.
Veröffentlicht: (2025)
von: Ju, Yiming, et al.
Veröffentlicht: (2025)
TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
von: Chen, Jiaben, et al.
Veröffentlicht: (2025)
von: Chen, Jiaben, et al.
Veröffentlicht: (2025)
Can Text-to-Video Generation help Video-Language Alignment?
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
CoMo: Compositional Motion Customization for Text-to-Video Generation
von: Xu, Youcan, et al.
Veröffentlicht: (2025)
von: Xu, Youcan, et al.
Veröffentlicht: (2025)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
von: Wang, Zhao, et al.
Veröffentlicht: (2024)
SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
von: Peng, Bo, et al.
Veröffentlicht: (2023)
von: Peng, Bo, et al.
Veröffentlicht: (2023)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2025)
von: He, Haibin, et al.
Veröffentlicht: (2025)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
von: Yang, Shuyu, et al.
Veröffentlicht: (2025)
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
von: Yang, Pei, et al.
Veröffentlicht: (2025)
von: Yang, Pei, et al.
Veröffentlicht: (2025)
UW-VOS: A Large-Scale Dataset for Underwater Video Object Segmentation
von: Zhao, Hongshen, et al.
Veröffentlicht: (2026)
von: Zhao, Hongshen, et al.
Veröffentlicht: (2026)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
Reinforcing Video Reasoning with Focused Thinking
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
von: Wang, Wenhao, et al.
Veröffentlicht: (2024) -
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
von: Wang, Wenhao, et al.
Veröffentlicht: (2024) -
DeMamba: AI-Generated Video Detection on Million-Scale GenVideo Benchmark
von: Chen, Haoxing, et al.
Veröffentlicht: (2024) -
Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation
von: Zhou, Yufan, et al.
Veröffentlicht: (2024) -
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
von: Yuan, Shenghai, et al.
Veröffentlicht: (2025)