MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Zile, Chen, Zhan, Zhu, Enze, Wei, Kan, Zou, Yongkang, Liu, Xiaoxuan, Wang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
by: Chen, Zhan, et al.
Published: (2025)
by: Chen, Zhan, et al.
Published: (2025)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
by: Qiao, Xiangshuo, et al.
Published: (2024)
by: Qiao, Xiangshuo, et al.
Published: (2024)
Exposure Completing for Temporally Consistent Neural High Dynamic Range Video Rendering
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
by: Zhang, Yuang, et al.
Published: (2024)
by: Zhang, Yuang, et al.
Published: (2024)
EVA: An Embodied World Model for Future Video Anticipation
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
by: Tan, Chaolei, et al.
Published: (2024)
by: Tan, Chaolei, et al.
Published: (2024)
Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline
by: Wu, Xuecheng, et al.
Published: (2023)
by: Wu, Xuecheng, et al.
Published: (2023)
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
by: Watanabe, Mitsuki, et al.
Published: (2025)
by: Watanabe, Mitsuki, et al.
Published: (2025)
VCEval: Rethinking What is a Good Educational Video and How to Automatically Evaluate It
by: Zhu, Xiaoxuan, et al.
Published: (2024)
by: Zhu, Xiaoxuan, et al.
Published: (2024)
Building and Evaluating a Realistic Virtual World for Large Scale Urban Exploration from 360° Videos
by: Takenawa, Mizuki, et al.
Published: (2025)
by: Takenawa, Mizuki, et al.
Published: (2025)
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
by: Chai, Zenghao, et al.
Published: (2024)
by: Chai, Zenghao, et al.
Published: (2024)
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
by: Hoang, Huong, et al.
Published: (2025)
by: Hoang, Huong, et al.
Published: (2025)
TrueFake: A Real World Case Dataset of Last Generation Fake Images also Shared on Social Networks
by: Dell'Anna, Stefano, et al.
Published: (2025)
by: Dell'Anna, Stefano, et al.
Published: (2025)
Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
by: Jiang, Chen, et al.
Published: (2023)
by: Jiang, Chen, et al.
Published: (2023)
MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
by: Zhao, Haochen, et al.
Published: (2025)
by: Zhao, Haochen, et al.
Published: (2025)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
by: Liu, Yuan, et al.
Published: (2024)
by: Liu, Yuan, et al.
Published: (2024)
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation
by: Chen, Lizhi, et al.
Published: (2025)
by: Chen, Lizhi, et al.
Published: (2025)
WoW: Towards a World omniscient World model Through Embodied Interaction
by: Chi, Xiaowei, et al.
Published: (2025)
by: Chi, Xiaowei, et al.
Published: (2025)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
by: Zhang, Zhongwei, et al.
Published: (2025)
by: Zhang, Zhongwei, et al.
Published: (2025)
Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
by: Wang, Yuze, et al.
Published: (2025)
by: Wang, Yuze, et al.
Published: (2025)
Towards Real-World Stickers Use: A New Dataset for Multi-Tag Sticker Recognition
by: Wang, Bingbing, et al.
Published: (2024)
by: Wang, Bingbing, et al.
Published: (2024)
Human Motion Video Generation: A Survey
by: Xue, Haiwei, et al.
Published: (2025)
by: Xue, Haiwei, et al.
Published: (2025)
Seeing World Dynamics in a Nutshell
by: Shen, Qiuhong, et al.
Published: (2025)
by: Shen, Qiuhong, et al.
Published: (2025)
L-LBVC: Long-Term Motion Estimation and Prediction for Learned Bi-Directional Video Compression
by: Zhai, Yongqi, et al.
Published: (2025)
by: Zhai, Yongqi, et al.
Published: (2025)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
by: Chu, Meng, et al.
Published: (2025)
by: Chu, Meng, et al.
Published: (2025)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
SFFNet: Synergistic Feature Fusion Network With Dual-Domain Edge Enhancement for UAV Image Object Detection
by: Zhang, Wenfeng, et al.
Published: (2026)
by: Zhang, Wenfeng, et al.
Published: (2026)
Waymo-3DSkelMo: A Multi-Agent 3D Skeletal Motion Dataset for Pedestrian Interaction Modeling in Autonomous Driving
by: Zhu, Guangxun, et al.
Published: (2025)
by: Zhu, Guangxun, et al.
Published: (2025)
CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos
by: Panambur, Tejas, et al.
Published: (2025)
by: Panambur, Tejas, et al.
Published: (2025)
PhyWorld: Physics-Faithful World Model for Video Generation
by: Zhao, Pu, et al.
Published: (2026)
by: Zhao, Pu, et al.
Published: (2026)
EyeNavGS: A 6-DoF Navigation Dataset and Record-n-Replay Software for Real-World 3DGS Scenes in VR
by: Ding, Zihao, et al.
Published: (2025)
by: Ding, Zihao, et al.
Published: (2025)
LookupForensics: A Large-Scale Multi-Task Dataset for Multi-Phase Image-Based Fact Verification
by: Cui, Shuhan, et al.
Published: (2024)
by: Cui, Shuhan, et al.
Published: (2024)
A Tri-Dynamic Preprocessing Framework for UGC Video Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
Arena: A Patch-of-Interest ViT Inference Acceleration System for Edge-Assisted Video Analytics
by: Peng, Haosong, et al.
Published: (2024)
by: Peng, Haosong, et al.
Published: (2024)
Med-Banana-50K: A Cross-modality Large-Scale Dataset for Text-guided Medical Image Editing
by: Chen, Zhihui, et al.
Published: (2025)
by: Chen, Zhihui, et al.
Published: (2025)
Failures to Surface Harmful Contents in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2025)
by: Cao, Yuxin, et al.
Published: (2025)
Context-Enhanced Video Moment Retrieval with Large Language Models
by: Liu, Weijia, et al.
Published: (2024)
by: Liu, Weijia, et al.
Published: (2024)
Similar Items
-
RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
by: Chen, Zhan, et al.
Published: (2025) -
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
by: Qiao, Xiangshuo, et al.
Published: (2024) -
Exposure Completing for Temporally Consistent Neural High Dynamic Range Video Rendering
by: Cui, Jiahao, et al.
Published: (2024) -
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
by: Zhang, Yuang, et al.
Published: (2024) -
EVA: An Embodied World Model for Future Video Anticipation
by: Chi, Xiaowei, et al.
Published: (2024)