Video Creation by Demonstration
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Yihong, Zhou, Hao, Yuan, Liangzhe, Sun, Jennifer J., Li, Yandong, Jia, Xuhui, Adam, Hartwig, Hariharan, Bharath, Zhao, Long, Liu, Ting |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOD-UV: Learning Mobile Object Detectors from Unlabeled Videos
by: Sun, Yihong, et al.
Published: (2024)
by: Sun, Yihong, et al.
Published: (2024)
Live Interactive Training for Video Segmentation
by: Yang, Xinyu, et al.
Published: (2026)
by: Yang, Xinyu, et al.
Published: (2026)
Tracking and Understanding Object Transformations
by: Sun, Yihong, et al.
Published: (2025)
by: Sun, Yihong, et al.
Published: (2025)
Distilling Vision-Language Models on Millions of Videos
by: Zhao, Yue, et al.
Published: (2024)
by: Zhao, Yue, et al.
Published: (2024)
VideoGLUE: Video General Understanding Evaluation of Foundation Models
by: Yuan, Liangzhe, et al.
Published: (2023)
by: Yuan, Liangzhe, et al.
Published: (2023)
Epsilon-VAE: Denoising as Visual Decoding
by: Zhao, Long, et al.
Published: (2024)
by: Zhao, Long, et al.
Published: (2024)
VideoPrism: A Foundational Visual Encoder for Video Understanding
by: Zhao, Long, et al.
Published: (2024)
by: Zhao, Long, et al.
Published: (2024)
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
by: Xiong, Yuanhao, et al.
Published: (2023)
by: Xiong, Yuanhao, et al.
Published: (2023)
Learning Feature Descriptors using Camera Pose Supervision
by: Wang, Qianqian, et al.
Published: (2020)
by: Wang, Qianqian, et al.
Published: (2020)
MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
by: Shaar, Shaden, et al.
Published: (2026)
by: Shaar, Shaden, et al.
Published: (2026)
Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
by: Peng, Wenxuan, et al.
Published: (2026)
by: Peng, Wenxuan, et al.
Published: (2026)
What's in a Name? Beyond Class Indices for Image Recognition
by: Han, Kai, et al.
Published: (2023)
by: Han, Kai, et al.
Published: (2023)
C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
by: Huang, Kuan Wei, et al.
Published: (2025)
by: Huang, Kuan Wei, et al.
Published: (2025)
KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos
by: Chou, Gene, et al.
Published: (2024)
by: Chou, Gene, et al.
Published: (2024)
Think over Trajectories: Leveraging Video Generation to Reconstruct GPS Trajectories from Cellular Signaling
by: Zhang, Ruixing, et al.
Published: (2026)
by: Zhang, Ruixing, et al.
Published: (2026)
Sports Re-ID: Improving Re-Identification Of Players In Broadcast Videos Of Team Sports
by: Comandur, Bharath
Published: (2022)
by: Comandur, Bharath
Published: (2022)
Learning 3D Perception from Others' Predictions
by: Yoo, Jinsu, et al.
Published: (2024)
by: Yoo, Jinsu, et al.
Published: (2024)
Color Bind: Exploring Color Perception in Text-to-Image Models
by: Shomer-Chai, Shay, et al.
Published: (2025)
by: Shomer-Chai, Shay, et al.
Published: (2025)
$Δ$ynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos
by: Kao, Chia-Hsiang, et al.
Published: (2026)
by: Kao, Chia-Hsiang, et al.
Published: (2026)
FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution
by: Chou, Gene, et al.
Published: (2025)
by: Chou, Gene, et al.
Published: (2025)
Image Diffusion Preview with Consistency Solver
by: Wang, Fu-Yun, et al.
Published: (2025)
by: Wang, Fu-Yun, et al.
Published: (2025)
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
by: Revankar, Shreelekha, et al.
Published: (2025)
by: Revankar, Shreelekha, et al.
Published: (2025)
Scale-Aware Recognition in Satellite Images under Resource Constraints
by: Revankar, Shreelekha, et al.
Published: (2024)
by: Revankar, Shreelekha, et al.
Published: (2024)
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
by: Chou, Gene, et al.
Published: (2026)
by: Chou, Gene, et al.
Published: (2026)
Deep Understanding of Soccer Match Videos
by: Xu, Shikun, et al.
Published: (2024)
by: Xu, Shikun, et al.
Published: (2024)
PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild
by: Yuan, Kun, et al.
Published: (2024)
by: Yuan, Kun, et al.
Published: (2024)
Crowded Video Individual Counting Informed by Social Grouping and Spatial-Temporal Displacement Priors
by: Lu, Hao, et al.
Published: (2026)
by: Lu, Hao, et al.
Published: (2026)
Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
by: Cheng, Dabing, et al.
Published: (2025)
by: Cheng, Dabing, et al.
Published: (2025)
Rethinking Preference Alignment for Diffusion Models with Classifier-Free Guidance
by: Jiang, Zhou, et al.
Published: (2026)
by: Jiang, Zhou, et al.
Published: (2026)
VACE: All-in-One Video Creation and Editing
by: Jiang, Zeyinzi, et al.
Published: (2025)
by: Jiang, Zeyinzi, et al.
Published: (2025)
Video Individual Counting With Implicit One-to-Many Matching
by: Zhu, Xuhui, et al.
Published: (2025)
by: Zhu, Xuhui, et al.
Published: (2025)
LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
by: Song, Wenhui, et al.
Published: (2025)
by: Song, Wenhui, et al.
Published: (2025)
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
by: Zhou, Huayi, et al.
Published: (2025)
by: Zhou, Huayi, et al.
Published: (2025)
ObjectCarver: Semi-automatic segmentation, reconstruction and separation of 3D objects
by: Hassena, Gemmechu, et al.
Published: (2024)
by: Hassena, Gemmechu, et al.
Published: (2024)
Streaming Autoregressive Video Generation via Diagonal Distillation
by: Liu, Jinxiu, et al.
Published: (2026)
by: Liu, Jinxiu, et al.
Published: (2026)
AllClear: A Comprehensive Dataset and Benchmark for Cloud Removal in Satellite Imagery
by: Zhou, Hangyu, et al.
Published: (2024)
by: Zhou, Hangyu, et al.
Published: (2024)
Vidi2.5: Large Multimodal Models for Video Understanding and Creation
by: Vidi Team, et al.
Published: (2025)
by: Vidi Team, et al.
Published: (2025)
Biomedical SAM 2: Segment Anything in Biomedical Images and Videos
by: Yan, Zhiling, et al.
Published: (2024)
by: Yan, Zhiling, et al.
Published: (2024)
Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning
by: Hu, Zhangchi, et al.
Published: (2026)
by: Hu, Zhangchi, et al.
Published: (2026)
Similar Items
-
MOD-UV: Learning Mobile Object Detectors from Unlabeled Videos
by: Sun, Yihong, et al.
Published: (2024) -
Live Interactive Training for Video Segmentation
by: Yang, Xinyu, et al.
Published: (2026) -
Tracking and Understanding Object Transformations
by: Sun, Yihong, et al.
Published: (2025) -
Distilling Vision-Language Models on Millions of Videos
by: Zhao, Yue, et al.
Published: (2024) -
VideoGLUE: Video General Understanding Evaluation of Foundation Models
by: Yuan, Liangzhe, et al.
Published: (2023)