Scaling Zero-Shot Reference-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Zijian, Liu, Shikun, Liu, Haozhe, Qiu, Haonan, An, Zhaochong, Ren, Weiming, Liu, Zhiheng, Huang, Xiaoke, Ng, Kam Woh, Xie, Tian, Han, Xiao, Cong, Yuren, Li, Hang, Zhu, Chuyan, Patel, Aditya, Xiang, Tao, He, Sen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
VecGlypher: Unified Vector Glyph Generation with Language Models
by: Huang, Xiaoke, et al.
Published: (2026)
by: Huang, Xiaoke, et al.
Published: (2026)
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
by: An, Zhaochong, et al.
Published: (2025)
by: An, Zhaochong, et al.
Published: (2025)
Learning Flow Fields in Attention for Controllable Person Image Generation
by: Zhou, Zijian, et al.
Published: (2024)
by: Zhou, Zijian, et al.
Published: (2024)
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
by: Liu, Haozhe, et al.
Published: (2025)
by: Liu, Haozhe, et al.
Published: (2025)
Scaling Sequence-to-Sequence Generative Neural Rendering
by: Liu, Shikun, et al.
Published: (2025)
by: Liu, Shikun, et al.
Published: (2025)
Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories
by: Jang, Wonbong, et al.
Published: (2026)
by: Jang, Wonbong, et al.
Published: (2026)
ConceptHash: Interpretable Fine-Grained Hashing via Concept Discovery
by: Ng, Kam Woh, et al.
Published: (2024)
by: Ng, Kam Woh, et al.
Published: (2024)
PartCraft: Crafting Creative Objects by Parts
by: Ng, Kam Woh, et al.
Published: (2024)
by: Ng, Kam Woh, et al.
Published: (2024)
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
by: Liu, Zhiheng, et al.
Published: (2026)
by: Liu, Zhiheng, et al.
Published: (2026)
Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models
by: Cao, Cong, et al.
Published: (2026)
by: Cao, Cong, et al.
Published: (2026)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
by: Wang, Haozhe, et al.
Published: (2026)
by: Wang, Haozhe, et al.
Published: (2026)
Zero-Shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
by: Cao, Cong, et al.
Published: (2024)
by: Cao, Cong, et al.
Published: (2024)
DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control
by: Wei, Yujie, et al.
Published: (2024)
by: Wei, Yujie, et al.
Published: (2024)
Learning by Imagining: Debiased Feature Augmentation for Compositional Zero-Shot Learning
by: Zhang, Haozhe, et al.
Published: (2025)
by: Zhang, Haozhe, et al.
Published: (2025)
IPR-NeRF: Ownership Verification meets Neural Radiance Field
by: Ong, Win Kent, et al.
Published: (2024)
by: Ong, Win Kent, et al.
Published: (2024)
VAInpaint: Zero-Shot Video-Audio inpainting framework with LLMs-driven Module
by: Wu, Kam Man, et al.
Published: (2025)
by: Wu, Kam Man, et al.
Published: (2025)
One-Shot Multilingual Font Generation Via ViT
by: Wang, Zhiheng, et al.
Published: (2024)
by: Wang, Zhiheng, et al.
Published: (2024)
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
by: Qiu, Lu, et al.
Published: (2025)
by: Qiu, Lu, et al.
Published: (2025)
CrowdMoGen: Zero-Shot Text-Driven Collective Motion Generation
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
by: Liu, Jeffrey, et al.
Published: (2025)
by: Liu, Jeffrey, et al.
Published: (2025)
Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
by: Su, Tongtong, et al.
Published: (2025)
by: Su, Tongtong, et al.
Published: (2025)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation
by: Liu, Ting, et al.
Published: (2025)
by: Liu, Ting, et al.
Published: (2025)
LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation
by: Li, Jiachen, et al.
Published: (2025)
by: Li, Jiachen, et al.
Published: (2025)
Context-based and Diversity-driven Specificity in Compositional Zero-Shot Learning
by: Li, Yun, et al.
Published: (2024)
by: Li, Yun, et al.
Published: (2024)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
by: She, D., et al.
Published: (2025)
by: She, D., et al.
Published: (2025)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
DriveVA: Video Action Models are Zero-Shot Drivers
by: Liu, Mengmeng, et al.
Published: (2026)
by: Liu, Mengmeng, et al.
Published: (2026)
Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
by: Liu, Kaiting, et al.
Published: (2026)
by: Liu, Kaiting, et al.
Published: (2026)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
by: Ren, Weiming, et al.
Published: (2024)
by: Ren, Weiming, et al.
Published: (2024)
Bus-Conditioned Zero-Shot Trajectory Generation via Task Arithmetic
by: Liu, Shuai, et al.
Published: (2026)
by: Liu, Shuai, et al.
Published: (2026)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
by: Li, Xiaolou, et al.
Published: (2024)
by: Li, Xiaolou, et al.
Published: (2024)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
by: Ku, Max, et al.
Published: (2024)
by: Ku, Max, et al.
Published: (2024)
NVS-Solver: Video Diffusion Model as Zero-Shot Novel View Synthesizer
by: You, Meng, et al.
Published: (2024)
by: You, Meng, et al.
Published: (2024)
Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
Binary Verification for Zero-Shot Vision
by: Hu, Rongbin, et al.
Published: (2025)
by: Hu, Rongbin, et al.
Published: (2025)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
EZSR: Event-based Zero-Shot Recognition
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
Similar Items
-
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025) -
VecGlypher: Unified Vector Glyph Generation with Language Models
by: Huang, Xiaoke, et al.
Published: (2026) -
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
by: An, Zhaochong, et al.
Published: (2025) -
Learning Flow Fields in Attention for Controllable Person Image Generation
by: Zhou, Zijian, et al.
Published: (2024) -
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
by: Liu, Zhiheng, et al.
Published: (2025)