Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Boyang, Chen, Xuweiyi, Gadelha, Matheus, Cheng, Zezhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic-Free Procedural 3D Shapes Are Surprisingly Good Teachers
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2026)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2026)
Probing the Mid-level Vision Capabilities of Self-Supervised Learning
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2025)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2025)
Open Vocabulary Monocular 3D Object Detection
von: Yao, Jin, et al.
Veröffentlicht: (2024)
von: Yao, Jin, et al.
Veröffentlicht: (2024)
Empowering Dynamic Urban Navigation with Stereo and Mid-Level Vision
von: Zhou, Wentao, et al.
Veröffentlicht: (2025)
von: Zhou, Wentao, et al.
Veröffentlicht: (2025)
SAB3R: Semantic-Augmented Backbone in 3D Reconstruction
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2025)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2025)
UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control
von: Xia, Tian, et al.
Veröffentlicht: (2024)
von: Xia, Tian, et al.
Veröffentlicht: (2024)
GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
von: Fortier-Chouinard, Frédéric, et al.
Veröffentlicht: (2025)
von: Fortier-Chouinard, Frédéric, et al.
Veröffentlicht: (2025)
Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image Sets
von: Decatur, Dale, et al.
Veröffentlicht: (2025)
von: Decatur, Dale, et al.
Veröffentlicht: (2025)
OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer
von: Wang, Boyang, et al.
Veröffentlicht: (2026)
von: Wang, Boyang, et al.
Veröffentlicht: (2026)
Learning Continuous 3D Words for Text-to-Image Generation
von: Cheng, Ta-Ying, et al.
Veröffentlicht: (2024)
von: Cheng, Ta-Ying, et al.
Veröffentlicht: (2024)
PSTF-AttControl: Per-Subject-Tuning-Free Personalized Image Generation with Controllable Face Attributes
von: liu, Xiang, et al.
Veröffentlicht: (2025)
von: liu, Xiang, et al.
Veröffentlicht: (2025)
3D Space as a Scratchpad for Editable Text-to-Image Generation
von: Saha, Oindrila, et al.
Veröffentlicht: (2026)
von: Saha, Oindrila, et al.
Veröffentlicht: (2026)
SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
von: Saha, Oindrila, et al.
Veröffentlicht: (2025)
von: Saha, Oindrila, et al.
Veröffentlicht: (2025)
Autoregressive Video Generation beyond Next Frames Prediction
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
von: Vandersanden, Jente, et al.
Veröffentlicht: (2026)
von: Vandersanden, Jente, et al.
Veröffentlicht: (2026)
Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation
von: Wang, Xiaojuan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaojuan, et al.
Veröffentlicht: (2024)
Material Magic Wand: Material-Aware Grouping of 3D Parts in Untextured Meshes
von: Jain, Umangi, et al.
Veröffentlicht: (2026)
von: Jain, Umangi, et al.
Veröffentlicht: (2026)
PreciseCam: Precise Camera Control for Text-to-Image Generation
von: Bernal-Berdun, Edurne, et al.
Veröffentlicht: (2025)
von: Bernal-Berdun, Edurne, et al.
Veröffentlicht: (2025)
FrameBridge: Improving Image-to-Video Generation with Bridge Models
von: Wang, Yuji, et al.
Veröffentlicht: (2024)
von: Wang, Yuji, et al.
Veröffentlicht: (2024)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
von: Yang, Shaoshu, et al.
Veröffentlicht: (2025)
von: Yang, Shaoshu, et al.
Veröffentlicht: (2025)
Generative Gaussian Splatting for Unbounded 3D City Generation
von: Xie, Haozhe, et al.
Veröffentlicht: (2024)
von: Xie, Haozhe, et al.
Veröffentlicht: (2024)
Compositional Generative Model of Unbounded 4D Cities
von: Xie, Haozhe, et al.
Veröffentlicht: (2025)
von: Xie, Haozhe, et al.
Veröffentlicht: (2025)
Generative Inbetweening through Frame-wise Conditions-Driven Video Generation
von: Zhu, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhu, Tianyi, et al.
Veröffentlicht: (2024)
InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models
von: Lu, Yifan, et al.
Veröffentlicht: (2024)
von: Lu, Yifan, et al.
Veröffentlicht: (2024)
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
von: Zhang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiyuan, et al.
Veröffentlicht: (2025)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
von: Luo, Rundong, et al.
Veröffentlicht: (2025)
von: Luo, Rundong, et al.
Veröffentlicht: (2025)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
VGDFR: Diffusion-based Video Generation with Dynamic Latent Frame Rate
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
FC-VFI: Faithful and Consistent Video Frame Interpolation for High-FPS Slow Motion Video Generation
von: Ding, Ganggui, et al.
Veröffentlicht: (2026)
von: Ding, Ganggui, et al.
Veröffentlicht: (2026)
Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization
von: Ali, Muhammad Kashif, et al.
Veröffentlicht: (2025)
von: Ali, Muhammad Kashif, et al.
Veröffentlicht: (2025)
Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery
von: Menn, Dennis, et al.
Veröffentlicht: (2026)
von: Menn, Dennis, et al.
Veröffentlicht: (2026)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
von: Peng, Bohao, et al.
Veröffentlicht: (2024)
von: Peng, Bohao, et al.
Veröffentlicht: (2024)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
von: Kizil, Muhammed Burak, et al.
Veröffentlicht: (2026)
von: Kizil, Muhammed Burak, et al.
Veröffentlicht: (2026)
Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction
von: Ma, Jiahao, et al.
Veröffentlicht: (2025)
von: Ma, Jiahao, et al.
Veröffentlicht: (2025)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Semantic-Free Procedural 3D Shapes Are Surprisingly Good Teachers
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024) -
WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2026) -
Probing the Mid-level Vision Capabilities of Self-Supervised Learning
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024) -
Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2025) -
Open Vocabulary Monocular 3D Object Detection
von: Yao, Jin, et al.
Veröffentlicht: (2024)