SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Namekata, Koichi, Bahmani, Sherwin, Wu, Ziyi, Kant, Yash, Gilitschenski, Igor, Lindell, David B. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars
by: Taubner, Felix, et al.
Published: (2025)
by: Taubner, Felix, et al.
Published: (2025)
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
by: Bahmani, Sherwin, et al.
Published: (2024)
by: Bahmani, Sherwin, et al.
Published: (2024)
TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
by: Mohammadi, Mohammad, et al.
Published: (2025)
by: Mohammadi, Mohammad, et al.
Published: (2025)
Realistic Evaluation of Model Merging for Compositional Generalization
by: Tam, Derek, et al.
Published: (2024)
by: Tam, Derek, et al.
Published: (2024)
Grow with the Flow: 4D Reconstruction of Growing Plants with Gaussian Flow Fields
by: Luo, Weihan, et al.
Published: (2026)
by: Luo, Weihan, et al.
Published: (2026)
S$^2$Edit: Text-Guided Image Editing with Precise Semantic and Spatial Control
by: Liu, Xudong, et al.
Published: (2025)
by: Liu, Xudong, et al.
Published: (2025)
Mind the Time: Temporally-Controlled Multi-Event Video Generation
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
GStex: Per-Primitive Texturing of 2D Gaussian Splatting for Decoupled Appearance and Geometry Modeling
by: Rong, Victor, et al.
Published: (2024)
by: Rong, Victor, et al.
Published: (2024)
TC4D: Trajectory-Conditioned Text-to-4D Generation
by: Bahmani, Sherwin, et al.
Published: (2024)
by: Bahmani, Sherwin, et al.
Published: (2024)
Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
by: Bahmani, Sherwin, et al.
Published: (2025)
by: Bahmani, Sherwin, et al.
Published: (2025)
VibES: Induced Vibration for Persistent Event-Based Sensing
by: Polizzi, Vincenzo, et al.
Published: (2025)
by: Polizzi, Vincenzo, et al.
Published: (2025)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
by: Bahmani, Sherwin, et al.
Published: (2024)
by: Bahmani, Sherwin, et al.
Published: (2024)
SPAD : Spatially Aware Multiview Diffusers
by: Kant, Yash, et al.
Published: (2024)
by: Kant, Yash, et al.
Published: (2024)
Semantic Self-adaptation: Enhancing Generalization with a Single Sample
by: Bahmani, Sherwin, et al.
Published: (2022)
by: Bahmani, Sherwin, et al.
Published: (2022)
Vista4D: Video Reshooting with 4D Point Clouds
by: Lin, Kuan Heng, et al.
Published: (2026)
by: Lin, Kuan Heng, et al.
Published: (2026)
Pippo: High-Resolution Multi-View Humans from a Single Image
by: Kant, Yash, et al.
Published: (2025)
by: Kant, Yash, et al.
Published: (2025)
4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling
by: Bahmani, Sherwin, et al.
Published: (2023)
by: Bahmani, Sherwin, et al.
Published: (2023)
ActionParty: Multi-Subject Action Binding in Generative Video Games
by: Pondaven, Alexander, et al.
Published: (2026)
by: Pondaven, Alexander, et al.
Published: (2026)
LEOD: Label-Efficient Object Detection for Event Cameras
by: Wu, Ziyi, et al.
Published: (2023)
by: Wu, Ziyi, et al.
Published: (2023)
CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D Diffusion
by: He, Kai, et al.
Published: (2024)
by: He, Kai, et al.
Published: (2024)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
by: Wu, Ziyi, et al.
Published: (2025)
by: Wu, Ziyi, et al.
Published: (2025)
Producing and Leveraging Online Map Uncertainty in Trajectory Prediction
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
GaussianCut: Interactive segmentation via graph cut for 3D Gaussian Splatting
by: Jain, Umangi, et al.
Published: (2024)
by: Jain, Umangi, et al.
Published: (2024)
EventSplat: 3D Gaussian Splatting from Moving Event Cameras for Real-time Rendering
by: Yura, Toshiya, et al.
Published: (2024)
by: Yura, Toshiya, et al.
Published: (2024)
EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion Models
by: Namekata, Koichi, et al.
Published: (2024)
by: Namekata, Koichi, et al.
Published: (2024)
CamI2V: Camera-Controlled Image-to-Video Diffusion Model
by: Zheng, Guangcong, et al.
Published: (2024)
by: Zheng, Guangcong, et al.
Published: (2024)
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
by: Zhang, Zhiyuan, et al.
Published: (2025)
by: Zhang, Zhiyuan, et al.
Published: (2025)
FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
by: Ge, Yunyang, et al.
Published: (2025)
by: Ge, Yunyang, et al.
Published: (2025)
Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
by: Zheng, Shuhong, et al.
Published: (2025)
by: Zheng, Shuhong, et al.
Published: (2025)
Addressing Image Authenticity When Cameras Use Generative AI
by: Masud, Umar, et al.
Published: (2026)
by: Masud, Umar, et al.
Published: (2026)
Generating HDR Video from SDR Video
by: Tedla, SaiKiran, et al.
Published: (2026)
by: Tedla, SaiKiran, et al.
Published: (2026)
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling
by: Shi, Xiaoyu, et al.
Published: (2024)
by: Shi, Xiaoyu, et al.
Published: (2024)
SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
by: Rajabi, Javad, et al.
Published: (2026)
by: Rajabi, Javad, et al.
Published: (2026)
RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
by: Li, Teng, et al.
Published: (2025)
by: Li, Teng, et al.
Published: (2025)
Towards Unsupervised Blind Face Restoration using Diffusion Prior
by: Kuai, Tianshu, et al.
Published: (2024)
by: Kuai, Tianshu, et al.
Published: (2024)
Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior
by: Guo, Chen, et al.
Published: (2025)
by: Guo, Chen, et al.
Published: (2025)
SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
by: Sivakumar, Ssharvien Kumar, et al.
Published: (2025)
by: Sivakumar, Ssharvien Kumar, et al.
Published: (2025)
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
by: Shen, Guibao, et al.
Published: (2024)
by: Shen, Guibao, et al.
Published: (2024)
Pixel-Aligned Multi-View Generation with Depth Guided Decoder
by: Tang, Zhenggang, et al.
Published: (2024)
by: Tang, Zhenggang, et al.
Published: (2024)
SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned Diffusion
by: Xiang, Zhengkang, et al.
Published: (2025)
by: Xiang, Zhengkang, et al.
Published: (2025)
Similar Items
-
MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars
by: Taubner, Felix, et al.
Published: (2025) -
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
by: Bahmani, Sherwin, et al.
Published: (2024) -
TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
by: Mohammadi, Mohammad, et al.
Published: (2025) -
Realistic Evaluation of Model Merging for Compositional Generalization
by: Tam, Derek, et al.
Published: (2024) -
Grow with the Flow: 4D Reconstruction of Growing Plants with Gaussian Flow Fields
by: Luo, Weihan, et al.
Published: (2026)