SPATIALALIGN: Aligning Dynamic Spatial Relationships in Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Fengming, Cham, Tat-Jen, Zheng, Chuanxia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One-shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing
by: Ji, Yuzhu, et al.
Published: (2024)
by: Ji, Yuzhu, et al.
Published: (2024)
PanoDiffusion: 360-degree Panorama Outpainting via Diffusion
by: Wu, Tianhao, et al.
Published: (2023)
by: Wu, Tianhao, et al.
Published: (2023)
ClusteringSDF: Self-Organized Neural Implicit Surfaces for 3D Decomposition
by: Wu, Tianhao, et al.
Published: (2024)
by: Wu, Tianhao, et al.
Published: (2024)
Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
by: Wu, Tianhao, et al.
Published: (2025)
by: Wu, Tianhao, et al.
Published: (2025)
Semantix: An Energy Guided Sampler for Semantic Style Transfer
by: He, Huiang, et al.
Published: (2025)
by: He, Huiang, et al.
Published: (2025)
Explicit Correspondence Matching for Generalizable Neural Radiance Fields
by: Chen, Yuedong, et al.
Published: (2023)
by: Chen, Yuedong, et al.
Published: (2023)
Cocktail: Mixing Multi-Modality Controls for Text-Conditional Image Generation
by: Hu, Minghui, et al.
Published: (2023)
by: Hu, Minghui, et al.
Published: (2023)
MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views
by: Chen, Yuedong, et al.
Published: (2024)
by: Chen, Yuedong, et al.
Published: (2024)
3iGS: Factorised Tensorial Illumination for 3D Gaussian Splatting
by: Tang, Zhe Jun, et al.
Published: (2024)
by: Tang, Zhe Jun, et al.
Published: (2024)
MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images
by: Chen, Yuedong, et al.
Published: (2024)
by: Chen, Yuedong, et al.
Published: (2024)
Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level Dynamics
by: Li, Ruining, et al.
Published: (2024)
by: Li, Ruining, et al.
Published: (2024)
GeoConv: Geodesic Guided Convolution for Facial Action Unit Recognition
by: Chen, Yuedong, et al.
Published: (2020)
by: Chen, Yuedong, et al.
Published: (2020)
Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation
by: Hu, Minghui, et al.
Published: (2021)
by: Hu, Minghui, et al.
Published: (2021)
Free3D: Consistent Novel View Synthesis without 3D Representation
by: Zheng, Chuanxia, et al.
Published: (2023)
by: Zheng, Chuanxia, et al.
Published: (2023)
DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness
by: Li, Ruining, et al.
Published: (2025)
by: Li, Ruining, et al.
Published: (2025)
D-LORD for Motion Stylization
by: Gupta, Meenakshi, et al.
Published: (2024)
by: Gupta, Meenakshi, et al.
Published: (2024)
Trajectory Consistency Distillation: Improved Latent Consistency Distillation by Semi-Linear Consistency Function with Trajectory Mapping
by: Zheng, Jianbin, et al.
Published: (2024)
by: Zheng, Jianbin, et al.
Published: (2024)
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
by: Jiang, Zeren, et al.
Published: (2026)
by: Jiang, Zeren, et al.
Published: (2026)
Extending Visual Dynamics for Video-to-Music Generation
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
by: Jiang, Zeren, et al.
Published: (2025)
by: Jiang, Zeren, et al.
Published: (2025)
Measuring 3D Spatial Geometric Consistency in Dynamic Generated Videos
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
Amodal Ground Truth and Completion in the Wild
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
DragAPart: Learning a Part-Level Motion Prior for Articulated Objects
by: Li, Ruining, et al.
Published: (2024)
by: Li, Ruining, et al.
Published: (2024)
Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
by: Lu, Jiahao, et al.
Published: (2024)
by: Lu, Jiahao, et al.
Published: (2024)
Vision Language Models: A Survey of 26K Papers
by: Lin, Fengming
Published: (2025)
by: Lin, Fengming
Published: (2025)
FSDETR: Frequency-Spatial Feature Enhancement for Small Object Detection
by: Huang, Jianchao, et al.
Published: (2026)
by: Huang, Jianchao, et al.
Published: (2026)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
by: Zheng, Xinyi, et al.
Published: (2026)
by: Zheng, Xinyi, et al.
Published: (2026)
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
by: Chen, Weirong, et al.
Published: (2026)
by: Chen, Weirong, et al.
Published: (2026)
Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual Foraging
by: Wang, Bo, et al.
Published: (2024)
by: Wang, Bo, et al.
Published: (2024)
Aligning Anime Video Generation with Human Feedback
by: Zhu, Bingwen, et al.
Published: (2025)
by: Zhu, Bingwen, et al.
Published: (2025)
RefAlign: Representation Alignment for Reference-to-Video Generation
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs
by: Smart, Brandon, et al.
Published: (2024)
by: Smart, Brandon, et al.
Published: (2024)
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
by: He, Dailan, et al.
Published: (2026)
by: He, Dailan, et al.
Published: (2026)
VMBench: A Benchmark for Perception-Aligned Video Motion Generation
by: Ling, Xinran, et al.
Published: (2025)
by: Ling, Xinran, et al.
Published: (2025)
AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation
by: Si, Chen, et al.
Published: (2026)
by: Si, Chen, et al.
Published: (2026)
Similar Items
-
One-shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing
by: Ji, Yuzhu, et al.
Published: (2024) -
PanoDiffusion: 360-degree Panorama Outpainting via Diffusion
by: Wu, Tianhao, et al.
Published: (2023) -
ClusteringSDF: Self-Organized Neural Implicit Surfaces for 3D Decomposition
by: Wu, Tianhao, et al.
Published: (2024) -
Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
by: Wu, Tianhao, et al.
Published: (2025) -
Semantix: An Energy Guided Sampler for Semantic Style Transfer
by: He, Huiang, et al.
Published: (2025)