VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Du, Hongyang, Ye, Junjie, Cong, Xiaoyan, Li, Runhao, Ni, Jingcheng, Agarwal, Aman, Zhou, Zeqi, Li, Zekun, Balestriero, Randall, Wang, Yue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
por: Patel, Niket, et al.
Publicado: (2025)
por: Patel, Niket, et al.
Publicado: (2025)
PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
por: Rai, Aashish, et al.
Publicado: (2026)
por: Rai, Aashish, et al.
Publicado: (2026)
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
por: Xu, Tian-Xing, et al.
Publicado: (2025)
por: Xu, Tian-Xing, et al.
Publicado: (2025)
Post-Hoc Guidance for Consistency Models by Joint Flow Distribution Learning
por: Hsu, Chia-Hong, et al.
Publicado: (2026)
por: Hsu, Chia-Hong, et al.
Publicado: (2026)
On the Geometry of Deep Learning
por: Balestriero, Randall, et al.
Publicado: (2024)
por: Balestriero, Randall, et al.
Publicado: (2024)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
por: Balestriero, Randall, et al.
Publicado: (2023)
por: Balestriero, Randall, et al.
Publicado: (2023)
GenHSI: Controllable Generation of Human-Scene Interaction Videos
por: Li, Zekun, et al.
Publicado: (2025)
por: Li, Zekun, et al.
Publicado: (2025)
Self-supervised Video Instance Segmentation Can Boost Geographic Entity Alignment in Historical Maps
por: Xia, Xue, et al.
Publicado: (2024)
por: Xia, Xue, et al.
Publicado: (2024)
Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models
por: Peng, Ying, et al.
Publicado: (2025)
por: Peng, Ying, et al.
Publicado: (2025)
Beyond and Free from Diffusion: Invertible Guided Consistency Training
por: Hsu, Chia-Hong, et al.
Publicado: (2025)
por: Hsu, Chia-Hong, et al.
Publicado: (2025)
Fast Multi-view Consistent 3D Editing with Video Priors
por: Chen, Liyi, et al.
Publicado: (2025)
por: Chen, Liyi, et al.
Publicado: (2025)
Learning Temporally Consistent Video Depth from Video Diffusion Priors
por: Shao, Jiahao, et al.
Publicado: (2024)
por: Shao, Jiahao, et al.
Publicado: (2024)
No Location Left Behind: Measuring and Improving the Fairness of Implicit Representations for Earth Data
por: Cai, Daniel, et al.
Publicado: (2025)
por: Cai, Daniel, et al.
Publicado: (2025)
Eidetic Learning: an Efficient and Provable Solution to Catastrophic Forgetting
por: Dronen, Nicholas, et al.
Publicado: (2025)
por: Dronen, Nicholas, et al.
Publicado: (2025)
SAFE: A Novel Approach to AI Weather Evaluation through Stratified Assessments of Forecasts over Earth
por: Masi, Nick, et al.
Publicado: (2025)
por: Masi, Nick, et al.
Publicado: (2025)
ALLoRA: Adaptive Learning Rate Mitigates LoRA Fatal Flaws
por: Huang, Hai, et al.
Publicado: (2024)
por: Huang, Hai, et al.
Publicado: (2024)
GPS-SSL: Guided Positive Sampling to Inject Prior Into Self-Supervised Learning
por: Feizi, Aarash, et al.
Publicado: (2024)
por: Feizi, Aarash, et al.
Publicado: (2024)
SplineCam: Exact Visualization and Characterization of Deep Network Geometry and Decision Boundaries
por: Humayun, Ahmed Imtiaz, et al.
Publicado: (2023)
por: Humayun, Ahmed Imtiaz, et al.
Publicado: (2023)
Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors
por: Wang, Rong, et al.
Publicado: (2026)
por: Wang, Rong, et al.
Publicado: (2026)
Interpreting Physics in Video World Models
por: Joseph, Sonia, et al.
Publicado: (2026)
por: Joseph, Sonia, et al.
Publicado: (2026)
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
por: Zheng, Duo, et al.
Publicado: (2025)
por: Zheng, Duo, et al.
Publicado: (2025)
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
por: Wu, Haoyu, et al.
Publicado: (2025)
por: Wu, Haoyu, et al.
Publicado: (2025)
GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion
por: Kim, Gwanghyun, et al.
Publicado: (2025)
por: Kim, Gwanghyun, et al.
Publicado: (2025)
Fast and Exact Enumeration of Deep Networks Partitions Regions
por: Balestriero, Randall, et al.
Publicado: (2024)
por: Balestriero, Randall, et al.
Publicado: (2024)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
por: Balestriero, Randall, et al.
Publicado: (2025)
por: Balestriero, Randall, et al.
Publicado: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
por: Balestriero, Randall, et al.
Publicado: (2024)
por: Balestriero, Randall, et al.
Publicado: (2024)
Controllable Video Object Insertion via Multiview Priors
por: Qi, Xia, et al.
Publicado: (2026)
por: Qi, Xia, et al.
Publicado: (2026)
LMP: Leveraging Motion Prior in Zero-Shot Video Generation with Diffusion Transformer
por: Chen, Changgu, et al.
Publicado: (2025)
por: Chen, Changgu, et al.
Publicado: (2025)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
por: Ren, Weiming, et al.
Publicado: (2024)
por: Ren, Weiming, et al.
Publicado: (2024)
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
por: Zhang, Runze, et al.
Publicado: (2025)
por: Zhang, Runze, et al.
Publicado: (2025)
VISReg: Variance-Invariance-Sketching Regularization for JEPA training
por: Wu, Haiyu, et al.
Publicado: (2026)
por: Wu, Haiyu, et al.
Publicado: (2026)
Curvature Tuning: Provable Training-free Model Steering From a Single Parameter
por: Hu, Leyang, et al.
Publicado: (2025)
por: Hu, Leyang, et al.
Publicado: (2025)
Self-Supervised Anomaly Detection in the Wild: Favor Joint Embeddings Methods
por: Otero, Daniel, et al.
Publicado: (2024)
por: Otero, Daniel, et al.
Publicado: (2024)
The Fair Language Model Paradox
por: Pinto, Andrea, et al.
Publicado: (2024)
por: Pinto, Andrea, et al.
Publicado: (2024)
Occam's Razor for Self Supervised Learning: What is Sufficient to Learn Good Representations?
por: Ibrahim, Mark, et al.
Publicado: (2024)
por: Ibrahim, Mark, et al.
Publicado: (2024)
GPA resistance dataset
por: Slavenko, Alex, et al.
Publicado: (2026)
por: Slavenko, Alex, et al.
Publicado: (2026)
DC-VSR: Spatially and Temporally Consistent Video Super-Resolution with Video Diffusion Prior
por: Han, Janghyeok, et al.
Publicado: (2025)
por: Han, Janghyeok, et al.
Publicado: (2025)
Towards Long Video Understanding via Fine-detailed Video Story Generation
por: You, Zeng, et al.
Publicado: (2024)
por: You, Zeng, et al.
Publicado: (2024)
Art3D: Training-Free 3D Generation from Flat-Colored Illustration
por: Cong, Xiaoyan, et al.
Publicado: (2025)
por: Cong, Xiaoyan, et al.
Publicado: (2025)
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
por: Fang, Shaoheng, et al.
Publicado: (2025)
por: Fang, Shaoheng, et al.
Publicado: (2025)
Ejemplares similares
-
Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
por: Patel, Niket, et al.
Publicado: (2025) -
PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
por: Rai, Aashish, et al.
Publicado: (2026) -
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
por: Xu, Tian-Xing, et al.
Publicado: (2025) -
Post-Hoc Guidance for Consistency Models by Joint Flow Distribution Learning
por: Hsu, Chia-Hong, et al.
Publicado: (2026) -
On the Geometry of Deep Learning
por: Balestriero, Randall, et al.
Publicado: (2024)