Epipolar Geometry Improves Video Generation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kupyn, Orest, Manhardt, Fabian, Tombari, Federico, Rupprecht, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dataset Enhancement with Instance-Level Augmentations
by: Kupyn, Orest, et al.
Published: (2024)
by: Kupyn, Orest, et al.
Published: (2024)
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
by: Kupyn, Orest, et al.
Published: (2025)
by: Kupyn, Orest, et al.
Published: (2025)
VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
by: Kupyn, Orest, et al.
Published: (2024)
by: Kupyn, Orest, et al.
Published: (2024)
PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models
by: Prospero, Lorenza, et al.
Published: (2026)
by: Prospero, Lorenza, et al.
Published: (2026)
Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas
by: Wimbauer, Felix, et al.
Published: (2026)
by: Wimbauer, Felix, et al.
Published: (2026)
Denoising Diffusion via Image-Based Rendering
by: Anciukevičius, Titas, et al.
Published: (2024)
by: Anciukevičius, Titas, et al.
Published: (2024)
CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation
by: Kalischek, Nikolai, et al.
Published: (2025)
by: Kalischek, Nikolai, et al.
Published: (2025)
A Taxonomy and Library for Visualizing Learned Features in Convolutional Neural Networks
by: Grün, Felix, et al.
Published: (2016)
by: Grün, Felix, et al.
Published: (2016)
OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views
by: Engelmann, Francis, et al.
Published: (2024)
by: Engelmann, Francis, et al.
Published: (2024)
Mixed Diffusion for 3D Indoor Scene Synthesis
by: Hu, Siyi, et al.
Published: (2024)
by: Hu, Siyi, et al.
Published: (2024)
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
by: An, Zhaochong, et al.
Published: (2026)
by: An, Zhaochong, et al.
Published: (2026)
KP-RED: Exploiting Semantic Keypoints for Joint 3D Shape Retrieval and Deformation
by: Zhang, Ruida, et al.
Published: (2024)
by: Zhang, Ruida, et al.
Published: (2024)
3D scene generation from scene graphs and self-attention
by: Bonazzi, Pietro, et al.
Published: (2024)
by: Bonazzi, Pietro, et al.
Published: (2024)
Probing into Camera Control of Video Models
by: Hou, Chen, et al.
Published: (2026)
by: Hou, Chen, et al.
Published: (2026)
LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering
by: Kulhanek, Jonas, et al.
Published: (2025)
by: Kulhanek, Jonas, et al.
Published: (2025)
SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs
by: Zhai, Guangyao, et al.
Published: (2023)
by: Zhai, Guangyao, et al.
Published: (2023)
DAD-3DHeads: A Large-scale Dense, Accurate and Diverse Dataset for 3D Head Alignment from a Single Image
by: Martyniuk, Tetiana, et al.
Published: (2022)
by: Martyniuk, Tetiana, et al.
Published: (2022)
SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation
by: Chen, Yamei, et al.
Published: (2023)
by: Chen, Yamei, et al.
Published: (2023)
D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction
by: Fu, Bowen, et al.
Published: (2023)
by: Fu, Bowen, et al.
Published: (2023)
HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape Generation
by: Leng, Zhiying, et al.
Published: (2024)
by: Leng, Zhiying, et al.
Published: (2024)
CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion
by: Ji, Chenhao, et al.
Published: (2025)
by: Ji, Chenhao, et al.
Published: (2025)
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
by: Chang, Xinhai, et al.
Published: (2024)
by: Chang, Xinhai, et al.
Published: (2024)
MOHO: Learning Single-view Hand-held Object Reconstruction with Multi-view Occlusion-Aware Supervision
by: Zhang, Chenyangguang, et al.
Published: (2023)
by: Zhang, Chenyangguang, et al.
Published: (2023)
Learning Neural Exposure Fields for View Synthesis
by: Niemeyer, Michael, et al.
Published: (2025)
by: Niemeyer, Michael, et al.
Published: (2025)
EpiDiffVO: Geometry-Aware Epipolar Diffusion for Robust Visual Odometry
by: Rao, Prateeth
Published: (2026)
by: Rao, Prateeth
Published: (2026)
EG-Gaussian: Epipolar Geometry and Graph Network Enhanced 3D Gaussian Splatting
by: Zhao, Beizhen, et al.
Published: (2025)
by: Zhao, Beizhen, et al.
Published: (2025)
RadSplat: Radiance Field-Informed Gaussian Splatting for Robust Real-Time Rendering with 900+ FPS
by: Niemeyer, Michael, et al.
Published: (2024)
by: Niemeyer, Michael, et al.
Published: (2024)
RePLAy: Remove Projective LiDAR Depthmap Artifacts via Exploiting Epipolar Geometry
by: Zhu, Shengjie, et al.
Published: (2024)
by: Zhu, Shengjie, et al.
Published: (2024)
Pixel-Accurate Epipolar Guided Matching
by: Nasypanyi, Oleksii, et al.
Published: (2026)
by: Nasypanyi, Oleksii, et al.
Published: (2026)
Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
Text-Conditioned Resampler For Long Form Video Understanding
by: Korbar, Bruno, et al.
Published: (2023)
by: Korbar, Bruno, et al.
Published: (2023)
GeoGaussian: Geometry-aware Gaussian Splatting for Scene Rendering
by: Li, Yanyan, et al.
Published: (2024)
by: Li, Yanyan, et al.
Published: (2024)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
by: Metzger, Nando, et al.
Published: (2025)
by: Metzger, Nando, et al.
Published: (2025)
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
by: Plizzari, Chiara, et al.
Published: (2025)
by: Plizzari, Chiara, et al.
Published: (2025)
Towards Real-Time Open-Vocabulary Video Instance Segmentation
by: Yan, Bin, et al.
Published: (2024)
by: Yan, Bin, et al.
Published: (2024)
LoRACLR: Contrastive Adaptation for Customization of Diffusion Models
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
Epipolar Attention Field Transformers for Bird's Eye View Semantic Segmentation
by: Witte, Christian, et al.
Published: (2024)
by: Witte, Christian, et al.
Published: (2024)
Parallax-Tolerant Image Stitching with Epipolar Displacement Field
by: Yu, Jian, et al.
Published: (2023)
by: Yu, Jian, et al.
Published: (2023)
Similar Items
-
Dataset Enhancement with Instance-Level Augmentations
by: Kupyn, Orest, et al.
Published: (2024) -
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
by: Kupyn, Orest, et al.
Published: (2025) -
VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
by: Kupyn, Orest, et al.
Published: (2024) -
PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models
by: Prospero, Lorenza, et al.
Published: (2026) -
Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas
by: Wimbauer, Felix, et al.
Published: (2026)