M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
Fuente:
arXiv
Saved in:
| Main Authors: | Shvetsova, Nina, Bhat, Goutam, Truong, Prune, Kuehne, Hilde, Tombari, Federico |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
by: Metzger, Nando, et al.
Published: (2025)
by: Metzger, Nando, et al.
Published: (2025)
Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025)
by: Vogel, Felix, et al.
Published: (2025)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
by: Shvetsova, Nina, et al.
Published: (2023)
by: Shvetsova, Nina, et al.
Published: (2023)
One2Any: One-Reference 6D Pose Estimation for Any Object
by: Liu, Mengya, et al.
Published: (2025)
by: Liu, Mengya, et al.
Published: (2025)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
by: Chen, Brian, et al.
Published: (2023)
by: Chen, Brian, et al.
Published: (2023)
TimeLogic: A Temporal Logic Benchmark for Video QA
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
SpatialMe: Stereo Video Conversion Using Depth-Warping and Blend-Inpainting
by: Zhang, Jiale, et al.
Published: (2024)
by: Zhang, Jiale, et al.
Published: (2024)
ImmersePro: End-to-End Stereo Video Synthesis Via Implicit Disparity Learning
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decomposition
by: Hu, Yihan, et al.
Published: (2025)
by: Hu, Yihan, et al.
Published: (2025)
Stitched Value Model for Diffusion Alignment
by: Go, Hyojun, et al.
Published: (2026)
by: Go, Hyojun, et al.
Published: (2026)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
DreamStereo: Towards Real-Time Stereo Inpainting for HD Videos
by: Huang, Yuan, et al.
Published: (2026)
by: Huang, Yuan, et al.
Published: (2026)
Modeling Stereo-Confidence Out of the End-to-End Stereo-Matching Network via Disparity Plane Sweep
by: Lee, Jae Young, et al.
Published: (2024)
by: Lee, Jae Young, et al.
Published: (2024)
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
by: Xing, Ke, et al.
Published: (2025)
by: Xing, Ke, et al.
Published: (2025)
End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
by: Behrens, Tjark, et al.
Published: (2025)
by: Behrens, Tjark, et al.
Published: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
Dynamic Gaussian Marbles for Novel View Synthesis of Casual Monocular Videos
by: Stearns, Colton, et al.
Published: (2024)
by: Stearns, Colton, et al.
Published: (2024)
AnyUp: Universal Feature Upsampling
by: Wimmer, Thomas, et al.
Published: (2025)
by: Wimmer, Thomas, et al.
Published: (2025)
Beyond Imperfections: A Conditional Inpainting Approach for End-to-End Artifact Removal in VTON and Pose Transfer
by: Tabatabaei, Aref, et al.
Published: (2024)
by: Tabatabaei, Aref, et al.
Published: (2024)
Eye2Eye: A Simple Approach for Monocular-to-Stereo Video Synthesis
by: Geyer, Michal, et al.
Published: (2025)
by: Geyer, Michal, et al.
Published: (2025)
Mono2Stereo: Monocular Knowledge Transfer for Enhanced Stereo Matching
by: Wang, Yuran, et al.
Published: (2024)
by: Wang, Yuran, et al.
Published: (2024)
End-to-End Shared Attention Estimation via Group Detection with Feedback Refinement
by: Nakatani, Chihiro, et al.
Published: (2026)
by: Nakatani, Chihiro, et al.
Published: (2026)
CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
by: Xie, Jingchao, et al.
Published: (2025)
by: Xie, Jingchao, et al.
Published: (2025)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)
by: Jahagirdar, Soumya, et al.
Published: (2026)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
by: Huang, Jiaxin, et al.
Published: (2025)
by: Huang, Jiaxin, et al.
Published: (2025)
DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving
by: Yin, Liuhan, et al.
Published: (2025)
by: Yin, Liuhan, et al.
Published: (2025)
MonoGSDF: Exploring Monocular Geometric Cues for Gaussian Splatting-Guided Implicit Surface Reconstruction
by: Li, Kunyi, et al.
Published: (2024)
by: Li, Kunyi, et al.
Published: (2024)
LMVC: An End-to-End Learned Multiview Video Coding Framework
by: Sheng, Xihua, et al.
Published: (2025)
by: Sheng, Xihua, et al.
Published: (2025)
An End-to-End Framework for Video Multi-Person Pose Estimation
by: Wei, Zhihong
Published: (2025)
by: Wei, Zhihong
Published: (2025)
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021)
by: Shi, Fengyuan, et al.
Published: (2021)
PatchRefiner: Leveraging Synthetic Data for Real-Domain High-Resolution Monocular Metric Depth Estimation
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
by: Guo, Yuwei, et al.
Published: (2025)
by: Guo, Yuwei, et al.
Published: (2025)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)
by: Zhang, Jinrong, et al.
Published: (2023)
Boundary-Refined Prototype Generation: A General End-to-End Paradigm for Semi-Supervised Semantic Segmentation
by: Dong, Junhao, et al.
Published: (2023)
by: Dong, Junhao, et al.
Published: (2023)
Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
by: Chen, Kaihua, et al.
Published: (2025)
by: Chen, Kaihua, et al.
Published: (2025)
Similar Items
-
Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
by: Metzger, Nando, et al.
Published: (2025) -
Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
by: Go, Hyojun, et al.
Published: (2025) -
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025) -
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025) -
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
by: Shvetsova, Nina, et al.
Published: (2023)