View-Consistent Diffusion Representations for 3D-Consistent Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Danier, Duolikun, Gao, Ge, McDonagh, Steven, Li, Changjian, Bilen, Hakan, Mac Aodha, Oisin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
by: Danier, Duolikun, et al.
Published: (2024)
by: Danier, Duolikun, et al.
Published: (2024)
Improving Semantic Correspondence with Viewpoint-Guided Spherical Maps
by: Mariotti, Octave, et al.
Published: (2023)
by: Mariotti, Octave, et al.
Published: (2023)
MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
by: Du, Zhipeng, et al.
Published: (2025)
by: Du, Zhipeng, et al.
Published: (2025)
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
by: Tsagkas, Nikolaos, et al.
Published: (2025)
by: Tsagkas, Nikolaos, et al.
Published: (2025)
Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence
by: Mariotti, Octave, et al.
Published: (2025)
by: Mariotti, Octave, et al.
Published: (2025)
SAOR: Single-View Articulated Object Reconstruction
by: Aygün, Mehmet, et al.
Published: (2023)
by: Aygün, Mehmet, et al.
Published: (2023)
Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
by: Tsagkas, Nikolaos, et al.
Published: (2025)
by: Tsagkas, Nikolaos, et al.
Published: (2025)
Looking 3D: Anomaly Detection with 2D-3D Alignment
by: Bhunia, Ankan, et al.
Published: (2024)
by: Bhunia, Ankan, et al.
Published: (2024)
LDMVFI: Video Frame Interpolation with Latent Diffusion Models
by: Danier, Duolikun, et al.
Published: (2023)
by: Danier, Duolikun, et al.
Published: (2023)
Enhancing Deformable Convolution based Video Frame Interpolation with Coarse-to-fine 3D CNN
by: Danier, Duolikun, et al.
Published: (2022)
by: Danier, Duolikun, et al.
Published: (2022)
BVI-VFI: A Video Quality Database for Video Frame Interpolation
by: Danier, Duolikun, et al.
Published: (2022)
by: Danier, Duolikun, et al.
Published: (2022)
AirPlanes: Accurate Plane Estimation via 3D-Consistent Embeddings
by: Watson, Jamie, et al.
Published: (2024)
by: Watson, Jamie, et al.
Published: (2024)
Enhancing 2D Representation Learning with a 3D Prior
by: Aygün, Mehmet, et al.
Published: (2024)
by: Aygün, Mehmet, et al.
Published: (2024)
Interpretable Text-Guided Image Clustering via Iterative Search
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
A Subjective Quality Study for Video Frame Interpolation
by: Danier, Duolikun, et al.
Published: (2022)
by: Danier, Duolikun, et al.
Published: (2022)
Beyond Pixel Histories: World Models with Persistent 3D State
by: Garcin, Samuel, et al.
Published: (2026)
by: Garcin, Samuel, et al.
Published: (2026)
Odd-One-Out: Anomaly Detection by Comparing with Neighbors
by: Bhunia, Ankan, et al.
Published: (2024)
by: Bhunia, Ankan, et al.
Published: (2024)
CrossSDF: 3D Reconstruction of Thin Structures From Cross-Sections
by: Walker, Thomas, et al.
Published: (2024)
by: Walker, Thomas, et al.
Published: (2024)
HumMorph: Generalized Dynamic Human Neural Fields from Few Views
by: Zadrożny, Jakub, et al.
Published: (2025)
by: Zadrożny, Jakub, et al.
Published: (2025)
Representational Similarity via Interpretable Visual Concepts
by: Kondapaneni, Neehar, et al.
Published: (2025)
by: Kondapaneni, Neehar, et al.
Published: (2025)
Representational Difference Explanations
by: Kondapaneni, Neehar, et al.
Published: (2025)
by: Kondapaneni, Neehar, et al.
Published: (2025)
GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting
by: Agarwal, Madhav, et al.
Published: (2025)
by: Agarwal, Madhav, et al.
Published: (2025)
VesselSDF: Distance Field Priors for Vascular Network Reconstruction
by: Esposito, Salvatore, et al.
Published: (2025)
by: Esposito, Salvatore, et al.
Published: (2025)
ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation
by: Danier, Duolikun, et al.
Published: (2021)
by: Danier, Duolikun, et al.
Published: (2021)
GFix: Perceptually Enhanced Gaussian Splatting Video Compression
by: Teng, Siyue, et al.
Published: (2025)
by: Teng, Siyue, et al.
Published: (2025)
Improving Object Detection via Local-global Contrastive Learning
by: Triantafyllidou, Danai, et al.
Published: (2024)
by: Triantafyllidou, Danai, et al.
Published: (2024)
BVI-Artefact: An Artefact Detection Benchmark Dataset for Streamed Videos
by: Feng, Chen, et al.
Published: (2023)
by: Feng, Chen, et al.
Published: (2023)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
by: Stogiannidis, Ilias, et al.
Published: (2025)
by: Stogiannidis, Ilias, et al.
Published: (2025)
Less is More: Discovering Concise Network Explanations
by: Kondapaneni, Neehar, et al.
Published: (2024)
by: Kondapaneni, Neehar, et al.
Published: (2024)
Labeled Data Selection for Category Discovery
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
Generating Binary Species Range Maps
by: Dorm, Filip, et al.
Published: (2024)
by: Dorm, Filip, et al.
Published: (2024)
Concept-based Adversarial Attack: a Probabilistic Perspective
by: Zhang, Andi, et al.
Published: (2025)
by: Zhang, Andi, et al.
Published: (2025)
3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement
by: Luo, Yihang, et al.
Published: (2024)
by: Luo, Yihang, et al.
Published: (2024)
Label-Efficient Object Detection via Region Proposal Network Pre-Training
by: Dong, Nanqing, et al.
Published: (2022)
by: Dong, Nanqing, et al.
Published: (2022)
Consistent-1-to-3: Consistent Image to 3D View Synthesis via Geometry-aware Diffusion Models
by: Ye, Jianglong, et al.
Published: (2023)
by: Ye, Jianglong, et al.
Published: (2023)
Free3D: Consistent Novel View Synthesis without 3D Representation
by: Zheng, Chuanxia, et al.
Published: (2023)
by: Zheng, Chuanxia, et al.
Published: (2023)
MVAD: A Multiple Visual Artifact Detector for Video Streaming
by: Feng, Chen, et al.
Published: (2024)
by: Feng, Chen, et al.
Published: (2024)
Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors
by: Tsagkas, Nikolaos, et al.
Published: (2024)
by: Tsagkas, Nikolaos, et al.
Published: (2024)
MotionPhysics: Learnable Motion Distillation for Text-Guided Simulation
by: Wang, Miaowei, et al.
Published: (2026)
by: Wang, Miaowei, et al.
Published: (2026)
No time to train! Training-Free Reference-Based Instance Segmentation
by: Espinosa, Miguel, et al.
Published: (2025)
by: Espinosa, Miguel, et al.
Published: (2025)
Similar Items
-
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
by: Danier, Duolikun, et al.
Published: (2024) -
Improving Semantic Correspondence with Viewpoint-Guided Spherical Maps
by: Mariotti, Octave, et al.
Published: (2023) -
MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
by: Du, Zhipeng, et al.
Published: (2025) -
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
by: Tsagkas, Nikolaos, et al.
Published: (2025) -
Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence
by: Mariotti, Octave, et al.
Published: (2025)