MetricGold: Leveraging Text-To-Image Latent Diffusion Models for Metric Depth Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Shah, Ansh, Krishna, K Madhava |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey on Quality Metrics for Text-to-Image Generation
by: Hartwig, Sebastian, et al.
Published: (2024)
by: Hartwig, Sebastian, et al.
Published: (2024)
DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion
by: Ye, Weicai, et al.
Published: (2024)
by: Ye, Weicai, et al.
Published: (2024)
Leveraging Foundation Models To learn the shape of semi-fluid deformable objects
by: Assal, Omar El, et al.
Published: (2024)
by: Assal, Omar El, et al.
Published: (2024)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
by: Ray, Arijit, et al.
Published: (2024)
by: Ray, Arijit, et al.
Published: (2024)
Infinite Leagues Under the Sea: Photorealistic 3D Underwater Terrain Generation by Latent Fractal Diffusion Models
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Free-Moving Object Reconstruction and Pose Estimation with Virtual Camera
by: Shi, Haixin, et al.
Published: (2024)
by: Shi, Haixin, et al.
Published: (2024)
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
by: Yang, Jianing, et al.
Published: (2025)
by: Yang, Jianing, et al.
Published: (2025)
Evaluating Design Video Generation: Metrics for Compositional Fidelity
by: Deganutti, Adrienne, et al.
Published: (2026)
by: Deganutti, Adrienne, et al.
Published: (2026)
Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera
by: Guo, Yuliang, et al.
Published: (2025)
by: Guo, Yuliang, et al.
Published: (2025)
Go-SLAM: Grounded Object Segmentation and Localization with Gaussian Splatting SLAM
by: Pham, Phu, et al.
Published: (2024)
by: Pham, Phu, et al.
Published: (2024)
Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects
by: Weng, Yijia, et al.
Published: (2024)
by: Weng, Yijia, et al.
Published: (2024)
AddBiomechanics Dataset: Capturing the Physics of Human Motion at Scale
by: Werling, Keenon, et al.
Published: (2024)
by: Werling, Keenon, et al.
Published: (2024)
ImDy: Human Inverse Dynamics from Imitated Observations
by: Liu, Xinpeng, et al.
Published: (2024)
by: Liu, Xinpeng, et al.
Published: (2024)
Generalizable Vision-Language Few-Shot Adaptation with Predictive Prompts and Negative Learning
by: Mandalika, Sriram
Published: (2025)
by: Mandalika, Sriram
Published: (2025)
Physics-Based Motion Imitation with Adversarial Differential Discriminators
by: Zhang, Ziyu, et al.
Published: (2025)
by: Zhang, Ziyu, et al.
Published: (2025)
Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset
by: Dong, Zhao, et al.
Published: (2025)
by: Dong, Zhao, et al.
Published: (2025)
SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes
by: Pfaff, Nicholas, et al.
Published: (2026)
by: Pfaff, Nicholas, et al.
Published: (2026)
VibraVerse: A Large-Scale Geometry-Acoustics Alignment Dataset for Physically-Consistent Multimodal Learning
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
SMP: Reusable Score-Matching Motion Priors for Physics-Based Character Control
by: Mu, Yuxuan, et al.
Published: (2025)
by: Mu, Yuxuan, et al.
Published: (2025)
Icy Moon Surface Simulation and Stereo Depth Estimation for Sampling Autonomy
by: Bhaskara, Ramchander, et al.
Published: (2024)
by: Bhaskara, Ramchander, et al.
Published: (2024)
LEAD: Latent Realignment for Human Motion Diffusion
by: Andreou, Nefeli, et al.
Published: (2024)
by: Andreou, Nefeli, et al.
Published: (2024)
Hollowed Net for On-Device Personalization of Text-to-Image Diffusion Models
by: Cho, Wonguk, et al.
Published: (2024)
by: Cho, Wonguk, et al.
Published: (2024)
SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing
by: Hong, Seokhyeon, et al.
Published: (2025)
by: Hong, Seokhyeon, et al.
Published: (2025)
RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffusion
by: Shriram, Jaidev, et al.
Published: (2024)
by: Shriram, Jaidev, et al.
Published: (2024)
MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control
by: Li, Bin, et al.
Published: (2026)
by: Li, Bin, et al.
Published: (2026)
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
by: Zafar, Oz, et al.
Published: (2024)
by: Zafar, Oz, et al.
Published: (2024)
Enhanced Controllability of Diffusion Models via Feature Disentanglement and Realism-Enhanced Sampling Methods
by: Cho, Wonwoong, et al.
Published: (2023)
by: Cho, Wonwoong, et al.
Published: (2023)
GENMO: A GENeralist Model for Human MOtion
by: Li, Jiefeng, et al.
Published: (2025)
by: Li, Jiefeng, et al.
Published: (2025)
Controlling Text-to-Image Diffusion by Orthogonal Finetuning
by: Qiu, Zeju, et al.
Published: (2023)
by: Qiu, Zeju, et al.
Published: (2023)
Learning Latent Representations for Image Translation using Frequency Distributed CycleGAN
by: Nigam, Shivangi, et al.
Published: (2025)
by: Nigam, Shivangi, et al.
Published: (2025)
ReLumix: Extending Image Relighting to Video via Video Diffusion Models
by: Wang, Lezhong, et al.
Published: (2025)
by: Wang, Lezhong, et al.
Published: (2025)
Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models
by: Liu, Bingchen, et al.
Published: (2024)
by: Liu, Bingchen, et al.
Published: (2024)
LuxDiT: Lighting Estimation with Video Diffusion Transformer
by: Liang, Ruofan, et al.
Published: (2025)
by: Liang, Ruofan, et al.
Published: (2025)
Spatiotemporally Consistent Indoor Lighting Estimation with Diffusion Priors
by: Tong, Mutian, et al.
Published: (2025)
by: Tong, Mutian, et al.
Published: (2025)
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
by: Yin, Zixin, et al.
Published: (2025)
by: Yin, Zixin, et al.
Published: (2025)
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
by: Hu, Wenbo, et al.
Published: (2024)
by: Hu, Wenbo, et al.
Published: (2024)
Lazy Diffusion Transformer for Interactive Image Editing
by: Nitzan, Yotam, et al.
Published: (2024)
by: Nitzan, Yotam, et al.
Published: (2024)
Make It Count: Text-to-Image Generation with an Accurate Number of Objects
by: Binyamin, Lital, et al.
Published: (2024)
by: Binyamin, Lital, et al.
Published: (2024)
Key-Locked Rank One Editing for Text-to-Image Personalization
by: Tewel, Yoad, et al.
Published: (2023)
by: Tewel, Yoad, et al.
Published: (2023)
BootPIG: Bootstrapping Zero-shot Personalized Image Generation Capabilities in Pretrained Diffusion Models
by: Purushwalkam, Senthil, et al.
Published: (2024)
by: Purushwalkam, Senthil, et al.
Published: (2024)
Similar Items
-
A Survey on Quality Metrics for Text-to-Image Generation
by: Hartwig, Sebastian, et al.
Published: (2024) -
DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion
by: Ye, Weicai, et al.
Published: (2024) -
Leveraging Foundation Models To learn the shape of semi-fluid deformable objects
by: Assal, Omar El, et al.
Published: (2024) -
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
by: Ray, Arijit, et al.
Published: (2024) -
Infinite Leagues Under the Sea: Photorealistic 3D Underwater Terrain Generation by Latent Fractal Diffusion Models
by: Zhang, Tianyi, et al.
Published: (2025)