Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Lihe, Kang, Bingyi, Huang, Zilong, Xu, Xiaogang, Feng, Jiashi, Zhao, Hengshuang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Depth Anything V2
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025)
by: Chen, Sili, et al.
Published: (2025)
Depth Anything with Any Prior
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial Augmentation
by: Cao, Zidong, et al.
Published: (2024)
by: Cao, Zidong, et al.
Published: (2024)
Classification Done Right for Vision-Language Pre-Training
by: Huang, Zilong, et al.
Published: (2024)
by: Huang, Zilong, et al.
Published: (2024)
Depth Anything 3: Recovering the Visual Space from Any Views
by: Lin, Haotong, et al.
Published: (2025)
by: Lin, Haotong, et al.
Published: (2025)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
by: Ren, Zhongwei, et al.
Published: (2025)
by: Ren, Zhongwei, et al.
Published: (2025)
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
by: Lin, Haotong, et al.
Published: (2024)
by: Lin, Haotong, et al.
Published: (2024)
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
Distinguish Any Fake Videos: Unleashing the Power of Large-scale Data and Motion Features
by: Ji, Lichuan, et al.
Published: (2024)
by: Ji, Lichuan, et al.
Published: (2024)
Image Understanding Makes for A Good Tokenizer for Image Generation
by: Wang, Luting, et al.
Published: (2024)
by: Wang, Luting, et al.
Published: (2024)
SuperCLIP: CLIP with Simple Classification Supervision
by: Zhao, Weiheng, et al.
Published: (2025)
by: Zhao, Weiheng, et al.
Published: (2025)
Towards Unified 3D Object Detection via Algorithm and Data Unification
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
by: Xiong, Tianwei, et al.
Published: (2025)
by: Xiong, Tianwei, et al.
Published: (2025)
A Lightweight Clustering Framework for Unsupervised Semantic Segmentation
by: Cheung, Yau Shing Jonathan, et al.
Published: (2023)
by: Cheung, Yau Shing Jonathan, et al.
Published: (2023)
Unleashing Unlabeled Data: A Paradigm for Cross-View Geo-Localization
by: Li, Guopeng, et al.
Published: (2024)
by: Li, Guopeng, et al.
Published: (2024)
LMSeg: Unleashing the Power of Large-Scale Models for Open-Vocabulary Semantic Segmentation
by: Tang, Huadong, et al.
Published: (2024)
by: Tang, Huadong, et al.
Published: (2024)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
How Far is Video Generation from World Model: A Physical Law Perspective
by: Kang, Bingyi, et al.
Published: (2024)
by: Kang, Bingyi, et al.
Published: (2024)
Depth Anything in $360^\circ$: Towards Scale Invariance in the Wild
by: Jiang, Hualie, et al.
Published: (2025)
by: Jiang, Hualie, et al.
Published: (2025)
DiffCamera: Arbitrary Refocusing on Images
by: Wang, Yiyang, et al.
Published: (2025)
by: Wang, Yiyang, et al.
Published: (2025)
Unlock the Power of Unlabeled Data in Language Driving Model
by: Wang, Chaoqun, et al.
Published: (2025)
by: Wang, Chaoqun, et al.
Published: (2025)
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
by: Ren, Zhongwei, et al.
Published: (2026)
by: Ren, Zhongwei, et al.
Published: (2026)
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
In Pursuit of Pixel Supervision for Visual Pre-training
by: Yang, Lihe, et al.
Published: (2025)
by: Yang, Lihe, et al.
Published: (2025)
FashionComposer: Compositional Fashion Image Generation
by: Ji, Sihui, et al.
Published: (2024)
by: Ji, Sihui, et al.
Published: (2024)
Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data Augmentation
by: Wang, Ning-Hsu, et al.
Published: (2024)
by: Wang, Ning-Hsu, et al.
Published: (2024)
GDRO: Group-level Reward Post-training Suitable for Diffusion Models
by: Wang, Yiyang, et al.
Published: (2026)
by: Wang, Yiyang, et al.
Published: (2026)
Stereo Anything: Unifying Zero-shot Stereo Matching with Large-Scale Mixed Data
by: Guo, Xianda, et al.
Published: (2024)
by: Guo, Xianda, et al.
Published: (2024)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
by: Xiong, Tianwei, et al.
Published: (2026)
by: Xiong, Tianwei, et al.
Published: (2026)
Unleashing the Power of Motion and Depth: A Selective Fusion Strategy for RGB-D Video Salient Object Detection
by: He, Jiahao, et al.
Published: (2025)
by: He, Jiahao, et al.
Published: (2025)
4th PVUW MeViS 3rd Place Report: Sa2VA
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
Amodal Depth Anything: Amodal Depth Estimation in the Wild
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
DA$^{2}$: Depth Anything in Any Direction
by: Li, Haodong, et al.
Published: (2025)
by: Li, Haodong, et al.
Published: (2025)
MoonAnything: A Vision Benchmark with Large-Scale Lunar Supervised Data
by: Grethen, Clémentine, et al.
Published: (2026)
by: Grethen, Clémentine, et al.
Published: (2026)
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
by: Ma, Baorui, et al.
Published: (2026)
by: Ma, Baorui, et al.
Published: (2026)
RigAnyFace: Scaling Neural Facial Mesh Auto-Rigging with Unlabeled Data
by: Ma, Wenchao, et al.
Published: (2025)
by: Ma, Wenchao, et al.
Published: (2025)
Similar Items
-
Depth Anything V2
by: Yang, Lihe, et al.
Published: (2024) -
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025) -
Depth Anything with Any Prior
by: Wang, Zehan, et al.
Published: (2025) -
PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial Augmentation
by: Cao, Zidong, et al.
Published: (2024) -
Classification Done Right for Vision-Language Pre-Training
by: Huang, Zilong, et al.
Published: (2024)