Depth Anything V2
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Lihe, Kang, Bingyi, Huang, Zilong, Zhao, Zhen, Xu, Xiaogang, Feng, Jiashi, Zhao, Hengshuang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025)
by: Chen, Sili, et al.
Published: (2025)
UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Depth Anything with Any Prior
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
Classification Done Right for Vision-Language Pre-Training
by: Huang, Zilong, et al.
Published: (2024)
by: Huang, Zilong, et al.
Published: (2024)
Depth Anything 3: Recovering the Visual Space from Any Views
by: Lin, Haotong, et al.
Published: (2025)
by: Lin, Haotong, et al.
Published: (2025)
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
by: Lin, Haotong, et al.
Published: (2024)
by: Lin, Haotong, et al.
Published: (2024)
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
SuperCLIP: CLIP with Simple Classification Supervision
by: Zhao, Weiheng, et al.
Published: (2025)
by: Zhao, Weiheng, et al.
Published: (2025)
Image Understanding Makes for A Good Tokenizer for Image Generation
by: Wang, Luting, et al.
Published: (2024)
by: Wang, Luting, et al.
Published: (2024)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
A Lightweight Clustering Framework for Unsupervised Semantic Segmentation
by: Cheung, Yau Shing Jonathan, et al.
Published: (2023)
by: Cheung, Yau Shing Jonathan, et al.
Published: (2023)
PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial Augmentation
by: Cao, Zidong, et al.
Published: (2024)
by: Cao, Zidong, et al.
Published: (2024)
Towards Unified 3D Object Detection via Algorithm and Data Unification
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
by: Ren, Zhongwei, et al.
Published: (2025)
by: Ren, Zhongwei, et al.
Published: (2025)
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
by: Ren, Zhongwei, et al.
Published: (2026)
by: Ren, Zhongwei, et al.
Published: (2026)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
How Far is Video Generation from World Model: A Physical Law Perspective
by: Kang, Bingyi, et al.
Published: (2024)
by: Kang, Bingyi, et al.
Published: (2024)
DiffCamera: Arbitrary Refocusing on Images
by: Wang, Yiyang, et al.
Published: (2025)
by: Wang, Yiyang, et al.
Published: (2025)
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
In Pursuit of Pixel Supervision for Visual Pre-training
by: Yang, Lihe, et al.
Published: (2025)
by: Yang, Lihe, et al.
Published: (2025)
FashionComposer: Compositional Fashion Image Generation
by: Ji, Sihui, et al.
Published: (2024)
by: Ji, Sihui, et al.
Published: (2024)
Orient Anything V2: Unifying Orientation and Rotation Understanding
by: Wang, Zehan, et al.
Published: (2026)
by: Wang, Zehan, et al.
Published: (2026)
GDRO: Group-level Reward Post-training Suitable for Diffusion Models
by: Wang, Yiyang, et al.
Published: (2026)
by: Wang, Yiyang, et al.
Published: (2026)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
by: Xiong, Tianwei, et al.
Published: (2025)
by: Xiong, Tianwei, et al.
Published: (2025)
DA$^{2}$: Depth Anything in Any Direction
by: Li, Haodong, et al.
Published: (2025)
by: Li, Haodong, et al.
Published: (2025)
4th PVUW MeViS 3rd Place Report: Sa2VA
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
DiffDoctor: Diagnosing Image Diffusion Models Before Treating
by: Wang, Yiyang, et al.
Published: (2025)
by: Wang, Yiyang, et al.
Published: (2025)
MESA: Matching Everything by Segmenting Anything
by: Zhang, Yesheng, et al.
Published: (2024)
by: Zhang, Yesheng, et al.
Published: (2024)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
by: Xiong, Tianwei, et al.
Published: (2026)
by: Xiong, Tianwei, et al.
Published: (2026)
Exploring Deeper! Segment Anything Model with Depth Perception for Camouflaged Object Detection
by: Yu, Zhenni, et al.
Published: (2024)
by: Yu, Zhenni, et al.
Published: (2024)
Amodal Depth Anything: Amodal Depth Estimation in the Wild
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Depth Anything in $360^\circ$: Towards Scale Invariance in the Wild
by: Jiang, Hualie, et al.
Published: (2025)
by: Jiang, Hualie, et al.
Published: (2025)
Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding
by: Yuan, Zhihao, et al.
Published: (2023)
by: Yuan, Zhihao, et al.
Published: (2023)
S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything without Supervision
by: Xu, Huihui, et al.
Published: (2025)
by: Xu, Huihui, et al.
Published: (2025)
Any to Full: Prompting Depth Anything for Depth Completion in One Stage
by: Zhou, Zhiyuan, et al.
Published: (2026)
by: Zhou, Zhiyuan, et al.
Published: (2026)
The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
by: Lei, Weixian, et al.
Published: (2025)
by: Lei, Weixian, et al.
Published: (2025)
Similar Items
-
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
by: Yang, Lihe, et al.
Published: (2024) -
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025) -
UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation
by: Yang, Lihe, et al.
Published: (2024) -
Depth Anything with Any Prior
by: Wang, Zehan, et al.
Published: (2025) -
Classification Done Right for Vision-Language Pre-Training
by: Huang, Zilong, et al.
Published: (2024)