Large Language Models Can Understanding Depth from Monocular Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xia, Zhongyi, Wu, Tianzhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language-Based Depth Hints for Monocular Depth Estimation
von: Auty, Dylan, et al.
Veröffentlicht: (2024)
von: Auty, Dylan, et al.
Veröffentlicht: (2024)
Advancing Depth Anything Model for Unsupervised Monocular Depth Estimation in Endoscopy
von: Li, Bojian, et al.
Veröffentlicht: (2024)
von: Li, Bojian, et al.
Veröffentlicht: (2024)
Focusable Monocular Depth Estimation
von: Du, Yuxin, et al.
Veröffentlicht: (2026)
von: Du, Yuxin, et al.
Veröffentlicht: (2026)
BRIDGE -- Building Reinforcement-Learning Depth-to-Image Data Generation Engine for Monocular Depth Estimation
von: Liu, Dingning, et al.
Veröffentlicht: (2025)
von: Liu, Dingning, et al.
Veröffentlicht: (2025)
Can Multimodal Large Language Models Truly Understand Small Objects?
von: Han, Fujun, et al.
Veröffentlicht: (2026)
von: Han, Fujun, et al.
Veröffentlicht: (2026)
Shedding Light on Depth: Explainability Assessment in Monocular Depth Estimation
von: Cirillo, Lorenzo, et al.
Veröffentlicht: (2025)
von: Cirillo, Lorenzo, et al.
Veröffentlicht: (2025)
DepthDark: Robust Monocular Depth Estimation for Low-Light Environments
von: Zeng, Longjian, et al.
Veröffentlicht: (2025)
von: Zeng, Longjian, et al.
Veröffentlicht: (2025)
CLIP Can Understand Depth
von: Kim, Sohee, et al.
Veröffentlicht: (2024)
von: Kim, Sohee, et al.
Veröffentlicht: (2024)
Always Clear Depth: Robust Monocular Depth Estimation under Adverse Weather
von: Jiang, Kui, et al.
Veröffentlicht: (2025)
von: Jiang, Kui, et al.
Veröffentlicht: (2025)
See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
von: Yang, Kunyi, et al.
Veröffentlicht: (2025)
von: Yang, Kunyi, et al.
Veröffentlicht: (2025)
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
von: Chen, Pingyi, et al.
Veröffentlicht: (2025)
von: Chen, Pingyi, et al.
Veröffentlicht: (2025)
SM4Depth: Seamless Monocular Metric Depth Estimation across Multiple Cameras and Scenes by One Model
von: Liu, Yihao, et al.
Veröffentlicht: (2024)
von: Liu, Yihao, et al.
Veröffentlicht: (2024)
WorDepth: Variational Language Prior for Monocular Depth Estimation
von: Zeng, Ziyao, et al.
Veröffentlicht: (2024)
von: Zeng, Ziyao, et al.
Veröffentlicht: (2024)
ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth Estimation
von: Patni, Suraj, et al.
Veröffentlicht: (2024)
von: Patni, Suraj, et al.
Veröffentlicht: (2024)
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology
von: Zhang, Lina, et al.
Veröffentlicht: (2026)
von: Zhang, Lina, et al.
Veröffentlicht: (2026)
AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
AuxDepthNet: Real-Time Monocular 3D Object Detection with Depth-Sensitive Features
von: Zhang, Ruochen, et al.
Veröffentlicht: (2025)
von: Zhang, Ruochen, et al.
Veröffentlicht: (2025)
UMono: Physical Model Informed Hybrid CNN-Transformer Framework for Underwater Monocular Depth Estimation
von: Wang, Jian, et al.
Veröffentlicht: (2024)
von: Wang, Jian, et al.
Veröffentlicht: (2024)
Real-time Accident Anticipation for Autonomous Driving Through Monocular Depth-Enhanced 3D Modeling
von: Liao, Haicheng, et al.
Veröffentlicht: (2024)
von: Liao, Haicheng, et al.
Veröffentlicht: (2024)
Can Vision-Language Models Understand Construction Workers? An Exploratory Study
von: Bui, Hieu, et al.
Veröffentlicht: (2026)
von: Bui, Hieu, et al.
Veröffentlicht: (2026)
Deep Neighbor Layer Aggregation for Lightweight Self-Supervised Monocular Depth Estimation
von: Boya, Wang, et al.
Veröffentlicht: (2023)
von: Boya, Wang, et al.
Veröffentlicht: (2023)
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
von: Li, Qingmei, et al.
Veröffentlicht: (2025)
von: Li, Qingmei, et al.
Veröffentlicht: (2025)
Can Vision Language Models Understand Mimed Actions?
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
Can Large Language Models Understand Symbolic Graphics Programs?
von: Qiu, Zeju, et al.
Veröffentlicht: (2024)
von: Qiu, Zeju, et al.
Veröffentlicht: (2024)
Can Large Vision-Language Models Detect Images Copyright Infringement from GenAI?
von: Xu, Qipan, et al.
Veröffentlicht: (2025)
von: Xu, Qipan, et al.
Veröffentlicht: (2025)
Review of Hallucination Understanding in Large Language and Vision Models
von: Ho, Zhengyi, et al.
Veröffentlicht: (2025)
von: Ho, Zhengyi, et al.
Veröffentlicht: (2025)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
von: Li, Xiujun, et al.
Veröffentlicht: (2023)
von: Li, Xiujun, et al.
Veröffentlicht: (2023)
On Large Visual Language Models for Medical Imaging Analysis: An Empirical Study
von: Van, Minh-Hao, et al.
Veröffentlicht: (2024)
von: Van, Minh-Hao, et al.
Veröffentlicht: (2024)
Adaptive Discrete Disparity Volume for Self-supervised Monocular Depth Estimation
von: Ren, Jianwei
Veröffentlicht: (2024)
von: Ren, Jianwei
Veröffentlicht: (2024)
MOSABench: Multi-Object Sentiment Analysis Benchmark for Evaluating Multimodal Large Language Models Understanding of Complex Image
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
A Critical Synthesis of Uncertainty Quantification and Foundation Models in Monocular Depth Estimation
von: Landgraf, Steven, et al.
Veröffentlicht: (2025)
von: Landgraf, Steven, et al.
Veröffentlicht: (2025)
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
von: He, Zhihao, et al.
Veröffentlicht: (2026)
von: He, Zhihao, et al.
Veröffentlicht: (2026)
EndoGMDE: Generalizable Monocular Depth Estimation with Mixture of Low-Rank Experts for Diverse Endoscopic Scenes
von: Shao, Liangjing, et al.
Veröffentlicht: (2025)
von: Shao, Liangjing, et al.
Veröffentlicht: (2025)
No Pose Estimation? No Problem: Pose-Agnostic and Instance-Aware Test-Time Adaptation for Monocular Depth Estimation
von: Sung, Mingyu, et al.
Veröffentlicht: (2025)
von: Sung, Mingyu, et al.
Veröffentlicht: (2025)
$D^3$-RSMDE: 40$\times$ Faster and High-Fidelity Remote Sensing Monocular Depth Estimation
von: Wang, Ruizhi, et al.
Veröffentlicht: (2026)
von: Wang, Ruizhi, et al.
Veröffentlicht: (2026)
S3MOT: Monocular 3D Object Tracking with Selective State Space Model
von: Yan, Zhuohao, et al.
Veröffentlicht: (2025)
von: Yan, Zhuohao, et al.
Veröffentlicht: (2025)
Scalable Cloud-Native Pipeline for Efficient 3D Model Reconstruction from Monocular Smartphone Images
von: Aghilar, Potito, et al.
Veröffentlicht: (2024)
von: Aghilar, Potito, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Language-Based Depth Hints for Monocular Depth Estimation
von: Auty, Dylan, et al.
Veröffentlicht: (2024) -
Advancing Depth Anything Model for Unsupervised Monocular Depth Estimation in Endoscopy
von: Li, Bojian, et al.
Veröffentlicht: (2024) -
Focusable Monocular Depth Estimation
von: Du, Yuxin, et al.
Veröffentlicht: (2026) -
BRIDGE -- Building Reinforcement-Learning Depth-to-Image Data Generation Engine for Monocular Depth Estimation
von: Liu, Dingning, et al.
Veröffentlicht: (2025) -
Can Multimodal Large Language Models Truly Understand Small Objects?
von: Han, Fujun, et al.
Veröffentlicht: (2026)