Lightweight Prompt-Guided CLIP Adaptation for Monocular Depth Estimation
Fuente:
arXiv
Salvato in:
| Autori principali: | Manghotay, Reyhaneh Ahani, Liang, Jie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
di: Bergkvist, Viktor, et al.
Pubblicazione: (2026)
di: Bergkvist, Viktor, et al.
Pubblicazione: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
di: Hou, Zhangcheng, et al.
Pubblicazione: (2026)
di: Hou, Zhangcheng, et al.
Pubblicazione: (2026)
Smooth regularization for efficient video recognition
di: Goldman, Gil, et al.
Pubblicazione: (2025)
di: Goldman, Gil, et al.
Pubblicazione: (2025)
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
di: Wu, Jason, et al.
Pubblicazione: (2026)
di: Wu, Jason, et al.
Pubblicazione: (2026)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
di: Huang, Yian, et al.
Pubblicazione: (2026)
di: Huang, Yian, et al.
Pubblicazione: (2026)
Caption-Driven Explainability: Probing CNNs for Bias via CLIP
di: Koller, Patrick, et al.
Pubblicazione: (2025)
di: Koller, Patrick, et al.
Pubblicazione: (2025)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
di: Boudras, Thomas, et al.
Pubblicazione: (2025)
di: Boudras, Thomas, et al.
Pubblicazione: (2025)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
di: Bartkowiak, Patryk, et al.
Pubblicazione: (2026)
di: Bartkowiak, Patryk, et al.
Pubblicazione: (2026)
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
di: Koroglu, Mathis, et al.
Pubblicazione: (2024)
di: Koroglu, Mathis, et al.
Pubblicazione: (2024)
Perceptual Flow Network for Visually Grounded Reasoning
di: Li, Yangfu, et al.
Pubblicazione: (2026)
di: Li, Yangfu, et al.
Pubblicazione: (2026)
Implementing Adaptations for Vision AutoRegressive Model
di: Shaikh, Kaif, et al.
Pubblicazione: (2025)
di: Shaikh, Kaif, et al.
Pubblicazione: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
di: Durrani, Hamza Ahmed, et al.
Pubblicazione: (2026)
di: Durrani, Hamza Ahmed, et al.
Pubblicazione: (2026)
UnCageNet: Tracking and Pose Estimation of Caged Animal
di: Dutta, Sayak, et al.
Pubblicazione: (2025)
di: Dutta, Sayak, et al.
Pubblicazione: (2025)
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
di: Lin, Xiaojian, et al.
Pubblicazione: (2025)
di: Lin, Xiaojian, et al.
Pubblicazione: (2025)
Car Object Counting and Position Estimation via Extension of the CLIP-EBC Framework
di: Jung, Seoik, et al.
Pubblicazione: (2025)
di: Jung, Seoik, et al.
Pubblicazione: (2025)
Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion
di: Zhu, Yu, et al.
Pubblicazione: (2025)
di: Zhu, Yu, et al.
Pubblicazione: (2025)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
di: Rajapaksha, Uchitha, et al.
Pubblicazione: (2024)
di: Rajapaksha, Uchitha, et al.
Pubblicazione: (2024)
Efficient Attention: Attention with Linear Complexities
di: Shen, Zhuoran, et al.
Pubblicazione: (2018)
di: Shen, Zhuoran, et al.
Pubblicazione: (2018)
Learning 3D object-centric representation through prediction
di: Day, John, et al.
Pubblicazione: (2024)
di: Day, John, et al.
Pubblicazione: (2024)
Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix
di: Gurbindo, Unai, et al.
Pubblicazione: (2025)
di: Gurbindo, Unai, et al.
Pubblicazione: (2025)
SPMamba-YOLO: An Underwater Object Detection Network Based on Multi-Scale Feature Enhancement and Global Context Modeling
di: Liao, Guanghao, et al.
Pubblicazione: (2026)
di: Liao, Guanghao, et al.
Pubblicazione: (2026)
Joint Learning of Depth, Pose, and Local Radiance Field for Large Scale Monocular 3D Reconstruction
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
An Analysis of Layer-Freezing Strategies for Enhanced Transfer Learning in YOLO Architectures
di: Dobrzycki, Andrzej D., et al.
Pubblicazione: (2025)
di: Dobrzycki, Andrzej D., et al.
Pubblicazione: (2025)
In Context Learning with Vision Transformers: Case Study
di: Zhao, Antony, et al.
Pubblicazione: (2025)
di: Zhao, Antony, et al.
Pubblicazione: (2025)
THIRDEYE: Cue-Aware Monocular Depth Estimation via Brain-Inspired Multi-Stage Fusion
di: Ioan, Calin Teodor
Pubblicazione: (2025)
di: Ioan, Calin Teodor
Pubblicazione: (2025)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
di: Shahin, Nada, et al.
Pubblicazione: (2025)
di: Shahin, Nada, et al.
Pubblicazione: (2025)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
di: Wang, Yiming, et al.
Pubblicazione: (2026)
di: Wang, Yiming, et al.
Pubblicazione: (2026)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
di: Brovko, D. V.
Pubblicazione: (2025)
di: Brovko, D. V.
Pubblicazione: (2025)
How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?
di: Costa, Daniel da Silva, et al.
Pubblicazione: (2026)
di: Costa, Daniel da Silva, et al.
Pubblicazione: (2026)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
di: Marian, Vasile, et al.
Pubblicazione: (2026)
di: Marian, Vasile, et al.
Pubblicazione: (2026)
Single-Shot Metric Depth from Focused Plenoptic Cameras
di: Lasheras-Hernandez, Blanca, et al.
Pubblicazione: (2024)
di: Lasheras-Hernandez, Blanca, et al.
Pubblicazione: (2024)
Synthetic-Child: An AIGC-Based Synthetic Data Pipeline for Privacy-Preserving Child Posture Estimation
di: Zeng, Taowen
Pubblicazione: (2026)
di: Zeng, Taowen
Pubblicazione: (2026)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
di: Arnaud, Sergio, et al.
Pubblicazione: (2025)
di: Arnaud, Sergio, et al.
Pubblicazione: (2025)
Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
di: Käs, Stephanie, et al.
Pubblicazione: (2025)
di: Käs, Stephanie, et al.
Pubblicazione: (2025)
Spiking Neural Networks for event-based action recognition: A new task to understand their advantage
di: Vicente-Sola, Alex, et al.
Pubblicazione: (2022)
di: Vicente-Sola, Alex, et al.
Pubblicazione: (2022)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
di: Shahin, Nada, et al.
Pubblicazione: (2025)
di: Shahin, Nada, et al.
Pubblicazione: (2025)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
di: Pu, Qingwen, et al.
Pubblicazione: (2026)
di: Pu, Qingwen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
di: Bergkvist, Viktor, et al.
Pubblicazione: (2026) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025) -
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
di: Hou, Zhangcheng, et al.
Pubblicazione: (2026) -
Smooth regularization for efficient video recognition
di: Goldman, Gil, et al.
Pubblicazione: (2025) -
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
di: Li, Zhuowei, et al.
Pubblicazione: (2025)