On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chatterjee, Agneet, Gokhale, Tejas, Baral, Chitta, Yang, Yezhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024)
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
von: Malaviya, Vatsal, et al.
Veröffentlicht: (2025)
von: Malaviya, Vatsal, et al.
Veröffentlicht: (2025)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
von: Patel, Maitreya, et al.
Veröffentlicht: (2023)
von: Patel, Maitreya, et al.
Veröffentlicht: (2023)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
Dual Caption Preference Optimization for Diffusion Models
von: Saeidi, Amir, et al.
Veröffentlicht: (2025)
von: Saeidi, Amir, et al.
Veröffentlicht: (2025)
Chimera: Compositional Image Generation using Part-based Concepting
von: Singh, Shivam, et al.
Veröffentlicht: (2025)
von: Singh, Shivam, et al.
Veröffentlicht: (2025)
Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
von: Luo, Yiran, et al.
Veröffentlicht: (2024)
von: Luo, Yiran, et al.
Veröffentlicht: (2024)
Getting it Right: Improving Spatial Consistency in Text-to-Image Models
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024)
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024)
Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2025)
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
von: Pathiraja, Bimsara, et al.
Veröffentlicht: (2025)
von: Pathiraja, Bimsara, et al.
Veröffentlicht: (2025)
The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
von: Anvekar, Tejas, et al.
Veröffentlicht: (2025)
von: Anvekar, Tejas, et al.
Veröffentlicht: (2025)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
von: Vani, Sameep, et al.
Veröffentlicht: (2025)
von: Vani, Sameep, et al.
Veröffentlicht: (2025)
EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
von: Kusumba, Abhiram, et al.
Veröffentlicht: (2025)
von: Kusumba, Abhiram, et al.
Veröffentlicht: (2025)
Improving Shift Invariance in Convolutional Neural Networks with Translation Invariant Polyphase Sampling
von: Saha, Sourajit, et al.
Veröffentlicht: (2024)
von: Saha, Sourajit, et al.
Veröffentlicht: (2024)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
Plane2Depth: Hierarchical Adaptive Plane Guidance for Monocular Depth Estimation
von: Liu, Li, et al.
Veröffentlicht: (2024)
von: Liu, Li, et al.
Veröffentlicht: (2024)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
von: Siingh, Shikhhar, et al.
Veröffentlicht: (2025)
von: Siingh, Shikhhar, et al.
Veröffentlicht: (2025)
Vision-Language Embodiment for Monocular Depth Estimation
von: Zhang, Jinchang, et al.
Veröffentlicht: (2025)
von: Zhang, Jinchang, et al.
Veröffentlicht: (2025)
R.A.C.E.: Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model
von: Kim, Changhoon, et al.
Veröffentlicht: (2024)
von: Kim, Changhoon, et al.
Veröffentlicht: (2024)
Side Effects of Erasing Concepts from Diffusion Models
von: Saha, Shaswati, et al.
Veröffentlicht: (2025)
von: Saha, Shaswati, et al.
Veröffentlicht: (2025)
MMTABREAL: Real-World Benchmark for Multimodal Table Understanding
von: Titiya, Prasham, et al.
Veröffentlicht: (2025)
von: Titiya, Prasham, et al.
Veröffentlicht: (2025)
DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
von: Dong, Yue-Jiang, et al.
Veröffentlicht: (2025)
von: Dong, Yue-Jiang, et al.
Veröffentlicht: (2025)
Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted Perturbations
von: Devulapally, Naresh Kumar, et al.
Veröffentlicht: (2025)
von: Devulapally, Naresh Kumar, et al.
Veröffentlicht: (2025)
DepthDark: Robust Monocular Depth Estimation for Low-Light Environments
von: Zeng, Longjian, et al.
Veröffentlicht: (2025)
von: Zeng, Longjian, et al.
Veröffentlicht: (2025)
DepthART: Monocular Depth Estimation as Autoregressive Refinement Task
von: Gabdullin, Bulat, et al.
Veröffentlicht: (2024)
von: Gabdullin, Bulat, et al.
Veröffentlicht: (2024)
Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
von: Xu, Zhen, et al.
Veröffentlicht: (2025)
Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
DEPTHOR++: Robust Depth Enhancement from a Real-World Lightweight dToF and RGB Guidance
von: Xiang, Jijun, et al.
Veröffentlicht: (2025)
von: Xiang, Jijun, et al.
Veröffentlicht: (2025)
Learning A Low-Level Vision Generalist via Visual Task Prompt
von: Chen, Xiangyu, et al.
Veröffentlicht: (2024)
von: Chen, Xiangyu, et al.
Veröffentlicht: (2024)
Geometric-Aware Low-Light Image and Video Enhancement via Depth Guidance
von: Lin, Yingqi, et al.
Veröffentlicht: (2023)
von: Lin, Yingqi, et al.
Veröffentlicht: (2023)
Language as Prior, Vision as Calibration: Metric Scale Recovery for Monocular Depth Estimation
von: Zhan, Mingxia, et al.
Veröffentlicht: (2026)
von: Zhan, Mingxia, et al.
Veröffentlicht: (2026)
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
von: Feinglass, Joshua, et al.
Veröffentlicht: (2024)
von: Feinglass, Joshua, et al.
Veröffentlicht: (2024)
DepthLM: Metric Depth From Vision Language Models
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024) -
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
von: Malaviya, Vatsal, et al.
Veröffentlicht: (2025) -
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
von: Patel, Maitreya, et al.
Veröffentlicht: (2023) -
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025) -
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)