Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Rajiv, Manjunath Prasad Holenarasipura, Vidyavathi, B. M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models
by: Ranasinghe, Yasiru, et al.
Published: (2025)
by: Ranasinghe, Yasiru, et al.
Published: (2025)
Enhancing Remote Sensing Vision-Language Models for Zero-Shot Scene Classification
by: Khoury, Karim El, et al.
Published: (2024)
by: Khoury, Karim El, et al.
Published: (2024)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
by: Chen, Kehan, et al.
Published: (2024)
by: Chen, Kehan, et al.
Published: (2024)
Towards Zero-Shot Annotation of the Built Environment with Vision-Language Models (Vision Paper)
by: Han, Bin, et al.
Published: (2024)
by: Han, Bin, et al.
Published: (2024)
Zero-Shot Head Swapping in Real-World Scenarios
by: Kang, Taewoong, et al.
Published: (2025)
by: Kang, Taewoong, et al.
Published: (2025)
ContextVLM: Zero-Shot and Few-Shot Context Understanding for Autonomous Driving using Vision Language Models
by: Sural, Shounak, et al.
Published: (2024)
by: Sural, Shounak, et al.
Published: (2024)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
Zero-Shot Scene Change Detection
by: Cho, Kyusik, et al.
Published: (2024)
by: Cho, Kyusik, et al.
Published: (2024)
From the Laboratory to Real-World Application: Evaluating Zero-Shot Scene Interpretation on Edge Devices for Mobile Robotics
by: Schuler, Nicolas, et al.
Published: (2025)
by: Schuler, Nicolas, et al.
Published: (2025)
Zero-Shot Multi-Object Scene Completion
by: Iwase, Shun, et al.
Published: (2024)
by: Iwase, Shun, et al.
Published: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
by: Luo, Kun, et al.
Published: (2026)
by: Luo, Kun, et al.
Published: (2026)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
Noise is an Efficient Learner for Zero-Shot Vision-Language Models
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
by: Qiao, Yanyuan, et al.
Published: (2024)
by: Qiao, Yanyuan, et al.
Published: (2024)
Zero-Shot Monocular Scene Flow Estimation in the Wild
by: Liang, Yiqing, et al.
Published: (2025)
by: Liang, Yiqing, et al.
Published: (2025)
ZeroSCD: Zero-Shot Street Scene Change Detection
by: Kannan, Shyam Sundar, et al.
Published: (2024)
by: Kannan, Shyam Sundar, et al.
Published: (2024)
Leveraging Vision-Language Embeddings for Zero-Shot Learning in Histopathology Images
by: Rahaman, Md Mamunur, et al.
Published: (2025)
by: Rahaman, Md Mamunur, et al.
Published: (2025)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions
by: Cao, Yongpeng, et al.
Published: (2026)
by: Cao, Yongpeng, et al.
Published: (2026)
SPADE: Sparsity Adaptive Depth Estimator for Zero-Shot, Real-Time, Monocular Depth Estimation in Underwater Environments
by: Zhang, Hongjie, et al.
Published: (2025)
by: Zhang, Hongjie, et al.
Published: (2025)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
Zero-Shot CFC: Fast Real-World Image Denoising based on Cross-Frequency Consistency
by: Jiang, Yanlin, et al.
Published: (2025)
by: Jiang, Yanlin, et al.
Published: (2025)
DDS: Decoupled Dynamic Scene-Graph Generation Network
by: Iftekhar, A S M, et al.
Published: (2023)
by: Iftekhar, A S M, et al.
Published: (2023)
RLM: A Vision-Language Model Approach for Radar Scene Understanding
by: Mishra, Pushkal, et al.
Published: (2025)
by: Mishra, Pushkal, et al.
Published: (2025)
A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
by: Zaouali, Mahmoud Chick, et al.
Published: (2025)
by: Zaouali, Mahmoud Chick, et al.
Published: (2025)
Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
by: Yin, Xiaojie, et al.
Published: (2025)
by: Yin, Xiaojie, et al.
Published: (2025)
Zero-Shot Robustness of Vision Language Models Via Confidence-Aware Weighting
by: Naghavian, Nikoo, et al.
Published: (2025)
by: Naghavian, Nikoo, et al.
Published: (2025)
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
by: Han, Chaolei, et al.
Published: (2025)
by: Han, Chaolei, et al.
Published: (2025)
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
by: Unmesh, Asim, et al.
Published: (2026)
by: Unmesh, Asim, et al.
Published: (2026)
Exploring the Zero-Shot Capabilities of Vision-Language Models for Improving Gaze Following
by: Gupta, Anshul, et al.
Published: (2024)
by: Gupta, Anshul, et al.
Published: (2024)
Parallel Neural Computing for Scene Understanding from LiDAR Perception in Autonomous Racing
by: Sah, Suwesh Prasad
Published: (2024)
by: Sah, Suwesh Prasad
Published: (2024)
VIZOR: Viewpoint-Invariant Zero-Shot Scene Graph Generation for 3D Scene Reasoning
by: Madhavaram, Vivek, et al.
Published: (2026)
by: Madhavaram, Vivek, et al.
Published: (2026)
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
by: Liu, Hongbo, et al.
Published: (2025)
by: Liu, Hongbo, et al.
Published: (2025)
OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
by: Liu, Katherine, et al.
Published: (2025)
by: Liu, Katherine, et al.
Published: (2025)
Open World Scene Graph Generation using Vision Language Models
by: Dutta, Amartya, et al.
Published: (2025)
by: Dutta, Amartya, et al.
Published: (2025)
Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click
by: Ruschel, Raphael, et al.
Published: (2025)
by: Ruschel, Raphael, et al.
Published: (2025)
Similar Items
-
Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025) -
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models
by: Ranasinghe, Yasiru, et al.
Published: (2025) -
Enhancing Remote Sensing Vision-Language Models for Zero-Shot Scene Classification
by: Khoury, Karim El, et al.
Published: (2024) -
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
by: Elhenawy, Mohammed, et al.
Published: (2025) -
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
by: Cao, Yue, et al.
Published: (2024)