Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ranasinghe, Yasiru, VS, Vibashan, Uplinger, James, De Melo, Celso, Patel, Vishal M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FaceXBench: Evaluating Multimodal LLMs on Face Understanding
by: Narayan, Kartik, et al.
Published: (2025)
by: Narayan, Kartik, et al.
Published: (2025)
Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception
by: Thushara, Rusiru, et al.
Published: (2026)
by: Thushara, Rusiru, et al.
Published: (2026)
Certainty and Uncertainty Guided Active Domain Adaptation
by: Safaei, Bardia, et al.
Published: (2025)
by: Safaei, Bardia, et al.
Published: (2025)
SegFace: Face Segmentation of Long-Tail Classes
by: Narayan, Kartik, et al.
Published: (2024)
by: Narayan, Kartik, et al.
Published: (2024)
FaceXFormer: A Unified Transformer for Facial Analysis
by: Narayan, Kartik, et al.
Published: (2024)
by: Narayan, Kartik, et al.
Published: (2024)
$CrowdDiff$: Multi-hypothesis Crowd Density Estimation using Diffusion Models
by: Ranasinghe, Yasiru, et al.
Published: (2023)
by: Ranasinghe, Yasiru, et al.
Published: (2023)
Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
by: Ranasinghe, Yasiru, et al.
Published: (2026)
by: Ranasinghe, Yasiru, et al.
Published: (2026)
PosSAM: Panoptic Open-vocabulary Segment Anything
by: VS, Vibashan, et al.
Published: (2024)
by: VS, Vibashan, et al.
Published: (2024)
SINR: Sparsity Driven Compressed Implicit Neural Representations
by: Jayasundara, Dhananjaya, et al.
Published: (2025)
by: Jayasundara, Dhananjaya, et al.
Published: (2025)
Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models
by: Xu, Jiacong, et al.
Published: (2025)
by: Xu, Jiacong, et al.
Published: (2025)
Enhancing Remote Sensing Vision-Language Models for Zero-Shot Scene Classification
by: Khoury, Karim El, et al.
Published: (2024)
by: Khoury, Karim El, et al.
Published: (2024)
An Application-Agnostic Automatic Target Recognition System Using Vision Language Models
by: Palladino, Anthony, et al.
Published: (2024)
by: Palladino, Anthony, et al.
Published: (2024)
Active Learning for Vision-Language Models
by: Safaei, Bardia, et al.
Published: (2024)
by: Safaei, Bardia, et al.
Published: (2024)
F-ViTA: Foundation Model Guided Visible to Thermal Translation
by: Paranjape, Jay N., et al.
Published: (2025)
by: Paranjape, Jay N., et al.
Published: (2025)
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
by: Unmesh, Asim, et al.
Published: (2026)
by: Unmesh, Asim, et al.
Published: (2026)
A Mamba-based Siamese Network for Remote Sensing Change Detection
by: Paranjape, Jay N., et al.
Published: (2024)
by: Paranjape, Jay N., et al.
Published: (2024)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
by: Xu, Zhenlin, et al.
Published: (2023)
by: Xu, Zhenlin, et al.
Published: (2023)
Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
ContextVLM: Zero-Shot and Few-Shot Context Understanding for Autonomous Driving using Vision Language Models
by: Sural, Shounak, et al.
Published: (2024)
by: Sural, Shounak, et al.
Published: (2024)
Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition
by: Ramirez, David F., et al.
Published: (2026)
by: Ramirez, David F., et al.
Published: (2026)
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
by: Miller, Kevin, et al.
Published: (2025)
by: Miller, Kevin, et al.
Published: (2025)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
by: Sharma, Saurav, et al.
Published: (2025)
by: Sharma, Saurav, et al.
Published: (2025)
Zero-Shot Monocular Scene Flow Estimation in the Wild
by: Liang, Yiqing, et al.
Published: (2025)
by: Liang, Yiqing, et al.
Published: (2025)
Deep Learning for Cross-Domain Few-Shot Visual Recognition: A Survey
by: Xu, Huali, et al.
Published: (2023)
by: Xu, Huali, et al.
Published: (2023)
Few-Shot Class-Incremental Learning For Efficient SAR Automatic Target Recognition
by: Karantaidis, George, et al.
Published: (2025)
by: Karantaidis, George, et al.
Published: (2025)
An Evaluation of Large Pre-Trained Models for Gesture Recognition using Synthetic Videos
by: Reddy, Arun, et al.
Published: (2024)
by: Reddy, Arun, et al.
Published: (2024)
Noise is an Efficient Learner for Zero-Shot Vision-Language Models
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
Understanding Long Videos with Multimodal Language Models
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
Zero-Shot Scene Change Detection
by: Cho, Kyusik, et al.
Published: (2024)
by: Cho, Kyusik, et al.
Published: (2024)
Enhancing Zero-Shot Image Recognition in Vision-Language Models through Human-like Concept Guidance
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
VIZOR: Viewpoint-Invariant Zero-Shot Scene Graph Generation for 3D Scene Reasoning
by: Madhavaram, Vivek, et al.
Published: (2026)
by: Madhavaram, Vivek, et al.
Published: (2026)
Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
by: Mohamud, Safaa Abdullahi Moallim, et al.
Published: (2025)
by: Mohamud, Safaa Abdullahi Moallim, et al.
Published: (2025)
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
by: Li, Ruihuang, et al.
Published: (2024)
by: Li, Ruihuang, et al.
Published: (2024)
Referring Change Detection in Remote Sensing Imagery
by: Korkmaz, Yilmaz, et al.
Published: (2025)
by: Korkmaz, Yilmaz, et al.
Published: (2025)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes
by: Rivera, Antonio Carlos, et al.
Published: (2024)
by: Rivera, Antonio Carlos, et al.
Published: (2024)
Similar Items
-
FaceXBench: Evaluating Multimodal LLMs on Face Understanding
by: Narayan, Kartik, et al.
Published: (2025) -
Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception
by: Thushara, Rusiru, et al.
Published: (2026) -
Certainty and Uncertainty Guided Active Domain Adaptation
by: Safaei, Bardia, et al.
Published: (2025) -
SegFace: Face Segmentation of Long-Tail Classes
by: Narayan, Kartik, et al.
Published: (2024) -
FaceXFormer: A Unified Transformer for Facial Analysis
by: Narayan, Kartik, et al.
Published: (2024)