Gespeichert in:
| Hauptverfasser: | Choi, Lucas, Greer, Ross |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2408.02244 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Cascaded Methods of Vision-Language Models for Zero-Shot Detection and Association of Hardhats for Increased Construction Safety
von: Choi, Lucas, et al.
Veröffentlicht: (2024)
von: Choi, Lucas, et al.
Veröffentlicht: (2024)
Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety
von: Shriram, Shashank, et al.
Veröffentlicht: (2025)
von: Shriram, Shashank, et al.
Veröffentlicht: (2025)
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
Driver Activity Classification Using Generalizable Representations from Vision-Language Models
von: Greer, Ross, et al.
Veröffentlicht: (2024)
von: Greer, Ross, et al.
Veröffentlicht: (2024)
Beyond General Prompts: Automated Prompt Refinement using Contrastive Class Alignment Scores for Disambiguating Objects in Vision-Language Models
von: Choi, Lucas, et al.
Veröffentlicht: (2025)
von: Choi, Lucas, et al.
Veröffentlicht: (2025)
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
von: Bendou, Yassir, et al.
Veröffentlicht: (2024)
von: Bendou, Yassir, et al.
Veröffentlicht: (2024)
DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving
von: Kirchner, Sven, et al.
Veröffentlicht: (2025)
von: Kirchner, Sven, et al.
Veröffentlicht: (2025)
Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models
von: Martinez-Sanchez, Angel, et al.
Veröffentlicht: (2026)
von: Martinez-Sanchez, Angel, et al.
Veröffentlicht: (2026)
doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
von: Roy, Parthib, et al.
Veröffentlicht: (2024)
von: Roy, Parthib, et al.
Veröffentlicht: (2024)
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
von: Greer, Ross, et al.
Veröffentlicht: (2024)
von: Greer, Ross, et al.
Veröffentlicht: (2024)
Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles
von: Bossen, Tonko E. W., et al.
Veröffentlicht: (2025)
von: Bossen, Tonko E. W., et al.
Veröffentlicht: (2025)
Towards Explainable, Safe Autonomous Driving with Language Embeddings for Novelty Identification and Active Learning: Framework and Experimental Analysis with Real-World Data Sets
von: Greer, Ross, et al.
Veröffentlicht: (2024)
von: Greer, Ross, et al.
Veröffentlicht: (2024)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
von: Zhang, Pu, et al.
Veröffentlicht: (2025)
von: Zhang, Pu, et al.
Veröffentlicht: (2025)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
von: Xu, Zhenlin, et al.
Veröffentlicht: (2023)
von: Xu, Zhenlin, et al.
Veröffentlicht: (2023)
Benchmarking Foundation Models for Zero-Shot Biometric Tasks
von: Sony, Redwan, et al.
Veröffentlicht: (2025)
von: Sony, Redwan, et al.
Veröffentlicht: (2025)
Language-Driven Active Learning for Diverse Open-Set 3D Object Detection
von: Greer, Ross, et al.
Veröffentlicht: (2024)
von: Greer, Ross, et al.
Veröffentlicht: (2024)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
Bayesian Modeling of Zero-Shot Classifications for Urban Flood Detection
von: Franchi, Matt, et al.
Veröffentlicht: (2025)
von: Franchi, Matt, et al.
Veröffentlicht: (2025)
Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection
von: Aqeel, Muhammad, et al.
Veröffentlicht: (2026)
von: Aqeel, Muhammad, et al.
Veröffentlicht: (2026)
Binary Verification for Zero-Shot Vision
von: Hu, Rongbin, et al.
Veröffentlicht: (2025)
von: Hu, Rongbin, et al.
Veröffentlicht: (2025)
TinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings
von: Wilson, Bibin
Veröffentlicht: (2026)
von: Wilson, Bibin
Veröffentlicht: (2026)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
von: Jeong, Seongjun, et al.
Veröffentlicht: (2024)
von: Jeong, Seongjun, et al.
Veröffentlicht: (2024)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
von: Yu, Lu, et al.
Veröffentlicht: (2024)
von: Yu, Lu, et al.
Veröffentlicht: (2024)
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation
von: Sajib, Rakib Hossain, et al.
Veröffentlicht: (2026)
von: Sajib, Rakib Hossain, et al.
Veröffentlicht: (2026)
Enhancing Construction Site Safety: A Lightweight Convolutional Network for Effective Helmet Detection
von: Alif, Mujadded Al Rabbani
Veröffentlicht: (2024)
von: Alif, Mujadded Al Rabbani
Veröffentlicht: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
von: Luo, Kun, et al.
Veröffentlicht: (2026)
von: Luo, Kun, et al.
Veröffentlicht: (2026)
TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation
von: Li, Dingbang, et al.
Veröffentlicht: (2024)
von: Li, Dingbang, et al.
Veröffentlicht: (2024)
HeatPrompt: Zero-Shot Vision-Language Modeling of Urban Heat Demand from Satellite Images
von: Thota, Kundan, et al.
Veröffentlicht: (2026)
von: Thota, Kundan, et al.
Veröffentlicht: (2026)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
Finding the Reflection Point: Unpadding Images to Remove Data Augmentation Artifacts in Large Open Source Image Datasets for Machine Learning
von: Choi, Lucas, et al.
Veröffentlicht: (2025)
von: Choi, Lucas, et al.
Veröffentlicht: (2025)
Optimizing Helmet Detection with Hybrid YOLO Pipelines: A Detailed Analysis
von: M, Vaikunth, et al.
Veröffentlicht: (2024)
von: M, Vaikunth, et al.
Veröffentlicht: (2024)
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
von: Metzen, Jan Hendrik, et al.
Veröffentlicht: (2023)
von: Metzen, Jan Hendrik, et al.
Veröffentlicht: (2023)
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
von: Wang, Haibo, et al.
Veröffentlicht: (2026)
von: Wang, Haibo, et al.
Veröffentlicht: (2026)
Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models
von: Xu, Zane, et al.
Veröffentlicht: (2025)
von: Xu, Zane, et al.
Veröffentlicht: (2025)
MTR-VP: Towards End-to-End Trajectory Planning through Context-Driven Image Encoding and Multiple Trajectory Prediction
von: Keskar, Maitrayee, et al.
Veröffentlicht: (2025)
von: Keskar, Maitrayee, et al.
Veröffentlicht: (2025)
VideoPoet: A Large Language Model for Zero-Shot Video Generation
von: Kondratyuk, Dan, et al.
Veröffentlicht: (2023)
von: Kondratyuk, Dan, et al.
Veröffentlicht: (2023)
Language as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity
von: Tang, MingZe, et al.
Veröffentlicht: (2025)
von: Tang, MingZe, et al.
Veröffentlicht: (2025)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
Towards Efficient and General-Purpose Few-Shot Misclassification Detection for Vision-Language Models
von: Zeng, Fanhu, et al.
Veröffentlicht: (2025)
von: Zeng, Fanhu, et al.
Veröffentlicht: (2025)
Zero-Training Task-Specific Model Synthesis for Few-Shot Medical Image Classification
von: Qin, Yao, et al.
Veröffentlicht: (2025)
von: Qin, Yao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating Cascaded Methods of Vision-Language Models for Zero-Shot Detection and Association of Hardhats for Increased Construction Safety
von: Choi, Lucas, et al.
Veröffentlicht: (2024) -
Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety
von: Shriram, Shashank, et al.
Veröffentlicht: (2025) -
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024) -
Driver Activity Classification Using Generalizable Representations from Vision-Language Models
von: Greer, Ross, et al.
Veröffentlicht: (2024) -
Beyond General Prompts: Automated Prompt Refinement using Contrastive Class Alignment Scores for Disambiguating Objects in Vision-Language Models
von: Choi, Lucas, et al.
Veröffentlicht: (2025)