Driver Activity Classification Using Generalizable Representations from Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Greer, Ross, Andersen, Mathias Viborg, Møgelmose, Andreas, Trivedi, Mohan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Find Missing Video Frames with Synthetic Data Augmentation: A General Framework and Application in Generating Thermal Images Using RGB Cameras
by: Andersen, Mathias Viborg, et al.
Published: (2024)
by: Andersen, Mathias Viborg, et al.
Published: (2024)
Language-Driven Active Learning for Diverse Open-Set 3D Object Detection
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
Towards Explainable, Safe Autonomous Driving with Language Embeddings for Novelty Identification and Active Learning: Framework and Experimental Analysis with Real-World Data Sets
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
The Why, When, and How to Use Active Learning in Large-Data-Driven 3D Object Detection for Safe Autonomous Driving: An Empirical Exploration
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
Vision and Language: Novel Representations and Artificial intelligence for Driving Scene Safety Assessment and Autonomous Vehicle Planning
by: Greer, Ross, et al.
Published: (2026)
by: Greer, Ross, et al.
Published: (2026)
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
by: Gopalkrishnan, Akshay, et al.
Published: (2024)
by: Gopalkrishnan, Akshay, et al.
Published: (2024)
Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles
by: Bossen, Tonko E. W., et al.
Published: (2025)
by: Bossen, Tonko E. W., et al.
Published: (2025)
Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models
by: Martinez-Sanchez, Angel, et al.
Published: (2026)
by: Martinez-Sanchez, Angel, et al.
Published: (2026)
MTR-VP: Towards End-to-End Trajectory Planning through Context-Driven Image Encoding and Multiple Trajectory Prediction
by: Keskar, Maitrayee, et al.
Published: (2025)
by: Keskar, Maitrayee, et al.
Published: (2025)
Evaluating Vision-Language Models for Zero-Shot Detection, Classification, and Association of Motorcycles, Passengers, and Helmets
by: Choi, Lucas, et al.
Published: (2024)
by: Choi, Lucas, et al.
Published: (2024)
Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making
by: Greer, Ross, et al.
Published: (2026)
by: Greer, Ross, et al.
Published: (2026)
ActiveAnno3D -- An Active Learning Framework for Multi-Modal 3D Object Detection
by: Ghita, Ahmed, et al.
Published: (2024)
by: Ghita, Ahmed, et al.
Published: (2024)
Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety
by: Shriram, Shashank, et al.
Published: (2025)
by: Shriram, Shashank, et al.
Published: (2025)
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
by: Chen, William, et al.
Published: (2024)
by: Chen, William, et al.
Published: (2024)
ALERT Open Dataset and Input-Size-Agnostic Vision Transformer for Driver Activity Recognition using IR-UWB
by: Park, Jeongjun, et al.
Published: (2025)
by: Park, Jeongjun, et al.
Published: (2025)
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
by: Abdollahi, Ali, et al.
Published: (2024)
by: Abdollahi, Ali, et al.
Published: (2024)
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
by: Roschmann, Simon, et al.
Published: (2025)
by: Roschmann, Simon, et al.
Published: (2025)
Wavelet-Driven Generalizable Framework for Deepfake Face Forgery Detection
by: Baru, Lalith Bharadwaj, et al.
Published: (2024)
by: Baru, Lalith Bharadwaj, et al.
Published: (2024)
Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents
by: Zhang, Zhizhen, et al.
Published: (2025)
by: Zhang, Zhizhen, et al.
Published: (2025)
BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning
by: Wu, Yizhou, et al.
Published: (2026)
by: Wu, Yizhou, et al.
Published: (2026)
Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
by: Yang, Zhijian, et al.
Published: (2025)
by: Yang, Zhijian, et al.
Published: (2025)
Transformer-Based Contrastive Meta-Learning For Low-Resource Generalizable Activity Recognition
by: Wang, Junyao, et al.
Published: (2024)
by: Wang, Junyao, et al.
Published: (2024)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
by: Grover, Shresth, et al.
Published: (2025)
by: Grover, Shresth, et al.
Published: (2025)
Hyperspectral Adapter for Semantic Segmentation with Vision Foundation Models
by: Hurtado, Juana Valeria, et al.
Published: (2025)
by: Hurtado, Juana Valeria, et al.
Published: (2025)
ForecastOcc: Vision-based Semantic Occupancy Forecasting
by: Mohan, Riya, et al.
Published: (2026)
by: Mohan, Riya, et al.
Published: (2026)
AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
by: Vasa, Santosh, et al.
Published: (2025)
by: Vasa, Santosh, et al.
Published: (2025)
FastVLM: Efficient Vision Encoding for Vision Language Models
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
On the Generalizability of ECG-based Stress Detection Models
by: Prajod, Pooja, et al.
Published: (2022)
by: Prajod, Pooja, et al.
Published: (2022)
Uncertainties of Latent Representations in Computer Vision
by: Kirchhof, Michael
Published: (2024)
by: Kirchhof, Michael
Published: (2024)
Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
by: Abid, Abderrazek, et al.
Published: (2025)
by: Abid, Abderrazek, et al.
Published: (2025)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
by: Chen, Yi-Chia, et al.
Published: (2024)
by: Chen, Yi-Chia, et al.
Published: (2024)
Adapting Vision-Language Models for Evaluating World Models
by: Hendriksen, Mariya, et al.
Published: (2025)
by: Hendriksen, Mariya, et al.
Published: (2025)
DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving
by: Kirchner, Sven, et al.
Published: (2025)
by: Kirchner, Sven, et al.
Published: (2025)
Multi-Modal Adapter for Vision-Language Models
by: Seputis, Dominykas, et al.
Published: (2024)
by: Seputis, Dominykas, et al.
Published: (2024)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
by: Lin, Zinan, et al.
Published: (2025)
by: Lin, Zinan, et al.
Published: (2025)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
by: Lian, Chenyu, et al.
Published: (2025)
by: Lian, Chenyu, et al.
Published: (2025)
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
by: Nguyen, Phu-Vinh, et al.
Published: (2025)
by: Nguyen, Phu-Vinh, et al.
Published: (2025)
Similar Items
-
Learning to Find Missing Video Frames with Synthetic Data Augmentation: A General Framework and Application in Generating Thermal Images Using RGB Cameras
by: Andersen, Mathias Viborg, et al.
Published: (2024) -
Language-Driven Active Learning for Diverse Open-Set 3D Object Detection
by: Greer, Ross, et al.
Published: (2024) -
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
by: Greer, Ross, et al.
Published: (2024) -
Towards Explainable, Safe Autonomous Driving with Language Embeddings for Novelty Identification and Active Learning: Framework and Experimental Analysis with Real-World Data Sets
by: Greer, Ross, et al.
Published: (2024) -
The Why, When, and How to Use Active Learning in Large-Data-Driven 3D Object Detection for Safe Autonomous Driving: An Empirical Exploration
by: Greer, Ross, et al.
Published: (2024)