How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
Fuente:
arXiv
Saved in:
| Main Authors: | Brusnicki, Roberto, Piccinini, Mattia, Betz, Johannes |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Isolated Sign Language Recognition with Segmentation and Pose Estimation
by: Perkins, Daniel, et al.
Published: (2025)
by: Perkins, Daniel, et al.
Published: (2025)
Interactive Image Selection and Training for Brain Tumor Segmentation Network
by: Cerqueira, Matheus A., et al.
Published: (2024)
by: Cerqueira, Matheus A., et al.
Published: (2024)
Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data
by: Richings, Jack, et al.
Published: (2025)
by: Richings, Jack, et al.
Published: (2025)
TWIG: Two-Step Image Generation using Segmentation Masks in Diffusion Models
by: Rakib, Mazharul Islam, et al.
Published: (2025)
by: Rakib, Mazharul Islam, et al.
Published: (2025)
RailSafeNet: Visual Scene Understanding for Tram Safety
by: Valach, Ondřej, et al.
Published: (2025)
by: Valach, Ondřej, et al.
Published: (2025)
Depth Priors in Removal Neural Radiance Fields
by: Guo, Zhihao, et al.
Published: (2024)
by: Guo, Zhihao, et al.
Published: (2024)
Z-Order Transformer for Feed-Forward Gaussian Splatting
by: Wang, Can, et al.
Published: (2026)
by: Wang, Can, et al.
Published: (2026)
EatGAN: An Edge-Attention Guided Generative Adversarial Network for Single Image Super-Resolution
by: Rao, Penghao, et al.
Published: (2025)
by: Rao, Penghao, et al.
Published: (2025)
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
by: Zhang, Shengkai, et al.
Published: (2024)
by: Zhang, Shengkai, et al.
Published: (2024)
Transformer-Based Model for Monocular Visual Odometry: A Video Understanding Approach
by: Françani, André O., et al.
Published: (2023)
by: Françani, André O., et al.
Published: (2023)
Motion Consistency Loss for Monocular Visual Odometry with Attention-Based Deep Learning
by: Françani, André O., et al.
Published: (2024)
by: Françani, André O., et al.
Published: (2024)
Building Brain Tumor Segmentation Networks with User-Assisted Filter Estimation and Selection
by: Cerqueira, Matheus A., et al.
Published: (2024)
by: Cerqueira, Matheus A., et al.
Published: (2024)
Does CLIP perceive art the same way we do?
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
Hybrid SIFT-SNN for Efficient Anomaly Detection of Traffic Flow-Control Infrastructure
by: Rathee, Munish, et al.
Published: (2025)
by: Rathee, Munish, et al.
Published: (2025)
SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval
by: Wu, Ren-Di, et al.
Published: (2025)
by: Wu, Ren-Di, et al.
Published: (2025)
Unlocking UML Class Diagram Understanding in Vision Language Models
by: Naboichenko, Artem, et al.
Published: (2026)
by: Naboichenko, Artem, et al.
Published: (2026)
Vision-Language Models for Acute Tuberculosis Diagnosis: A Multimodal Approach Combining Imaging and Clinical Data
by: Ganapthy, Ananya, et al.
Published: (2025)
by: Ganapthy, Ananya, et al.
Published: (2025)
Computational Imaging Priors for Wireless Capsule Endoscopy: Monte Carlo-Guided Hemoglobin Mapping for Rare-Anomaly Detection
by: Yang, Chengshuai, et al.
Published: (2026)
by: Yang, Chengshuai, et al.
Published: (2026)
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
by: Jin, Hang, et al.
Published: (2025)
by: Jin, Hang, et al.
Published: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
by: Ma, Yueen, et al.
Published: (2024)
by: Ma, Yueen, et al.
Published: (2024)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
by: Salehi, Pegah, et al.
Published: (2025)
by: Salehi, Pegah, et al.
Published: (2025)
Memory-augmented Online Video Anomaly Detection
by: Rossi, Leonardo, et al.
Published: (2023)
by: Rossi, Leonardo, et al.
Published: (2023)
CNNtention: Can CNNs do better with Attention?
by: Kapila, Nikhil, et al.
Published: (2024)
by: Kapila, Nikhil, et al.
Published: (2024)
Poisson Flow Consistency Training
by: Zhang, Anthony, et al.
Published: (2025)
by: Zhang, Anthony, et al.
Published: (2025)
Explainable Image Similarity: Integrating Siamese Networks and Grad-CAM
by: Livieris, Ioannis E., et al.
Published: (2023)
by: Livieris, Ioannis E., et al.
Published: (2023)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
by: Hassan, Umair
Published: (2025)
by: Hassan, Umair
Published: (2025)
JVLGS: Joint Vision-Language Gas Leak Segmentation
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
by: Cui, Hejie, et al.
Published: (2024)
by: Cui, Hejie, et al.
Published: (2024)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
by: Ouyang, Rongxin, et al.
Published: (2024)
by: Ouyang, Rongxin, et al.
Published: (2024)
MambaNetLK: Enhancing Colonoscopy Point Cloud Registration with Mamba
by: Jiang, Linzhe, et al.
Published: (2025)
by: Jiang, Linzhe, et al.
Published: (2025)
LRVS-Fashion: Extending Visual Search with Referring Instructions
by: Lepage, Simon, et al.
Published: (2023)
by: Lepage, Simon, et al.
Published: (2023)
E Pluribus Unum Interpretable Convolutional Neural Networks
by: Dimas, George, et al.
Published: (2022)
by: Dimas, George, et al.
Published: (2022)
Harmony: A Joint Self-Supervised and Weakly-Supervised Framework for Learning General Purpose Visual Representations
by: Baharoon, Mohammed, et al.
Published: (2024)
by: Baharoon, Mohammed, et al.
Published: (2024)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
Human Vision Constrained Super-Resolution
by: Karpenko, Volodymyr, et al.
Published: (2024)
by: Karpenko, Volodymyr, et al.
Published: (2024)
If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions
by: Barbano, Carlo Alberto, et al.
Published: (2024)
by: Barbano, Carlo Alberto, et al.
Published: (2024)
Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
by: Shin, Suyeon, et al.
Published: (2024)
by: Shin, Suyeon, et al.
Published: (2024)
Similar Items
-
Isolated Sign Language Recognition with Segmentation and Pose Estimation
by: Perkins, Daniel, et al.
Published: (2025) -
Interactive Image Selection and Training for Brain Tumor Segmentation Network
by: Cerqueira, Matheus A., et al.
Published: (2024) -
Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data
by: Richings, Jack, et al.
Published: (2025) -
TWIG: Two-Step Image Generation using Segmentation Masks in Diffusion Models
by: Rakib, Mazharul Islam, et al.
Published: (2025) -
RailSafeNet: Visual Scene Understanding for Tram Safety
by: Valach, Ondřej, et al.
Published: (2025)