Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix
Fuente:
arXiv
Guardado en:
| Autores principales: | Gurbindo, Unai, Brando, Axel, Abella, Jaume, König, Caroline |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
por: Wu, Jason, et al.
Publicado: (2026)
por: Wu, Jason, et al.
Publicado: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
por: Lin, Xiaojian, et al.
Publicado: (2025)
por: Lin, Xiaojian, et al.
Publicado: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
Implementing Adaptations for Vision AutoRegressive Model
por: Shaikh, Kaif, et al.
Publicado: (2025)
por: Shaikh, Kaif, et al.
Publicado: (2025)
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
por: Zhong, Zesen, et al.
Publicado: (2025)
por: Zhong, Zesen, et al.
Publicado: (2025)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
por: Brovko, D. V.
Publicado: (2025)
por: Brovko, D. V.
Publicado: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
por: Hou, Zhangcheng, et al.
Publicado: (2026)
por: Hou, Zhangcheng, et al.
Publicado: (2026)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
por: Pu, Qingwen, et al.
Publicado: (2026)
por: Pu, Qingwen, et al.
Publicado: (2026)
Smooth regularization for efficient video recognition
por: Goldman, Gil, et al.
Publicado: (2025)
por: Goldman, Gil, et al.
Publicado: (2025)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
por: Bergkvist, Viktor, et al.
Publicado: (2026)
por: Bergkvist, Viktor, et al.
Publicado: (2026)
Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
por: Jiang, Xintong, et al.
Publicado: (2025)
por: Jiang, Xintong, et al.
Publicado: (2025)
Salient Concept-Aware Generative Data Augmentation
por: Zhao, Tianchen, et al.
Publicado: (2025)
por: Zhao, Tianchen, et al.
Publicado: (2025)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
por: Huang, Yian, et al.
Publicado: (2026)
por: Huang, Yian, et al.
Publicado: (2026)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
por: Boudras, Thomas, et al.
Publicado: (2025)
por: Boudras, Thomas, et al.
Publicado: (2025)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
NV3D: Leveraging Spatial Shape Through Normal Vector-based 3D Object Detection
por: Chaowakarn, Krittin, et al.
Publicado: (2025)
por: Chaowakarn, Krittin, et al.
Publicado: (2025)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
por: Ji, Binbin, et al.
Publicado: (2025)
por: Ji, Binbin, et al.
Publicado: (2025)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
por: Chahine, Makram, et al.
Publicado: (2024)
por: Chahine, Makram, et al.
Publicado: (2024)
Caption-Driven Explainability: Probing CNNs for Bias via CLIP
por: Koller, Patrick, et al.
Publicado: (2025)
por: Koller, Patrick, et al.
Publicado: (2025)
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
por: Koroglu, Mathis, et al.
Publicado: (2024)
por: Koroglu, Mathis, et al.
Publicado: (2024)
Perceptual Flow Network for Visually Grounded Reasoning
por: Li, Yangfu, et al.
Publicado: (2026)
por: Li, Yangfu, et al.
Publicado: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
por: Liu, Bingnan, et al.
Publicado: (2026)
por: Liu, Bingnan, et al.
Publicado: (2026)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
por: Arnaud, Sergio, et al.
Publicado: (2025)
por: Arnaud, Sergio, et al.
Publicado: (2025)
Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models
por: Martin, Michael R., et al.
Publicado: (2025)
por: Martin, Michael R., et al.
Publicado: (2025)
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
por: Li, Zhuowei, et al.
Publicado: (2025)
por: Li, Zhuowei, et al.
Publicado: (2025)
Neural Attention: A Novel Mechanism for Enhanced Expressive Power in Transformer Models
por: DiGiugno, Andrew, et al.
Publicado: (2025)
por: DiGiugno, Andrew, et al.
Publicado: (2025)
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
por: Diller, Christian, et al.
Publicado: (2023)
por: Diller, Christian, et al.
Publicado: (2023)
Rethinking Visual Intelligence: Insights from Video Pretraining
por: Acuaviva, Pablo, et al.
Publicado: (2025)
por: Acuaviva, Pablo, et al.
Publicado: (2025)
Motion Attribution for Video Generation
por: Wu, Xindi, et al.
Publicado: (2026)
por: Wu, Xindi, et al.
Publicado: (2026)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
por: Gupta, Sunny, et al.
Publicado: (2024)
por: Gupta, Sunny, et al.
Publicado: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
por: Kashyap, Pankhi, et al.
Publicado: (2024)
por: Kashyap, Pankhi, et al.
Publicado: (2024)
TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition
por: Hassan, Imtiaz Ul, et al.
Publicado: (2026)
por: Hassan, Imtiaz Ul, et al.
Publicado: (2026)
DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification
por: Ho, Darryl, et al.
Publicado: (2025)
por: Ho, Darryl, et al.
Publicado: (2025)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
por: Peng, Hongxing, et al.
Publicado: (2025)
por: Peng, Hongxing, et al.
Publicado: (2025)
PhysicsNeRF: Physics-Guided 3D Reconstruction from Sparse Views
por: Barhdadi, Mohamed Rayan, et al.
Publicado: (2025)
por: Barhdadi, Mohamed Rayan, et al.
Publicado: (2025)
Efficient Attention: Attention with Linear Complexities
por: Shen, Zhuoran, et al.
Publicado: (2018)
por: Shen, Zhuoran, et al.
Publicado: (2018)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
por: Rajapaksha, Uchitha, et al.
Publicado: (2024)
por: Rajapaksha, Uchitha, et al.
Publicado: (2024)
Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
por: Lapin, Mykyta, et al.
Publicado: (2025)
por: Lapin, Mykyta, et al.
Publicado: (2025)
Ejemplares similares
-
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
por: Wu, Jason, et al.
Publicado: (2026) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025) -
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
por: Lin, Xiaojian, et al.
Publicado: (2025) -
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026) -
Implementing Adaptations for Vision AutoRegressive Model
por: Shaikh, Kaif, et al.
Publicado: (2025)