Saved in:
| Main Authors: | Crulis, Ben, De Runz, Cyril, Serres, Barthelemy, Venturini, Gilles |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.06298 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Few-shot target-driven instance detection based on open-vocabulary object detection models
by: Crulis, Ben, et al.
Published: (2024)
by: Crulis, Ben, et al.
Published: (2024)
An experimental comparative study of backpropagation and alternatives for training binary neural networks for image classification
by: Crulis, Ben, et al.
Published: (2024)
by: Crulis, Ben, et al.
Published: (2024)
Spectral Norm of Convolutional Layers with Circular and Zero Paddings
by: Delattre, Blaise, et al.
Published: (2024)
by: Delattre, Blaise, et al.
Published: (2024)
Robustness to distribution shifts of compressed networks for edge devices
by: Shen, Lulan, et al.
Published: (2024)
by: Shen, Lulan, et al.
Published: (2024)
Block Selective Reprogramming for On-device Training of Vision Transformers
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
Lotus: learning-based online thermal and latency variation management for two-stage detectors on edge devices
by: Gong, Yifan, et al.
Published: (2024)
by: Gong, Yifan, et al.
Published: (2024)
Learning without Forgetting for Vision-Language Models
by: Zhou, Da-Wei, et al.
Published: (2023)
by: Zhou, Da-Wei, et al.
Published: (2023)
Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling
by: Hazimeh, Adam, et al.
Published: (2025)
by: Hazimeh, Adam, et al.
Published: (2025)
Bridge the Modality and Capability Gaps in Vision-Language Model Selection
by: Yi, Chao, et al.
Published: (2024)
by: Yi, Chao, et al.
Published: (2024)
Vision Language Models are Biased
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
Demystifying Variational Diffusion Models
by: Ribeiro, Fabio De Sousa, et al.
Published: (2024)
by: Ribeiro, Fabio De Sousa, et al.
Published: (2024)
Understanding Task Transfer in Vision-Language Models
by: Sachdeva, Bhuvan, et al.
Published: (2025)
by: Sachdeva, Bhuvan, et al.
Published: (2025)
Neutral-Reference Prompting for Vision-Language Models
by: Tian, Senmao, et al.
Published: (2026)
by: Tian, Senmao, et al.
Published: (2026)
Post-hoc Probabilistic Vision-Language Models
by: Baumann, Anton, et al.
Published: (2024)
by: Baumann, Anton, et al.
Published: (2024)
Déjà Vu Memorization in Vision-Language Models
by: Jayaraman, Bargav, et al.
Published: (2024)
by: Jayaraman, Bargav, et al.
Published: (2024)
GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
by: Li, Ling, et al.
Published: (2024)
by: Li, Ling, et al.
Published: (2024)
Self-Evolving Visual Concept Library using Vision-Language Critics
by: Sehgal, Atharva, et al.
Published: (2025)
by: Sehgal, Atharva, et al.
Published: (2025)
Medical Vision Language Models as Policies for Robotic Surgery
by: Muppidi, Akshay, et al.
Published: (2025)
by: Muppidi, Akshay, et al.
Published: (2025)
Attention Guided Alignment in Efficient Vision-Language Models
by: Mahajan, Shweta, et al.
Published: (2025)
by: Mahajan, Shweta, et al.
Published: (2025)
Attribute-based Visual Reprogramming for Vision-Language Models
by: Cai, Chengyi, et al.
Published: (2025)
by: Cai, Chengyi, et al.
Published: (2025)
Quantifying Cross-Modality Memorization in Vision-Language Models
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
Improved Alignment of Modalities in Large Vision Language Models
by: Jangra, Kartik, et al.
Published: (2025)
by: Jangra, Kartik, et al.
Published: (2025)
Decoupling the components of geometric understanding in Vision Language Models
by: Kosoy, Eliza, et al.
Published: (2025)
by: Kosoy, Eliza, et al.
Published: (2025)
The Dual Mechanisms of Spatial Reasoning in Vision-Language Models
by: Cui, Kelly, et al.
Published: (2026)
by: Cui, Kelly, et al.
Published: (2026)
Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models
by: Kubaty, Piotr, et al.
Published: (2026)
by: Kubaty, Piotr, et al.
Published: (2026)
Interpreting Neurons in Deep Vision Networks with Language Models
by: Bai, Nicholas, et al.
Published: (2024)
by: Bai, Nicholas, et al.
Published: (2024)
Vision-Language Models are Strong Noisy Label Detectors
by: Wei, Tong, et al.
Published: (2024)
by: Wei, Tong, et al.
Published: (2024)
Leveraging Vision Language Models for Specialized Agricultural Tasks
by: Arshad, Muhammad Arbab, et al.
Published: (2024)
by: Arshad, Muhammad Arbab, et al.
Published: (2024)
Reinforcement Learning Friendly Vision-Language Model for Minecraft
by: Jiang, Haobin, et al.
Published: (2023)
by: Jiang, Haobin, et al.
Published: (2023)
Detecting and Preventing Hallucinations in Large Vision Language Models
by: Gunjal, Anisha, et al.
Published: (2023)
by: Gunjal, Anisha, et al.
Published: (2023)
MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
by: Zhou, Xunlan, et al.
Published: (2026)
by: Zhou, Xunlan, et al.
Published: (2026)
Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
by: Xu, Haochuan, et al.
Published: (2025)
by: Xu, Haochuan, et al.
Published: (2025)
On Vision Transformers for Classification Tasks in Side-Scan Sonar Imagery
by: Sheffield, BW, et al.
Published: (2024)
by: Sheffield, BW, et al.
Published: (2024)
The Neglected Tails in Vision-Language Models
by: Parashar, Shubham, et al.
Published: (2024)
by: Parashar, Shubham, et al.
Published: (2024)
MMRL: Multi-Modal Representation Learning for Vision-Language Models
by: Guo, Yuncheng, et al.
Published: (2025)
by: Guo, Yuncheng, et al.
Published: (2025)
Universal Camouflage Attack on Vision-Language Models for Autonomous Driving
by: Kong, Dehong, et al.
Published: (2025)
by: Kong, Dehong, et al.
Published: (2025)
FREE: Fast and Robust Vision Language Models with Early Exits
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Transferable Adversarial Attacks on Black-Box Vision-Language Models
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
Efficient Test-Time Scaling for Small Vision-Language Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
Synthetic Data is an Elegant GIFT for Continual Vision-Language Models
by: Wu, Bin, et al.
Published: (2025)
by: Wu, Bin, et al.
Published: (2025)
Similar Items
-
Few-shot target-driven instance detection based on open-vocabulary object detection models
by: Crulis, Ben, et al.
Published: (2024) -
An experimental comparative study of backpropagation and alternatives for training binary neural networks for image classification
by: Crulis, Ben, et al.
Published: (2024) -
Spectral Norm of Convolutional Layers with Circular and Zero Paddings
by: Delattre, Blaise, et al.
Published: (2024) -
Robustness to distribution shifts of compressed networks for edge devices
by: Shen, Lulan, et al.
Published: (2024) -
Block Selective Reprogramming for On-device Training of Vision Transformers
by: Sarkar, Sreetama, et al.
Published: (2024)