Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features
Fuente:
arXiv
Saved in:
| Main Authors: | Sengupta, Saurav, Moradinasab, Nazanin, Liu, Jiebei, Brown, Donald E. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions
by: Sengupta, Saurav, et al.
Published: (2025)
by: Sengupta, Saurav, et al.
Published: (2025)
Towards Robust Multimodal Representation: A Unified Approach with Adaptive Experts and Alignment
by: Moradinasab, Nazanin, et al.
Published: (2025)
by: Moradinasab, Nazanin, et al.
Published: (2025)
GenGMM: Generalized Gaussian-Mixture-based Domain Adaptation Model for Semantic Segmentation
by: Moradinasab, Nazanin, et al.
Published: (2024)
by: Moradinasab, Nazanin, et al.
Published: (2024)
Automatic Report Generation for Histopathology images using pre-trained Vision Transformers and BERT
by: Sengupta, Saurav, et al.
Published: (2023)
by: Sengupta, Saurav, et al.
Published: (2023)
ProtoGMM: Multi-prototype Gaussian-Mixture-based Domain Adaptation Model for Semantic Segmentation
by: Moradinasab, Nazanin, et al.
Published: (2024)
by: Moradinasab, Nazanin, et al.
Published: (2024)
Label-efficient Contrastive Learning-based model for nuclei detection and classification in 3D Cardiovascular Immunofluorescent Images
by: Moradinasab, Nazanin, et al.
Published: (2023)
by: Moradinasab, Nazanin, et al.
Published: (2023)
Vision-Language Models for Infrared Industrial Sensing in Additive Manufacturing Scene Description
by: Mahjourian, Nazanin, et al.
Published: (2025)
by: Mahjourian, Nazanin, et al.
Published: (2025)
Sanitizing Manufacturing Dataset Labels Using Vision-Language Models
by: Mahjourian, Nazanin, et al.
Published: (2025)
by: Mahjourian, Nazanin, et al.
Published: (2025)
CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language Models
by: Jha, Saurav, et al.
Published: (2024)
by: Jha, Saurav, et al.
Published: (2024)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
by: Sharma, Saurav, et al.
Published: (2025)
by: Sharma, Saurav, et al.
Published: (2025)
UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language Models
by: Liang, Jiachen, et al.
Published: (2024)
by: Liang, Jiachen, et al.
Published: (2024)
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
by: Huang, Huang, et al.
Published: (2025)
by: Huang, Huang, et al.
Published: (2025)
Emergence of Text Readability in Vision Language Models
by: Park, Jaeyoo, et al.
Published: (2025)
by: Park, Jaeyoo, et al.
Published: (2025)
MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving
by: Zhang, Enming, et al.
Published: (2024)
by: Zhang, Enming, et al.
Published: (2024)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
by: Waseda, Futa, et al.
Published: (2025)
by: Waseda, Futa, et al.
Published: (2025)
Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Text Prompt Injection of Vision Language Models
by: Zhu, Ruizhe
Published: (2025)
by: Zhu, Ruizhe
Published: (2025)
Multi-modal Attribute Prompting for Vision-Language Models
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
by: Jia, Mengzhao, et al.
Published: (2024)
by: Jia, Mengzhao, et al.
Published: (2024)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection
by: Zighem, Mohammed-En-Nadhir, et al.
Published: (2025)
by: Zighem, Mohammed-En-Nadhir, et al.
Published: (2025)
Detecting Text Manipulation in Images using Vision Language Models
by: Vidit, Vidit, et al.
Published: (2025)
by: Vidit, Vidit, et al.
Published: (2025)
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Text Promptable Surgical Instrument Segmentation with Vision-Language Models
by: Zhou, Zijian, et al.
Published: (2023)
by: Zhou, Zijian, et al.
Published: (2023)
An Examination of the Compositionality of Large Generative Vision-Language Models
by: Ma, Teli, et al.
Published: (2023)
by: Ma, Teli, et al.
Published: (2023)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
by: Liu, Qing'an, et al.
Published: (2026)
by: Liu, Qing'an, et al.
Published: (2026)
AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
by: Yue, Tongtian, et al.
Published: (2025)
by: Yue, Tongtian, et al.
Published: (2025)
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
by: Li, Haosheng, et al.
Published: (2026)
by: Li, Haosheng, et al.
Published: (2026)
Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models
by: Choi, In Chong, et al.
Published: (2026)
by: Choi, In Chong, et al.
Published: (2026)
UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models
by: Li, Yujie, et al.
Published: (2024)
by: Li, Yujie, et al.
Published: (2024)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
by: Hannan, Tanveer, et al.
Published: (2025)
by: Hannan, Tanveer, et al.
Published: (2025)
Collaborative Multi-Mode Pruning for Vision-Language Models
by: Wu, Zimeng, et al.
Published: (2026)
by: Wu, Zimeng, et al.
Published: (2026)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
by: Gao, Peng, et al.
Published: (2021)
by: Gao, Peng, et al.
Published: (2021)
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
by: Yakun, Cui, et al.
Published: (2026)
by: Yakun, Cui, et al.
Published: (2026)
Similar Items
-
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions
by: Sengupta, Saurav, et al.
Published: (2025) -
Towards Robust Multimodal Representation: A Unified Approach with Adaptive Experts and Alignment
by: Moradinasab, Nazanin, et al.
Published: (2025) -
GenGMM: Generalized Gaussian-Mixture-based Domain Adaptation Model for Semantic Segmentation
by: Moradinasab, Nazanin, et al.
Published: (2024) -
Automatic Report Generation for Histopathology images using pre-trained Vision Transformers and BERT
by: Sengupta, Saurav, et al.
Published: (2023) -
ProtoGMM: Multi-prototype Gaussian-Mixture-based Domain Adaptation Model for Semantic Segmentation
by: Moradinasab, Nazanin, et al.
Published: (2024)