Preserving Localized Patch Semantics in VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Esmaeilkhani, Parsa, Latecki, Longin Jan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Direct Visual Grounding by Directing Attention of Visual Tokens
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
Layout Stroke Imitation: A Layout Guided Handwriting Stroke Generation for Style Imitation with Diffusion Model
by: Hanif, Sidra, et al.
Published: (2025)
by: Hanif, Sidra, et al.
Published: (2025)
Learning Object Focused Attention
by: Trivedy, Vivek, et al.
Published: (2025)
by: Trivedy, Vivek, et al.
Published: (2025)
VAPO: Visibility-Aware Keypoint Localization for Efficient 6DoF Object Pose Estimation
by: Lian, Ruyi, et al.
Published: (2024)
by: Lian, Ruyi, et al.
Published: (2024)
Enhancing Renal Tumor Malignancy Prediction: Deep Learning with Automatic 3D CT Organ Focused Attention
by: Fan, Zhengkang, et al.
Published: (2026)
by: Fan, Zhengkang, et al.
Published: (2026)
FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding
by: Pan, Huitong, et al.
Published: (2024)
by: Pan, Huitong, et al.
Published: (2024)
From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance
by: Aryashad, Ardalan, et al.
Published: (2025)
by: Aryashad, Ardalan, et al.
Published: (2025)
Cross-Domain Semantic Segmentation on Inconsistent Taxonomy using VLMs
by: Lim, Jeongkee, et al.
Published: (2024)
by: Lim, Jeongkee, et al.
Published: (2024)
SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMs
by: Xu, Shuhan, et al.
Published: (2025)
by: Xu, Shuhan, et al.
Published: (2025)
Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data
by: Lei, Tianhao, et al.
Published: (2026)
by: Lei, Tianhao, et al.
Published: (2026)
VLMs Can Aggregate Scattered Training Patches
by: Zhou, Zhanhui, et al.
Published: (2025)
by: Zhou, Zhanhui, et al.
Published: (2025)
Teaching VLMs to Localize Specific Objects from In-context Examples
by: Doveh, Sivan, et al.
Published: (2024)
by: Doveh, Sivan, et al.
Published: (2024)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
by: Chinchure, Aditya, et al.
Published: (2025)
by: Chinchure, Aditya, et al.
Published: (2025)
MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
by: Liu, Liu, et al.
Published: (2025)
by: Liu, Liu, et al.
Published: (2025)
Structure-Preserving Patch Decoding for Efficient Neural Video Representation
by: Hayami, Taiga, et al.
Published: (2025)
by: Hayami, Taiga, et al.
Published: (2025)
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
by: Fang, Zhengyao, et al.
Published: (2026)
by: Fang, Zhengyao, et al.
Published: (2026)
Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
by: Benschop, Pascal, et al.
Published: (2026)
by: Benschop, Pascal, et al.
Published: (2026)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
by: Lee, Jaehoon, et al.
Published: (2026)
by: Lee, Jaehoon, et al.
Published: (2026)
Adaptive Patch Contrast for Weakly Supervised Semantic Segmentation
by: Wu, Wangyu, et al.
Published: (2024)
by: Wu, Wangyu, et al.
Published: (2024)
SVIP: Semantically Contextualized Visual Patches for Zero-Shot Learning
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
Adversarial Patch for 3D Local Feature Extractor
by: Pao, Yu Wen, et al.
Published: (2024)
by: Pao, Yu Wen, et al.
Published: (2024)
Patch Ranking: Efficient CLIP by Learning to Rank Local Patches
by: Wu, Cheng-En, et al.
Published: (2024)
by: Wu, Cheng-En, et al.
Published: (2024)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
Bridging Semantic Logic Gaps: A Cognition Inspired Multimodal Boundary Preserving Network for Image Manipulation Localization
by: Li, Songlin, et al.
Published: (2025)
by: Li, Songlin, et al.
Published: (2025)
Deep Pre-Alignment for VLMs
by: Yu, Tianyu, et al.
Published: (2026)
by: Yu, Tianyu, et al.
Published: (2026)
HPFF: Hierarchical Locally Supervised Learning with Patch Feature Fusion
by: Su, Junhao, et al.
Published: (2024)
by: Su, Junhao, et al.
Published: (2024)
Hierarchical Salient Patch Identification for Interpretable Fundus Disease Localization
by: Peng, Yitao, et al.
Published: (2024)
by: Peng, Yitao, et al.
Published: (2024)
Adversarial Patch Attack for Ship Detection via Localized Augmentation
by: Liu, Chun, et al.
Published: (2025)
by: Liu, Chun, et al.
Published: (2025)
PATS: Patch Area Transportation with Subdivision for Local Feature Matching
by: Ni, Junjie, et al.
Published: (2023)
by: Ni, Junjie, et al.
Published: (2023)
Privacy Preserving Ordinal-Meta Learning with VLMs for Fine-Grained Fruit Quality Prediction
by: Jain, Riddhi, et al.
Published: (2025)
by: Jain, Riddhi, et al.
Published: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
by: Wang, Dianyi, et al.
Published: (2025)
by: Wang, Dianyi, et al.
Published: (2025)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
by: Shahgir, Haz Sameen, et al.
Published: (2026)
by: Shahgir, Haz Sameen, et al.
Published: (2026)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Context Patch Fusion With Class Token Enhancement for Weakly Supervised Semantic Segmentation
by: Fu, Yiyang, et al.
Published: (2026)
by: Fu, Yiyang, et al.
Published: (2026)
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
by: Zhang, Dengke, et al.
Published: (2024)
by: Zhang, Dengke, et al.
Published: (2024)
Are VLMs Really Blind
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
by: Wen, Ziqi, et al.
Published: (2026)
by: Wen, Ziqi, et al.
Published: (2026)
Regression in EO: Are VLMs Up to the Challenge?
by: Xue, Xizhe, et al.
Published: (2025)
by: Xue, Xizhe, et al.
Published: (2025)
GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning
by: Peng, Tiantian, et al.
Published: (2025)
by: Peng, Tiantian, et al.
Published: (2025)
Similar Items
-
Direct Visual Grounding by Directing Attention of Visual Tokens
by: Esmaeilkhani, Parsa, et al.
Published: (2025) -
Layout Stroke Imitation: A Layout Guided Handwriting Stroke Generation for Style Imitation with Diffusion Model
by: Hanif, Sidra, et al.
Published: (2025) -
Learning Object Focused Attention
by: Trivedy, Vivek, et al.
Published: (2025) -
VAPO: Visibility-Aware Keypoint Localization for Efficient 6DoF Object Pose Estimation
by: Lian, Ruyi, et al.
Published: (2024) -
Enhancing Renal Tumor Malignancy Prediction: Deep Learning with Automatic 3D CT Organ Focused Attention
by: Fan, Zhengkang, et al.
Published: (2026)