Infusing fine-grained visual knowledge to Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Ypsilantis, Nikolaos-Antonios, Chen, Kaifeng, Araujo, André, Chum, Ondřej |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
UDON: Universal Dynamic Online distillatioN for generic image representations
por: Ypsilantis, Nikolaos-Antonios, et al.
Publicado: (2024)
por: Ypsilantis, Nikolaos-Antonios, et al.
Publicado: (2024)
Co-Segmentation without any Pixel-level Supervision with Application to Large-Scale Sketch Classification
por: Ypsilantis, Nikolaos-Antonios, et al.
Publicado: (2024)
por: Ypsilantis, Nikolaos-Antonios, et al.
Publicado: (2024)
ILIAS: Instance-Level Image retrieval At Scale
por: Kordopatis-Zilos, Giorgos, et al.
Publicado: (2025)
por: Kordopatis-Zilos, Giorgos, et al.
Publicado: (2025)
Dark Side Augmentation: Generating Diverse Night Examples for Metric Learning
por: Mohwald, Albert, et al.
Publicado: (2023)
por: Mohwald, Albert, et al.
Publicado: (2023)
Crafting Distribution Shifts for Validation and Training in Single Source Domain Generalization
por: Efthymiadis, Nikos, et al.
Publicado: (2024)
por: Efthymiadis, Nikos, et al.
Publicado: (2024)
InPK: Infusing Prior Knowledge into Prompt for Vision-Language Models
por: Zhou, Shuchang, et al.
Publicado: (2025)
por: Zhou, Shuchang, et al.
Publicado: (2025)
Composed Image Retrieval for Training-Free Domain Conversion
por: Efthymiadis, Nikos, et al.
Publicado: (2024)
por: Efthymiadis, Nikos, et al.
Publicado: (2024)
Composed Image Retrieval for Remote Sensing
por: Psomas, Bill, et al.
Publicado: (2024)
por: Psomas, Bill, et al.
Publicado: (2024)
CrossFlowDG: Bridging the Modality Gap with Cross-modal Flow Matching for Domain Generalization
por: Kritikos, Antonios, et al.
Publicado: (2026)
por: Kritikos, Antonios, et al.
Publicado: (2026)
Learning Vision from Models Rivals Learning Vision from Data
por: Tian, Yonglong, et al.
Publicado: (2023)
por: Tian, Yonglong, et al.
Publicado: (2023)
Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking
por: Aiger, Dror, et al.
Publicado: (2025)
por: Aiger, Dror, et al.
Publicado: (2025)
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
por: Jing, Liqiang, et al.
Publicado: (2023)
por: Jing, Liqiang, et al.
Publicado: (2023)
Instance-Level Composed Image Retrieval
por: Psomas, Bill, et al.
Publicado: (2025)
por: Psomas, Bill, et al.
Publicado: (2025)
Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
por: Chen, Honghao, et al.
Publicado: (2025)
por: Chen, Honghao, et al.
Publicado: (2025)
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
por: Wang, Chaoyang, et al.
Publicado: (2025)
por: Wang, Chaoyang, et al.
Publicado: (2025)
Visual RAG: Expanding MLLM visual knowledge without fine-tuning
por: Bonomo, Mirco, et al.
Publicado: (2025)
por: Bonomo, Mirco, et al.
Publicado: (2025)
Koo-Fu CLIP: Closed-Form Adaptation of Vision-Language Models via Fukunaga-Koontz Linear Discriminant Analysis
por: Suchanek, Matej, et al.
Publicado: (2026)
por: Suchanek, Matej, et al.
Publicado: (2026)
Benchmarking Composed Image Retrieval for Applied Earth Observation
por: Psomas, Bill, et al.
Publicado: (2026)
por: Psomas, Bill, et al.
Publicado: (2026)
Let's Roll a BiFTA: Bi-refinement for Fine-grained Text-visual Alignment in Vision-Language Models
por: Sun, Yuhao, et al.
Publicado: (2026)
por: Sun, Yuhao, et al.
Publicado: (2026)
Demographic-aware fine-grained visual recognition of pediatric wrist pathologies
por: Ahmed, Ammar, et al.
Publicado: (2025)
por: Ahmed, Ammar, et al.
Publicado: (2025)
Adapting Vision-Language Model with Fine-grained Semantics for Open-Vocabulary Segmentation
por: Chng, Yong Xien, et al.
Publicado: (2024)
por: Chng, Yong Xien, et al.
Publicado: (2024)
TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
por: Cao, Bingyi, et al.
Publicado: (2026)
por: Cao, Bingyi, et al.
Publicado: (2026)
Yes, we CANN: Constrained Approximate Nearest Neighbors for local feature-based visual localization
por: Aiger, Dror, et al.
Publicado: (2023)
por: Aiger, Dror, et al.
Publicado: (2023)
LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models
por: Gkalelis, Nikolaos, et al.
Publicado: (2026)
por: Gkalelis, Nikolaos, et al.
Publicado: (2026)
Vision Language Model for Interpretable and Fine-grained Detection of Safety Compliance in Diverse Workplaces
por: Chen, Zhiling, et al.
Publicado: (2024)
por: Chen, Zhiling, et al.
Publicado: (2024)
Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
por: Zhang, Wenyao, et al.
Publicado: (2025)
por: Zhang, Wenyao, et al.
Publicado: (2025)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
por: Sharma, Saurav, et al.
Publicado: (2025)
por: Sharma, Saurav, et al.
Publicado: (2025)
Context-Infused Visual Grounding for Art
por: Khan, Selina, et al.
Publicado: (2024)
por: Khan, Selina, et al.
Publicado: (2024)
Is CLIP the main roadblock for fine-grained open-world perception?
por: Bianchi, Lorenzo, et al.
Publicado: (2024)
por: Bianchi, Lorenzo, et al.
Publicado: (2024)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
por: Hong, Wenyi, et al.
Publicado: (2025)
por: Hong, Wenyi, et al.
Publicado: (2025)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
por: Jing, Liqiang, et al.
Publicado: (2024)
por: Jing, Liqiang, et al.
Publicado: (2024)
DAE-Net: Deforming Auto-Encoder for fine-grained shape co-segmentation
por: Chen, Zhiqin, et al.
Publicado: (2023)
por: Chen, Zhiqin, et al.
Publicado: (2023)
SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning
por: Thoker, Fida Mohammad, et al.
Publicado: (2025)
por: Thoker, Fida Mohammad, et al.
Publicado: (2025)
Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
por: Wang, Ruiyu, et al.
Publicado: (2025)
por: Wang, Ruiyu, et al.
Publicado: (2025)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
por: Demidov, Dmitry, et al.
Publicado: (2025)
por: Demidov, Dmitry, et al.
Publicado: (2025)
Specificity-aware reinforcement learning for fine-grained open-world classification
por: Angheben, Samuele, et al.
Publicado: (2026)
por: Angheben, Samuele, et al.
Publicado: (2026)
Large Language Models estimate fine-grained human color-concept associations
por: Mukherjee, Kushin, et al.
Publicado: (2024)
por: Mukherjee, Kushin, et al.
Publicado: (2024)
MExD: An Expert-Infused Diffusion Model for Whole-Slide Image Classification
por: Zhao, Jianwei, et al.
Publicado: (2025)
por: Zhao, Jianwei, et al.
Publicado: (2025)
Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention
por: Liu, Ying, et al.
Publicado: (2024)
por: Liu, Ying, et al.
Publicado: (2024)
An Inpainting-Infused Pipeline for Attire and Background Replacement
por: Perche-Mahlow, Felipe Rodrigues, et al.
Publicado: (2024)
por: Perche-Mahlow, Felipe Rodrigues, et al.
Publicado: (2024)
Ejemplares similares
-
UDON: Universal Dynamic Online distillatioN for generic image representations
por: Ypsilantis, Nikolaos-Antonios, et al.
Publicado: (2024) -
Co-Segmentation without any Pixel-level Supervision with Application to Large-Scale Sketch Classification
por: Ypsilantis, Nikolaos-Antonios, et al.
Publicado: (2024) -
ILIAS: Instance-Level Image retrieval At Scale
por: Kordopatis-Zilos, Giorgos, et al.
Publicado: (2025) -
Dark Side Augmentation: Generating Diverse Night Examples for Metric Learning
por: Mohwald, Albert, et al.
Publicado: (2023) -
Crafting Distribution Shifts for Validation and Training in Single Source Domain Generalization
por: Efthymiadis, Nikos, et al.
Publicado: (2024)