VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Ziyan, Meng, Rui, Yang, Xinyi, Yavuz, Semih, Zhou, Yingbo, Chen, Wenhu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
by: Meng, Rui, et al.
Published: (2025)
by: Meng, Rui, et al.
Published: (2025)
Retrofitting Small Multilingual Models for Retrieval: Matching 7B Performance with 300M Parameters
by: Tu, Lifu, et al.
Published: (2025)
by: Tu, Lifu, et al.
Published: (2025)
Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
Traffic Light or Light Traffic? Investigating Phrasal Semantics in Large Language Models
by: Meng, Rui, et al.
Published: (2024)
by: Meng, Rui, et al.
Published: (2024)
From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
by: Sheta, Hala, et al.
Published: (2025)
by: Sheta, Hala, et al.
Published: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
by: Yang, Jianjiang, et al.
Published: (2025)
by: Yang, Jianjiang, et al.
Published: (2025)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task
by: Tian, Yu, et al.
Published: (2024)
by: Tian, Yu, et al.
Published: (2024)
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
by: Pan, Xichen, et al.
Published: (2023)
by: Pan, Xichen, et al.
Published: (2023)
Investigating Factuality in Long-Form Text Generation: The Roles of Self-Known and Self-Unknown
by: Tu, Lifu, et al.
Published: (2024)
by: Tu, Lifu, et al.
Published: (2024)
VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
by: Aimar, Emanuel Sánchez, et al.
Published: (2025)
by: Aimar, Emanuel Sánchez, et al.
Published: (2025)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
by: Jia, Mengzhao, et al.
Published: (2024)
by: Jia, Mengzhao, et al.
Published: (2024)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
by: Chen, Xinyi, et al.
Published: (2023)
by: Chen, Xinyi, et al.
Published: (2023)
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
by: Qu, Kevin, et al.
Published: (2026)
by: Qu, Kevin, et al.
Published: (2026)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
by: Fan, Zhiwen, et al.
Published: (2025)
by: Fan, Zhiwen, et al.
Published: (2025)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
by: Wang, Zihang, et al.
Published: (2026)
by: Wang, Zihang, et al.
Published: (2026)
PersonaVLM: Long-Term Personalized Multimodal LLMs
by: Nie, Chang, et al.
Published: (2026)
by: Nie, Chang, et al.
Published: (2026)
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
by: Wang, Zekun, et al.
Published: (2025)
by: Wang, Zekun, et al.
Published: (2025)
OViP: Online Vision-Language Preference Learning for VLM Hallucination
by: Liu, Shujun, et al.
Published: (2025)
by: Liu, Shujun, et al.
Published: (2025)
SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
by: Lim, Gyubeum, et al.
Published: (2025)
by: Lim, Gyubeum, et al.
Published: (2025)
From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models
by: Wang, Qidong, et al.
Published: (2026)
by: Wang, Qidong, et al.
Published: (2026)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
by: Lokesh, K, et al.
Published: (2026)
by: Lokesh, K, et al.
Published: (2026)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
by: Wei, Xinyu, et al.
Published: (2025)
by: Wei, Xinyu, et al.
Published: (2025)
SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model
by: Deng, Guifeng, et al.
Published: (2026)
by: Deng, Guifeng, et al.
Published: (2026)
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
by: Shen, Haozhan, et al.
Published: (2025)
by: Shen, Haozhan, et al.
Published: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
Lost in Embeddings: Information Loss in Vision-Language Models
by: Li, Wenyan, et al.
Published: (2025)
by: Li, Wenyan, et al.
Published: (2025)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
by: Zhang, Zhaoyang, et al.
Published: (2023)
by: Zhang, Zhaoyang, et al.
Published: (2023)
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
by: Cho, Seunghyuk, et al.
Published: (2025)
by: Cho, Seunghyuk, et al.
Published: (2025)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
by: Zhu, Wenxin, et al.
Published: (2025)
by: Zhu, Wenxin, et al.
Published: (2025)
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
by: Ebouky, Brown, et al.
Published: (2026)
by: Ebouky, Brown, et al.
Published: (2026)
Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
by: Zhao, Pu, et al.
Published: (2025)
by: Zhao, Pu, et al.
Published: (2025)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
by: Kumar, Divake, et al.
Published: (2026)
by: Kumar, Divake, et al.
Published: (2026)
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
by: Li, Yanshu
Published: (2025)
by: Li, Yanshu
Published: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models Decoding
by: Tu, Lifu, et al.
Published: (2023)
by: Tu, Lifu, et al.
Published: (2023)
MIEB: Massive Image Embedding Benchmark
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
by: Zhang, Jipeng, et al.
Published: (2025)
by: Zhang, Jipeng, et al.
Published: (2025)
Similar Items
-
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
by: Meng, Rui, et al.
Published: (2025) -
Retrofitting Small Multilingual Models for Retrieval: Matching 7B Performance with 300M Parameters
by: Tu, Lifu, et al.
Published: (2025) -
Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
by: Thirukovalluru, Raghuveer, et al.
Published: (2025) -
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
by: Lu, Yujie, et al.
Published: (2024) -
Traffic Light or Light Traffic? Investigating Phrasal Semantics in Large Language Models
by: Meng, Rui, et al.
Published: (2024)