Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhixin, Zhang, Yiyuan, Ding, Xiaohan, Yue, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising
by: Fu, Chenghan, et al.
Published: (2025)
by: Fu, Chenghan, et al.
Published: (2025)
Closing the Modality Gap for Mixed Modality Search
by: Li, Binxu, et al.
Published: (2025)
by: Li, Binxu, et al.
Published: (2025)
GENIUS: A Generative Framework for Universal Multimodal Search
by: Kim, Sungyeon, et al.
Published: (2025)
by: Kim, Sungyeon, et al.
Published: (2025)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Scaling Up Your Kernels: Large Kernel Design in ConvNets towards Universal Representations
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
by: Williams-Lekuona, Mikel, et al.
Published: (2025)
by: Williams-Lekuona, Mikel, et al.
Published: (2025)
Learning-Based Hashing for ANN Search: Foundations and Early Advances
by: Moran, Sean
Published: (2025)
by: Moran, Sean
Published: (2025)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
V-Agent: An Interactive Video Search System Using Vision-Language Models
by: Park, SunYoung, et al.
Published: (2025)
by: Park, SunYoung, et al.
Published: (2025)
LookSync: Large-Scale Visual Product Search System for AI-Generated Fashion Looks
by: M, Pradeep, et al.
Published: (2025)
by: M, Pradeep, et al.
Published: (2025)
Character-based Outfit Generation with Vision-augmented Style Extraction via LLMs
by: Forouzandehmehr, Najmeh, et al.
Published: (2024)
by: Forouzandehmehr, Najmeh, et al.
Published: (2024)
Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships
by: Waseda, Futa, et al.
Published: (2024)
by: Waseda, Futa, et al.
Published: (2024)
When Search Engine Services meet Large Language Models: Visions and Challenges
by: Xiong, Haoyi, et al.
Published: (2024)
by: Xiong, Haoyi, et al.
Published: (2024)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Reasoning-Augmented Representations for Multimodal Retrieval
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Open Multimodal Retrieval-Augmented Factual Image Generation
by: Tian, Yang, et al.
Published: (2025)
by: Tian, Yang, et al.
Published: (2025)
EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
by: Wang, Peng, et al.
Published: (2023)
by: Wang, Peng, et al.
Published: (2023)
Smart Routing for Multimodal Video Retrieval: When to Search What
by: Rosa, Kevin Dela
Published: (2025)
by: Rosa, Kevin Dela
Published: (2025)
MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
by: Zhang, Daoze, et al.
Published: (2025)
by: Zhang, Daoze, et al.
Published: (2025)
MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding
by: Wu, Junxian, et al.
Published: (2026)
by: Wu, Junxian, et al.
Published: (2026)
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
by: Nie, Zhanheng, et al.
Published: (2025)
by: Nie, Zhanheng, et al.
Published: (2025)
Multimodal RAG Enhanced Visual Description
by: Jaiswal, Amit Kumar, et al.
Published: (2025)
by: Jaiswal, Amit Kumar, et al.
Published: (2025)
ImplicitAVE: An Open-Source Dataset and Multimodal LLMs Benchmark for Implicit Attribute Value Extraction
by: Zou, Henry Peng, et al.
Published: (2024)
by: Zou, Henry Peng, et al.
Published: (2024)
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
by: Duan, Yicheng, et al.
Published: (2025)
by: Duan, Yicheng, et al.
Published: (2025)
Knowledge-Augmented Vision Language Models for Underwater Bioacoustic Spectrogram Analysis
by: Nihal, Ragib Amin, et al.
Published: (2025)
by: Nihal, Ragib Amin, et al.
Published: (2025)
Siamese Content-based Search Engine for a More Transparent Skin and Breast Cancer Diagnosis through Histological Imaging
by: Tabatabaei, Zahra, et al.
Published: (2024)
by: Tabatabaei, Zahra, et al.
Published: (2024)
FeatUp: A Model-Agnostic Framework for Features at Any Resolution
by: Fu, Stephanie, et al.
Published: (2024)
by: Fu, Stephanie, et al.
Published: (2024)
Feedback-based Modal Mutual Search for Attacking Vision-Language Pre-training Models
by: Ding, Renhua, et al.
Published: (2024)
by: Ding, Renhua, et al.
Published: (2024)
WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
Read and Think: An Efficient Step-wise Multimodal Language Model for Document Understanding and Reasoning
by: Zhang, Jinxu
Published: (2024)
by: Zhang, Jinxu
Published: (2024)
RAVEN: Multitask Retrieval Augmented Vision-Language Learning
by: Rao, Varun Nagaraj, et al.
Published: (2024)
by: Rao, Varun Nagaraj, et al.
Published: (2024)
Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation
by: Luo, Weiqing, et al.
Published: (2026)
by: Luo, Weiqing, et al.
Published: (2026)
Modality-Aware Integration with Large Language Models for Knowledge-based Visual Question Answering
by: Dong, Junnan, et al.
Published: (2024)
by: Dong, Junnan, et al.
Published: (2024)
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
by: Fan, Wan-Cyuan, et al.
Published: (2024)
by: Fan, Wan-Cyuan, et al.
Published: (2024)
Infinite Video Understanding
by: Zhang, Dell, et al.
Published: (2025)
by: Zhang, Dell, et al.
Published: (2025)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
by: Lin, Sheng-Chieh, et al.
Published: (2024)
by: Lin, Sheng-Chieh, et al.
Published: (2024)
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
by: Guo, Zhuoning, et al.
Published: (2025)
by: Guo, Zhuoning, et al.
Published: (2025)
Online Vectorized HD Map Construction using Geometry
by: Zhang, Zhixin, et al.
Published: (2023)
by: Zhang, Zhixin, et al.
Published: (2023)
Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models
by: Yada, Yuki, et al.
Published: (2025)
by: Yada, Yuki, et al.
Published: (2025)
Similar Items
-
MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising
by: Fu, Chenghan, et al.
Published: (2025) -
Closing the Modality Gap for Mixed Modality Search
by: Li, Binxu, et al.
Published: (2025) -
GENIUS: A Generative Framework for Universal Multimodal Search
by: Kim, Sungyeon, et al.
Published: (2025) -
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024) -
Scaling Up Your Kernels: Large Kernel Design in ConvNets towards Universal Representations
by: Zhang, Yiyuan, et al.
Published: (2024)