Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers
Fuente:
arXiv
Salvato in:
| Autori principali: | Eltahir, Mohamed, Habibullah, Ali, Ayash, Lama, Hussain, Tanveer, Khan, Naeemullah |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs
di: Eltahir, Mohamed, et al.
Pubblicazione: (2026)
di: Eltahir, Mohamed, et al.
Pubblicazione: (2026)
VideoAtlas: Navigating Long-Form Video in Logarithmic Compute
di: Eltahir, Mohamed, et al.
Pubblicazione: (2026)
di: Eltahir, Mohamed, et al.
Pubblicazione: (2026)
PILOT: Policy-Informed Learned Optimization for Adaptive Deep Network Training
di: Altuuaim, Sattam, et al.
Pubblicazione: (2026)
di: Altuuaim, Sattam, et al.
Pubblicazione: (2026)
Zero-Shot Hashing Based on Reconstruction With Part Alignment
di: Jiang, Yan, et al.
Pubblicazione: (2025)
di: Jiang, Yan, et al.
Pubblicazione: (2025)
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
di: Tu, Rong-Cheng, et al.
Pubblicazione: (2025)
di: Tu, Rong-Cheng, et al.
Pubblicazione: (2025)
Visual Zero-Shot E-Commerce Product Attribute Value Extraction
di: Gong, Jiaying, et al.
Pubblicazione: (2025)
di: Gong, Jiaying, et al.
Pubblicazione: (2025)
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
di: Agnolucci, Lorenzo, et al.
Pubblicazione: (2024)
di: Agnolucci, Lorenzo, et al.
Pubblicazione: (2024)
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
di: Tu, Rong-Cheng, et al.
Pubblicazione: (2025)
di: Tu, Rong-Cheng, et al.
Pubblicazione: (2025)
UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
di: Le, Hoang-Bao, et al.
Pubblicazione: (2025)
di: Le, Hoang-Bao, et al.
Pubblicazione: (2025)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
di: Sun, Zelong, et al.
Pubblicazione: (2025)
di: Sun, Zelong, et al.
Pubblicazione: (2025)
Multimodal Language Models for Domain-Specific Procedural Video Summarization
di: Hussain, Nafisa
Pubblicazione: (2024)
di: Hussain, Nafisa
Pubblicazione: (2024)
Interactive Multi-Turn Retrieval for Health Videos
di: Wu, Chengzheng, et al.
Pubblicazione: (2026)
di: Wu, Chengzheng, et al.
Pubblicazione: (2026)
MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion
di: Samuel, Saron, et al.
Pubblicazione: (2025)
di: Samuel, Saron, et al.
Pubblicazione: (2025)
Which Country Is This? Automatic Country Ranking of Street View Photos
di: Menzner, Tim, et al.
Pubblicazione: (2024)
di: Menzner, Tim, et al.
Pubblicazione: (2024)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
di: Gao, Haowen, et al.
Pubblicazione: (2025)
di: Gao, Haowen, et al.
Pubblicazione: (2025)
UniNote: A Unified Embedding Model for Multimodal Representation and Ranking
di: Zhao, Jinghan, et al.
Pubblicazione: (2026)
di: Zhao, Jinghan, et al.
Pubblicazione: (2026)
RDP: Ranked Differential Privacy for Facial Feature Protection in Multiscale Sparsified Subspace
di: Ou, Lu, et al.
Pubblicazione: (2024)
di: Ou, Lu, et al.
Pubblicazione: (2024)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
di: Duan, Yicheng, et al.
Pubblicazione: (2025)
di: Duan, Yicheng, et al.
Pubblicazione: (2025)
Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification
di: Gustineli, Murilo, et al.
Pubblicazione: (2025)
di: Gustineli, Murilo, et al.
Pubblicazione: (2025)
Language-only Efficient Training of Zero-shot Composed Image Retrieval
di: Gu, Geonmo, et al.
Pubblicazione: (2023)
di: Gu, Geonmo, et al.
Pubblicazione: (2023)
Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval
di: Li, Haiwen, et al.
Pubblicazione: (2024)
di: Li, Haiwen, et al.
Pubblicazione: (2024)
MTMD: A Multi-Task Multi-Domain Framework for Unified Ad Lightweight Ranking at Pinterest
di: Yang, Xiao, et al.
Pubblicazione: (2025)
di: Yang, Xiao, et al.
Pubblicazione: (2025)
Advancing Re-Ranking with Multimodal Fusion and Target-Oriented Auxiliary Tasks in E-Commerce Search
di: Xu, Enqiang, et al.
Pubblicazione: (2024)
di: Xu, Enqiang, et al.
Pubblicazione: (2024)
Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
di: Luo, Enming, et al.
Pubblicazione: (2024)
di: Luo, Enming, et al.
Pubblicazione: (2024)
GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval
di: Sun, Chengsong, et al.
Pubblicazione: (2025)
di: Sun, Chengsong, et al.
Pubblicazione: (2025)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
di: Cai, Qifeng, et al.
Pubblicazione: (2025)
di: Cai, Qifeng, et al.
Pubblicazione: (2025)
Zero-shot Composed Image Retrieval Considering Query-target Relationship Leveraging Masked Image-text Pairs
di: Zhang, Huaying, et al.
Pubblicazione: (2024)
di: Zhang, Huaying, et al.
Pubblicazione: (2024)
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
di: Deng, Chenlong, et al.
Pubblicazione: (2026)
di: Deng, Chenlong, et al.
Pubblicazione: (2026)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
di: Sarwar, Nobin
Pubblicazione: (2025)
di: Sarwar, Nobin
Pubblicazione: (2025)
MammoWise: Multi-Model Local RAG Pipeline for Mammography Report Generation
di: Jahangir, Raiyan, et al.
Pubblicazione: (2026)
di: Jahangir, Raiyan, et al.
Pubblicazione: (2026)
BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment
di: Mounis, Mohamed Darwish, et al.
Pubblicazione: (2026)
di: Mounis, Mohamed Darwish, et al.
Pubblicazione: (2026)
HyM-UNet: Synergizing Local Texture and Global Context via Hybrid CNN-Mamba Architecture for Medical Image Segmentation
di: Chen, Haodong, et al.
Pubblicazione: (2025)
di: Chen, Haodong, et al.
Pubblicazione: (2025)
Bridge the Gap between Past and Future: Siamese Model Optimization for Context-Aware Document Ranking
di: Wu, Songhao, et al.
Pubblicazione: (2025)
di: Wu, Songhao, et al.
Pubblicazione: (2025)
VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
di: Aimar, Emanuel Sánchez, et al.
Pubblicazione: (2025)
di: Aimar, Emanuel Sánchez, et al.
Pubblicazione: (2025)
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
di: Zhu, Tianyu, et al.
Pubblicazione: (2024)
di: Zhu, Tianyu, et al.
Pubblicazione: (2024)
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
di: Li, Fanxiao, et al.
Pubblicazione: (2025)
di: Li, Fanxiao, et al.
Pubblicazione: (2025)
SignRAG: A Retrieval-Augmented System for Scalable Zero-Shot Road Sign Recognition
di: Zhu, Minghao, et al.
Pubblicazione: (2025)
di: Zhu, Minghao, et al.
Pubblicazione: (2025)
Sustainable transparency in Recommender Systems: Bayesian Ranking of Images for Explainability
di: Paz-Ruza, Jorge, et al.
Pubblicazione: (2023)
di: Paz-Ruza, Jorge, et al.
Pubblicazione: (2023)
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
di: Luo, Tianci, et al.
Pubblicazione: (2026)
di: Luo, Tianci, et al.
Pubblicazione: (2026)
Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
di: Tam, Jason Kahei, et al.
Pubblicazione: (2025)
di: Tam, Jason Kahei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs
di: Eltahir, Mohamed, et al.
Pubblicazione: (2026) -
VideoAtlas: Navigating Long-Form Video in Logarithmic Compute
di: Eltahir, Mohamed, et al.
Pubblicazione: (2026) -
PILOT: Policy-Informed Learned Optimization for Adaptive Deep Network Training
di: Altuuaim, Sattam, et al.
Pubblicazione: (2026) -
Zero-Shot Hashing Based on Reconstruction With Part Alignment
di: Jiang, Yan, et al.
Pubblicazione: (2025) -
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
di: Tu, Rong-Cheng, et al.
Pubblicazione: (2025)