Enregistré dans:
| Auteurs principaux: | Eltahir, Mohamed, Habibullah, Ali, Ayash, Lama, Hussain, Tanveer, Khan, Naeemullah |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2511.01617 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs
par: Eltahir, Mohamed, et autres
Publié: (2026)
par: Eltahir, Mohamed, et autres
Publié: (2026)
VideoAtlas: Navigating Long-Form Video in Logarithmic Compute
par: Eltahir, Mohamed, et autres
Publié: (2026)
par: Eltahir, Mohamed, et autres
Publié: (2026)
PILOT: Policy-Informed Learned Optimization for Adaptive Deep Network Training
par: Altuuaim, Sattam, et autres
Publié: (2026)
par: Altuuaim, Sattam, et autres
Publié: (2026)
Zero-Shot Hashing Based on Reconstruction With Part Alignment
par: Jiang, Yan, et autres
Publié: (2025)
par: Jiang, Yan, et autres
Publié: (2025)
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
par: Tu, Rong-Cheng, et autres
Publié: (2025)
par: Tu, Rong-Cheng, et autres
Publié: (2025)
Visual Zero-Shot E-Commerce Product Attribute Value Extraction
par: Gong, Jiaying, et autres
Publié: (2025)
par: Gong, Jiaying, et autres
Publié: (2025)
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
par: Agnolucci, Lorenzo, et autres
Publié: (2024)
par: Agnolucci, Lorenzo, et autres
Publié: (2024)
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
par: Tu, Rong-Cheng, et autres
Publié: (2025)
par: Tu, Rong-Cheng, et autres
Publié: (2025)
UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
par: Le, Hoang-Bao, et autres
Publié: (2025)
par: Le, Hoang-Bao, et autres
Publié: (2025)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
par: Sun, Zelong, et autres
Publié: (2025)
par: Sun, Zelong, et autres
Publié: (2025)
Multimodal Language Models for Domain-Specific Procedural Video Summarization
par: Hussain, Nafisa
Publié: (2024)
par: Hussain, Nafisa
Publié: (2024)
Interactive Multi-Turn Retrieval for Health Videos
par: Wu, Chengzheng, et autres
Publié: (2026)
par: Wu, Chengzheng, et autres
Publié: (2026)
MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion
par: Samuel, Saron, et autres
Publié: (2025)
par: Samuel, Saron, et autres
Publié: (2025)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
par: Duan, Yicheng, et autres
Publié: (2025)
par: Duan, Yicheng, et autres
Publié: (2025)
Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification
par: Gustineli, Murilo, et autres
Publié: (2025)
par: Gustineli, Murilo, et autres
Publié: (2025)
Which Country Is This? Automatic Country Ranking of Street View Photos
par: Menzner, Tim, et autres
Publié: (2024)
par: Menzner, Tim, et autres
Publié: (2024)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
par: Gao, Haowen, et autres
Publié: (2025)
par: Gao, Haowen, et autres
Publié: (2025)
UniNote: A Unified Embedding Model for Multimodal Representation and Ranking
par: Zhao, Jinghan, et autres
Publié: (2026)
par: Zhao, Jinghan, et autres
Publié: (2026)
Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
par: Luo, Enming, et autres
Publié: (2024)
par: Luo, Enming, et autres
Publié: (2024)
RDP: Ranked Differential Privacy for Facial Feature Protection in Multiscale Sparsified Subspace
par: Ou, Lu, et autres
Publié: (2024)
par: Ou, Lu, et autres
Publié: (2024)
Language-only Efficient Training of Zero-shot Composed Image Retrieval
par: Gu, Geonmo, et autres
Publié: (2023)
par: Gu, Geonmo, et autres
Publié: (2023)
Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval
par: Li, Haiwen, et autres
Publié: (2024)
par: Li, Haiwen, et autres
Publié: (2024)
MTMD: A Multi-Task Multi-Domain Framework for Unified Ad Lightweight Ranking at Pinterest
par: Yang, Xiao, et autres
Publié: (2025)
par: Yang, Xiao, et autres
Publié: (2025)
Advancing Re-Ranking with Multimodal Fusion and Target-Oriented Auxiliary Tasks in E-Commerce Search
par: Xu, Enqiang, et autres
Publié: (2024)
par: Xu, Enqiang, et autres
Publié: (2024)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
par: Sarwar, Nobin
Publié: (2025)
par: Sarwar, Nobin
Publié: (2025)
GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval
par: Sun, Chengsong, et autres
Publié: (2025)
par: Sun, Chengsong, et autres
Publié: (2025)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
par: Cai, Qifeng, et autres
Publié: (2025)
par: Cai, Qifeng, et autres
Publié: (2025)
Bridge the Gap between Past and Future: Siamese Model Optimization for Context-Aware Document Ranking
par: Wu, Songhao, et autres
Publié: (2025)
par: Wu, Songhao, et autres
Publié: (2025)
Zero-shot Composed Image Retrieval Considering Query-target Relationship Leveraging Masked Image-text Pairs
par: Zhang, Huaying, et autres
Publié: (2024)
par: Zhang, Huaying, et autres
Publié: (2024)
MammoWise: Multi-Model Local RAG Pipeline for Mammography Report Generation
par: Jahangir, Raiyan, et autres
Publié: (2026)
par: Jahangir, Raiyan, et autres
Publié: (2026)
BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment
par: Mounis, Mohamed Darwish, et autres
Publié: (2026)
par: Mounis, Mohamed Darwish, et autres
Publié: (2026)
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
par: Deng, Chenlong, et autres
Publié: (2026)
par: Deng, Chenlong, et autres
Publié: (2026)
VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
par: Aimar, Emanuel Sánchez, et autres
Publié: (2025)
par: Aimar, Emanuel Sánchez, et autres
Publié: (2025)
HyM-UNet: Synergizing Local Texture and Global Context via Hybrid CNN-Mamba Architecture for Medical Image Segmentation
par: Chen, Haodong, et autres
Publié: (2025)
par: Chen, Haodong, et autres
Publié: (2025)
SignRAG: A Retrieval-Augmented System for Scalable Zero-Shot Road Sign Recognition
par: Zhu, Minghao, et autres
Publié: (2025)
par: Zhu, Minghao, et autres
Publié: (2025)
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
par: Zhu, Tianyu, et autres
Publié: (2024)
par: Zhu, Tianyu, et autres
Publié: (2024)
Sustainable transparency in Recommender Systems: Bayesian Ranking of Images for Explainability
par: Paz-Ruza, Jorge, et autres
Publié: (2023)
par: Paz-Ruza, Jorge, et autres
Publié: (2023)
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
par: Li, Fanxiao, et autres
Publié: (2025)
par: Li, Fanxiao, et autres
Publié: (2025)
Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
par: Tam, Jason Kahei, et autres
Publié: (2025)
par: Tam, Jason Kahei, et autres
Publié: (2025)
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
par: Luo, Tianci, et autres
Publié: (2026)
par: Luo, Tianci, et autres
Publié: (2026)
Documents similaires
-
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs
par: Eltahir, Mohamed, et autres
Publié: (2026) -
VideoAtlas: Navigating Long-Form Video in Logarithmic Compute
par: Eltahir, Mohamed, et autres
Publié: (2026) -
PILOT: Policy-Informed Learned Optimization for Adaptive Deep Network Training
par: Altuuaim, Sattam, et autres
Publié: (2026) -
Zero-Shot Hashing Based on Reconstruction With Part Alignment
par: Jiang, Yan, et autres
Publié: (2025) -
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
par: Tu, Rong-Cheng, et autres
Publié: (2025)