Gespeichert in:
| Hauptverfasser: | Wang, Zhenyu, Li, Wenjia, Zhu, Pengyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.14332 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search
von: Tang, Hengzhu, et al.
Veröffentlicht: (2025)
von: Tang, Hengzhu, et al.
Veröffentlicht: (2025)
Music Recommendation Based on Facial Emotion Recognition
von: B, Rajesh, et al.
Veröffentlicht: (2024)
von: B, Rajesh, et al.
Veröffentlicht: (2024)
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
von: Meng, GuangHao, et al.
Veröffentlicht: (2025)
von: Meng, GuangHao, et al.
Veröffentlicht: (2025)
I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
Rethinking Detection Based Table Structure Recognition for Visually Rich Document Images
von: Xiao, Bin, et al.
Veröffentlicht: (2023)
von: Xiao, Bin, et al.
Veröffentlicht: (2023)
MTMD: A Multi-Task Multi-Domain Framework for Unified Ad Lightweight Ranking at Pinterest
von: Yang, Xiao, et al.
Veröffentlicht: (2025)
von: Yang, Xiao, et al.
Veröffentlicht: (2025)
DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model Generalization
von: Lv, Zheqi, et al.
Veröffentlicht: (2022)
von: Lv, Zheqi, et al.
Veröffentlicht: (2022)
UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
von: Le, Hoang-Bao, et al.
Veröffentlicht: (2025)
von: Le, Hoang-Bao, et al.
Veröffentlicht: (2025)
YOLO-Vehicle-Pro: A Cloud-Edge Collaborative Framework for Object Detection in Autonomous Driving under Adverse Weather Conditions
von: Li, Xiguang, et al.
Veröffentlicht: (2024)
von: Li, Xiguang, et al.
Veröffentlicht: (2024)
CoLLM: A Large Language Model for Composed Image Retrieval
von: Huynh, Chuong, et al.
Veröffentlicht: (2025)
von: Huynh, Chuong, et al.
Veröffentlicht: (2025)
FF-PNet: A Pyramid Network Based on Feature and Field for Brain Image Registration
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
von: Zhu, Yingjian, et al.
Veröffentlicht: (2026)
von: Zhu, Yingjian, et al.
Veröffentlicht: (2026)
Zero-Shot Hashing Based on Reconstruction With Part Alignment
von: Jiang, Yan, et al.
Veröffentlicht: (2025)
von: Jiang, Yan, et al.
Veröffentlicht: (2025)
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
von: De Nadai, Marco, et al.
Veröffentlicht: (2025)
von: De Nadai, Marco, et al.
Veröffentlicht: (2025)
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
Leveraging Foundation Models for Content-Based Image Retrieval in Radiology
von: Denner, Stefan, et al.
Veröffentlicht: (2024)
von: Denner, Stefan, et al.
Veröffentlicht: (2024)
UniNote: A Unified Embedding Model for Multimodal Representation and Ranking
von: Zhao, Jinghan, et al.
Veröffentlicht: (2026)
von: Zhao, Jinghan, et al.
Veröffentlicht: (2026)
Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal Framework
von: Ortego, Diego, et al.
Veröffentlicht: (2025)
von: Ortego, Diego, et al.
Veröffentlicht: (2025)
Personalized Video Summarization using Text-Based Queries and Conditional Modeling
von: Huang, Jia-Hong
Veröffentlicht: (2024)
von: Huang, Jia-Hong
Veröffentlicht: (2024)
Rethinking Sparse Lexical Representations for Image Retrieval in the Age of Rising Multi-Modal Large Language Models
von: Nakata, Kengo, et al.
Veröffentlicht: (2024)
von: Nakata, Kengo, et al.
Veröffentlicht: (2024)
Iterative Optimal Attention and Local Model for Single Image Rain Streak Removal
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
TrajSV: A Trajectory-based Model for Sports Video Representations and Applications
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
von: Wang, Zheng, et al.
Veröffentlicht: (2025)
MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark
von: Osmulski, Radek, et al.
Veröffentlicht: (2025)
von: Osmulski, Radek, et al.
Veröffentlicht: (2025)
EndoFinder: Online Image Retrieval for Explainable Colorectal Polyp Diagnosis
von: Yang, Ruijie, et al.
Veröffentlicht: (2024)
von: Yang, Ruijie, et al.
Veröffentlicht: (2024)
A Flexible and Scalable Framework for Video Moment Search
von: Zhang, Chongzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Chongzhi, et al.
Veröffentlicht: (2025)
Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval
von: Molina, Adrià, et al.
Veröffentlicht: (2024)
von: Molina, Adrià, et al.
Veröffentlicht: (2024)
GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval
von: Sun, Chengsong, et al.
Veröffentlicht: (2025)
von: Sun, Chengsong, et al.
Veröffentlicht: (2025)
SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model
von: Lin, Lin, et al.
Veröffentlicht: (2025)
von: Lin, Lin, et al.
Veröffentlicht: (2025)
SignRAG: A Retrieval-Augmented System for Scalable Zero-Shot Road Sign Recognition
von: Zhu, Minghao, et al.
Veröffentlicht: (2025)
von: Zhu, Minghao, et al.
Veröffentlicht: (2025)
E5-V: Universal Embeddings with Multimodal Large Language Models
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
Diff-VPS: Video Polyp Segmentation via a Multi-task Diffusion Network with Adversarial Temporal Reasoning
von: Lu, Yingling, et al.
Veröffentlicht: (2024)
von: Lu, Yingling, et al.
Veröffentlicht: (2024)
Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning
von: Kühn, Paul Julius, et al.
Veröffentlicht: (2026)
von: Kühn, Paul Julius, et al.
Veröffentlicht: (2026)
MealRec: Multi-granularity Sequential Modeling via Hierarchical Diffusion Models for Micro-Video Recommendation
von: Dong, Xinxin, et al.
Veröffentlicht: (2026)
von: Dong, Xinxin, et al.
Veröffentlicht: (2026)
NextAds: Towards Next-generation Personalized Video Advertising
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
Are They the Same Picture? Adapting Concept Bottleneck Models for Human-AI Collaboration in Image Retrieval
von: Balloli, Vaibhav, et al.
Veröffentlicht: (2024)
von: Balloli, Vaibhav, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
Large Language Model Informed Patent Image Retrieval
von: Lo, Hao-Cheng, et al.
Veröffentlicht: (2024)
von: Lo, Hao-Cheng, et al.
Veröffentlicht: (2024)
Knowledge Discovery in Optical Music Recognition: Enhancing Information Retrieval with Instance Segmentation
von: Shatri, Elona, et al.
Veröffentlicht: (2024)
von: Shatri, Elona, et al.
Veröffentlicht: (2024)
Automating Iconclass: LLMs and RAG for Large-Scale Classification of Religious Woodcuts
von: Thomas, Drew B.
Veröffentlicht: (2025)
von: Thomas, Drew B.
Veröffentlicht: (2025)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search
von: Tang, Hengzhu, et al.
Veröffentlicht: (2025) -
Music Recommendation Based on Facial Emotion Recognition
von: B, Rajesh, et al.
Veröffentlicht: (2024) -
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
von: Meng, GuangHao, et al.
Veröffentlicht: (2025) -
I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking
von: Liu, Ziyan, et al.
Veröffentlicht: (2025) -
Rethinking Detection Based Table Structure Recognition for Visually Rich Document Images
von: Xiao, Bin, et al.
Veröffentlicht: (2023)