PICS: Pipeline for Image Captioning and Search
Fuente:
arXiv
Salvato in:
| Autori principali: | Rosario, Grant, Noever, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
di: Xu, Yifan, et al.
Pubblicazione: (2024)
di: Xu, Yifan, et al.
Pubblicazione: (2024)
Siamese Content-based Search Engine for a More Transparent Skin and Breast Cancer Diagnosis through Histological Imaging
di: Tabatabaei, Zahra, et al.
Pubblicazione: (2024)
di: Tabatabaei, Zahra, et al.
Pubblicazione: (2024)
When & How to Write for Personalized Demand-aware Query Rewriting in Video Search
di: cheng, Cheng, et al.
Pubblicazione: (2025)
di: cheng, Cheng, et al.
Pubblicazione: (2025)
Electrooptical Image Synthesis from SAR Imagery Using Generative Adversarial Networks
di: Rosario, Grant, et al.
Pubblicazione: (2024)
di: Rosario, Grant, et al.
Pubblicazione: (2024)
LotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff Table
di: Matsui, Yusuke
Pubblicazione: (2025)
di: Matsui, Yusuke
Pubblicazione: (2025)
Evidential Transformers for Improved Image Retrieval
di: Dordevic, Danilo, et al.
Pubblicazione: (2024)
di: Dordevic, Danilo, et al.
Pubblicazione: (2024)
Image Fusion for Cross-Domain Sequential Recommendation
di: Wu, Wangyu, et al.
Pubblicazione: (2024)
di: Wu, Wangyu, et al.
Pubblicazione: (2024)
FOR: Finetuning for Object Level Open Vocabulary Image Retrieval
di: Levi, Hila, et al.
Pubblicazione: (2024)
di: Levi, Hila, et al.
Pubblicazione: (2024)
Image Outlier Detection Without Training using RANSAC
di: Tsai, Chen-Han, et al.
Pubblicazione: (2023)
di: Tsai, Chen-Han, et al.
Pubblicazione: (2023)
Sustainable transparency in Recommender Systems: Bayesian Ranking of Images for Explainability
di: Paz-Ruza, Jorge, et al.
Pubblicazione: (2023)
di: Paz-Ruza, Jorge, et al.
Pubblicazione: (2023)
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
di: Sogi, Naoya, et al.
Pubblicazione: (2024)
di: Sogi, Naoya, et al.
Pubblicazione: (2024)
Hierarchy-of-Visual-Words: a Learning-based Approach for Trademark Image Retrieval
di: Lourenço, Vítor N., et al.
Pubblicazione: (2019)
di: Lourenço, Vítor N., et al.
Pubblicazione: (2019)
Image Hashing via Cross-View Code Alignment in the Age of Foundation Models
di: Moummad, Ilyass, et al.
Pubblicazione: (2025)
di: Moummad, Ilyass, et al.
Pubblicazione: (2025)
Towards Resource-Efficient Streaming of Large-Scale Medical Image Datasets for Deep Learning
di: Kulkarni, Pranav, et al.
Pubblicazione: (2023)
di: Kulkarni, Pranav, et al.
Pubblicazione: (2023)
Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
di: Liu, Zheyuan, et al.
Pubblicazione: (2023)
di: Liu, Zheyuan, et al.
Pubblicazione: (2023)
Taxonomic Reasoning for Rare Arthropods: Combining Dense Image Captioning and RAG for Interpretable Classification
di: Lesperance, Nathaniel, et al.
Pubblicazione: (2025)
di: Lesperance, Nathaniel, et al.
Pubblicazione: (2025)
CoopHash: Cooperative Learning of Multipurpose Descriptor and Contrastive Pair Generator via Variational MCMC Teaching for Supervised Image Hashing
di: Doan, Khoa D., et al.
Pubblicazione: (2022)
di: Doan, Khoa D., et al.
Pubblicazione: (2022)
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding
di: Pal, Aniket, et al.
Pubblicazione: (2025)
di: Pal, Aniket, et al.
Pubblicazione: (2025)
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
di: Zhang, Zhixin, et al.
Pubblicazione: (2024)
di: Zhang, Zhixin, et al.
Pubblicazione: (2024)
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
di: Duan, Yue, et al.
Pubblicazione: (2024)
di: Duan, Yue, et al.
Pubblicazione: (2024)
CTR-Driven Advertising Image Generation with Multimodal Large Language Models
di: Chen, Xingye, et al.
Pubblicazione: (2025)
di: Chen, Xingye, et al.
Pubblicazione: (2025)
GENIUS: A Generative Framework for Universal Multimodal Search
di: Kim, Sungyeon, et al.
Pubblicazione: (2025)
di: Kim, Sungyeon, et al.
Pubblicazione: (2025)
Learning-Based Hashing for ANN Search: Foundations and Early Advances
di: Moran, Sean
Pubblicazione: (2025)
di: Moran, Sean
Pubblicazione: (2025)
MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising
di: Fu, Chenghan, et al.
Pubblicazione: (2025)
di: Fu, Chenghan, et al.
Pubblicazione: (2025)
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
di: Zhu, Tianyu, et al.
Pubblicazione: (2024)
di: Zhu, Tianyu, et al.
Pubblicazione: (2024)
Proceedings of the 6th International Workshop on Reading Music Systems
di: Calvo-Zaragoza, Jorge, et al.
Pubblicazione: (2024)
di: Calvo-Zaragoza, Jorge, et al.
Pubblicazione: (2024)
Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval
di: Sarkar, Rohan, et al.
Pubblicazione: (2024)
di: Sarkar, Rohan, et al.
Pubblicazione: (2024)
A Dataset and Framework for Learning State-invariant Object Representations
di: Sarkar, Rohan, et al.
Pubblicazione: (2024)
di: Sarkar, Rohan, et al.
Pubblicazione: (2024)
iRAG: Advancing RAG for Videos with an Incremental Approach
di: Arefeen, Md Adnan, et al.
Pubblicazione: (2024)
di: Arefeen, Md Adnan, et al.
Pubblicazione: (2024)
Multi-Label Plant Species Classification with Self-Supervised Vision Transformers
di: Gustineli, Murilo, et al.
Pubblicazione: (2024)
di: Gustineli, Murilo, et al.
Pubblicazione: (2024)
A Guide to Similarity Measures
di: Levy, Avivit, et al.
Pubblicazione: (2024)
di: Levy, Avivit, et al.
Pubblicazione: (2024)
Online Learning via Memory: Retrieval-Augmented Detector Adaptation
di: Jian, Yanan, et al.
Pubblicazione: (2024)
di: Jian, Yanan, et al.
Pubblicazione: (2024)
A Fashion Item Recommendation Model in Hyperbolic Space
di: Shimizu, Ryotaro, et al.
Pubblicazione: (2024)
di: Shimizu, Ryotaro, et al.
Pubblicazione: (2024)
InvDiff: Invariant Guidance for Bias Mitigation in Diffusion Models
di: Hou, Min, et al.
Pubblicazione: (2024)
di: Hou, Min, et al.
Pubblicazione: (2024)
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
di: Dave, Ishan Rajendrakumar, et al.
Pubblicazione: (2024)
They're All Doctors: Synthesizing Diverse Counterfactuals to Mitigate Associative Bias
di: Magid, Salma Abdel, et al.
Pubblicazione: (2024)
di: Magid, Salma Abdel, et al.
Pubblicazione: (2024)
Transfer Learning with Self-Supervised Vision Transformers for Snake Identification
di: Miyaguchi, Anthony, et al.
Pubblicazione: (2024)
di: Miyaguchi, Anthony, et al.
Pubblicazione: (2024)
VQPP: Video Query Performance Prediction Benchmark
di: Lutu, Adrian Catalin, et al.
Pubblicazione: (2026)
di: Lutu, Adrian Catalin, et al.
Pubblicazione: (2026)
SOLAR: SVD-Optimized Lifelong Attention for Recommendation
di: Zhang, Chenghao, et al.
Pubblicazione: (2026)
di: Zhang, Chenghao, et al.
Pubblicazione: (2026)
Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
di: Tam, Jason Kahei, et al.
Pubblicazione: (2025)
di: Tam, Jason Kahei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
di: Xu, Yifan, et al.
Pubblicazione: (2024) -
Siamese Content-based Search Engine for a More Transparent Skin and Breast Cancer Diagnosis through Histological Imaging
di: Tabatabaei, Zahra, et al.
Pubblicazione: (2024) -
When & How to Write for Personalized Demand-aware Query Rewriting in Video Search
di: cheng, Cheng, et al.
Pubblicazione: (2025) -
Electrooptical Image Synthesis from SAR Imagery Using Generative Adversarial Networks
di: Rosario, Grant, et al.
Pubblicazione: (2024) -
LotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff Table
di: Matsui, Yusuke
Pubblicazione: (2025)