Multimodal Neural Databases
Fuente:
arXiv
Guardado en:
| Autores principales: | Trappolini, Giovanni, Santilli, Andrea, Rodolà, Emanuele, Halevy, Alon, Silvestri, Fabrizio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Introduction of a tree-based technique for efficient and real-time label retrieval in the object tracking system
por: Benrazek, Ala-Eddine, et al.
Publicado: (2022)
por: Benrazek, Ala-Eddine, et al.
Publicado: (2022)
LazyVLM: Neuro-Symbolic Approach to Video Analytics
por: Jian, Xiangru, et al.
Publicado: (2025)
por: Jian, Xiangru, et al.
Publicado: (2025)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
por: Wu, Siwei, et al.
Publicado: (2024)
por: Wu, Siwei, et al.
Publicado: (2024)
CLOSP: A Unified Semantic Space for SAR, MSI, and Text in Remote Sensing
por: Cambrin, Daniele Rege, et al.
Publicado: (2025)
por: Cambrin, Daniele Rege, et al.
Publicado: (2025)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
por: Yang, Jheng-Hong, et al.
Publicado: (2024)
por: Yang, Jheng-Hong, et al.
Publicado: (2024)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
por: Shih, Yu-Fei, et al.
Publicado: (2025)
por: Shih, Yu-Fei, et al.
Publicado: (2025)
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
por: Fang, Xiang, et al.
Publicado: (2022)
por: Fang, Xiang, et al.
Publicado: (2022)
Joint-Dataset Learning and Cross-Consistent Regularization for Text-to-Motion Retrieval
por: Messina, Nicola, et al.
Publicado: (2024)
por: Messina, Nicola, et al.
Publicado: (2024)
Multimodal Inverse Attention Network with Intrinsic Discriminant Feature Exploitation for Fake News Detection
por: Zhang, Tianlin, et al.
Publicado: (2025)
por: Zhang, Tianlin, et al.
Publicado: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
por: Li, Yongqi, et al.
Publicado: (2024)
por: Li, Yongqi, et al.
Publicado: (2024)
DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
por: Yang, Mengzheng, et al.
Publicado: (2025)
por: Yang, Mengzheng, et al.
Publicado: (2025)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
por: Li, Po-han, et al.
Publicado: (2024)
por: Li, Po-han, et al.
Publicado: (2024)
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
por: Lian, Niu, et al.
Publicado: (2026)
por: Lian, Niu, et al.
Publicado: (2026)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
por: Kong, Fanheng, et al.
Publicado: (2025)
por: Kong, Fanheng, et al.
Publicado: (2025)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
por: Liu, Han, et al.
Publicado: (2025)
por: Liu, Han, et al.
Publicado: (2025)
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
por: Fu, Junchen, et al.
Publicado: (2026)
por: Fu, Junchen, et al.
Publicado: (2026)
From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion
por: Wang, Zheng, et al.
Publicado: (2025)
por: Wang, Zheng, et al.
Publicado: (2025)
SODIUM: From Open Web Data to Queryable Databases
por: Hu, Chuxuan, et al.
Publicado: (2026)
por: Hu, Chuxuan, et al.
Publicado: (2026)
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
por: Zhang, Kai, et al.
Publicado: (2024)
por: Zhang, Kai, et al.
Publicado: (2024)
NativE: Multi-modal Knowledge Graph Completion in the Wild
por: Zhang, Yichi, et al.
Publicado: (2024)
por: Zhang, Yichi, et al.
Publicado: (2024)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
por: Gatti, Prajwal, et al.
Publicado: (2025)
por: Gatti, Prajwal, et al.
Publicado: (2025)
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
por: Xu, Zhengfei, et al.
Publicado: (2024)
por: Xu, Zhengfei, et al.
Publicado: (2024)
VKIE: The Application of Key Information Extraction on Video Text
por: An, Siyu, et al.
Publicado: (2023)
por: An, Siyu, et al.
Publicado: (2023)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
por: Ning, Hailong, et al.
Publicado: (2025)
por: Ning, Hailong, et al.
Publicado: (2025)
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
por: Li, Jun, et al.
Publicado: (2025)
por: Li, Jun, et al.
Publicado: (2025)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
por: Xiao, Jian, et al.
Publicado: (2025)
por: Xiao, Jian, et al.
Publicado: (2025)
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
por: Zhang, Pingping, et al.
Publicado: (2024)
por: Zhang, Pingping, et al.
Publicado: (2024)
Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
por: Deng, Jiaqi, et al.
Publicado: (2025)
por: Deng, Jiaqi, et al.
Publicado: (2025)
VisTopics: A Visual Semantic Unsupervised Approach to Topic Modeling of Video and Image Data
por: Lokmanoglu, Ayse D, et al.
Publicado: (2025)
por: Lokmanoglu, Ayse D, et al.
Publicado: (2025)
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
por: Messina, Nicola, et al.
Publicado: (2024)
por: Messina, Nicola, et al.
Publicado: (2024)
Efficient Self-Supervised Video Hashing with Selective State Spaces
por: Wang, Jinpeng, et al.
Publicado: (2024)
por: Wang, Jinpeng, et al.
Publicado: (2024)
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
por: Li, Jun, et al.
Publicado: (2026)
por: Li, Jun, et al.
Publicado: (2026)
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
por: Luo, Tianci, et al.
Publicado: (2026)
por: Luo, Tianci, et al.
Publicado: (2026)
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
por: Deng, Jiaqi, et al.
Publicado: (2025)
por: Deng, Jiaqi, et al.
Publicado: (2025)
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
por: Hu, Fan, et al.
Publicado: (2025)
por: Hu, Fan, et al.
Publicado: (2025)
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
por: Li, Fanxiao, et al.
Publicado: (2025)
por: Li, Fanxiao, et al.
Publicado: (2025)
Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
por: Chen, Junyu, et al.
Publicado: (2025)
por: Chen, Junyu, et al.
Publicado: (2025)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
por: Li, Minghan, et al.
Publicado: (2026)
por: Li, Minghan, et al.
Publicado: (2026)
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
por: Lian, Niu, et al.
Publicado: (2025)
por: Lian, Niu, et al.
Publicado: (2025)
UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval
por: Jiang, Haoyu, et al.
Publicado: (2024)
por: Jiang, Haoyu, et al.
Publicado: (2024)
Ejemplares similares
-
Introduction of a tree-based technique for efficient and real-time label retrieval in the object tracking system
por: Benrazek, Ala-Eddine, et al.
Publicado: (2022) -
LazyVLM: Neuro-Symbolic Approach to Video Analytics
por: Jian, Xiangru, et al.
Publicado: (2025) -
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
por: Wu, Siwei, et al.
Publicado: (2024) -
CLOSP: A Unified Semantic Space for SAR, MSI, and Text in Remote Sensing
por: Cambrin, Daniele Rege, et al.
Publicado: (2025) -
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
por: Yang, Jheng-Hong, et al.
Publicado: (2024)