Guardado en:
| Autores principales: | Turnbull, Robert, Fitzgerald, Emily, Thompson, Karen, Birch, Joanne L. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2410.08740 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Benchmark Granularity and Model Robustness for Image-Text Retrieval
por: Hendriksen, Mariya, et al.
Publicado: (2024)
por: Hendriksen, Mariya, et al.
Publicado: (2024)
Smart Routing for Multimodal Video Retrieval: When to Search What
por: Rosa, Kevin Dela
Publicado: (2025)
por: Rosa, Kevin Dela
Publicado: (2025)
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
por: Lim, Ho Hung, et al.
Publicado: (2026)
por: Lim, Ho Hung, et al.
Publicado: (2026)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
por: Cui, Cheng, et al.
Publicado: (2026)
por: Cui, Cheng, et al.
Publicado: (2026)
Attribute-Aware Implicit Modality Alignment for Text Attribute Person Search
por: Wang, Xin, et al.
Publicado: (2024)
por: Wang, Xin, et al.
Publicado: (2024)
Weakly Supervised Deep Hyperspherical Quantization for Image Retrieval
por: Wang, Jinpeng, et al.
Publicado: (2024)
por: Wang, Jinpeng, et al.
Publicado: (2024)
Low-Data Classification of Historical Music Manuscripts: A Few-Shot Learning Approach
por: Shatri, Elona, et al.
Publicado: (2024)
por: Shatri, Elona, et al.
Publicado: (2024)
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
por: Horn, Pius, et al.
Publicado: (2025)
por: Horn, Pius, et al.
Publicado: (2025)
From Historical Tabular Image to Knowledge Graphs: A Provenance-Aware Modular Pipeline
por: Shoilee, Sarah Binta Alam, et al.
Publicado: (2026)
por: Shoilee, Sarah Binta Alam, et al.
Publicado: (2026)
Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search
por: Chen, Lei, et al.
Publicado: (2026)
por: Chen, Lei, et al.
Publicado: (2026)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
por: Rosa, Kevin Dela
Publicado: (2024)
por: Rosa, Kevin Dela
Publicado: (2024)
Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
por: Luo, Enming, et al.
Publicado: (2024)
por: Luo, Enming, et al.
Publicado: (2024)
Accelerating Flood Warnings by 10 Hours: The Power of River Network Topology in AI-enhanced Flood Forecasting
por: Wang, Hongjun, et al.
Publicado: (2024)
por: Wang, Hongjun, et al.
Publicado: (2024)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
por: Duan, Yicheng, et al.
Publicado: (2025)
por: Duan, Yicheng, et al.
Publicado: (2025)
Content-based 3D Image Retrieval and a ColBERT-inspired Re-ranking for Tumor Flagging and Staging
por: Jush, Farnaz Khun, et al.
Publicado: (2025)
por: Jush, Farnaz Khun, et al.
Publicado: (2025)
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
por: Long, Zijun, et al.
Publicado: (2025)
por: Long, Zijun, et al.
Publicado: (2025)
Your Embedding Model is SMARTer Than You Think
por: Zhang, Jianrui, et al.
Publicado: (2026)
por: Zhang, Jianrui, et al.
Publicado: (2026)
PHPQ: Pyramid Hybrid Pooling Quantization for Efficient Fine-Grained Image Retrieval
por: Zeng, Ziyun, et al.
Publicado: (2021)
por: Zeng, Ziyun, et al.
Publicado: (2021)
Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems
por: Zhang, Tuo, et al.
Publicado: (2025)
por: Zhang, Tuo, et al.
Publicado: (2025)
Efficient Logic Gate Networks for Video Copy Detection
por: Fojcik, Katarzyna
Publicado: (2026)
por: Fojcik, Katarzyna
Publicado: (2026)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
por: Yin, Xinlei, et al.
Publicado: (2026)
por: Yin, Xinlei, et al.
Publicado: (2026)
Good Scores, Bad Data: A Metric for Multimodal Coherence
por: Srinivasan, Vasundra
Publicado: (2026)
por: Srinivasan, Vasundra
Publicado: (2026)
RAVEN: Multitask Retrieval Augmented Vision-Language Learning
por: Rao, Varun Nagaraj, et al.
Publicado: (2024)
por: Rao, Varun Nagaraj, et al.
Publicado: (2024)
Are They the Same Picture? Adapting Concept Bottleneck Models for Human-AI Collaboration in Image Retrieval
por: Balloli, Vaibhav, et al.
Publicado: (2024)
por: Balloli, Vaibhav, et al.
Publicado: (2024)
Self-supervised Graph Neural Network for Mechanical CAD Retrieval
por: Quan, Yuhan, et al.
Publicado: (2024)
por: Quan, Yuhan, et al.
Publicado: (2024)
Content-Based Image Retrieval for Multi-Class Volumetric Radiology Images: A Benchmark Study
por: Jush, Farnaz Khun, et al.
Publicado: (2024)
por: Jush, Farnaz Khun, et al.
Publicado: (2024)
The future of document indexing: GPT and Donut revolutionize table of content processing
por: Feyisa, Degaga Wolde, et al.
Publicado: (2024)
por: Feyisa, Degaga Wolde, et al.
Publicado: (2024)
NovaLAD: A Fast, CPU-Optimized Document Extraction Pipeline for Generative AI and Data Intelligence
por: Ulla, Aman
Publicado: (2026)
por: Ulla, Aman
Publicado: (2026)
From Drawings to Decisions: A Hybrid Vision-Language Framework for Parsing 2D Engineering Drawings into Structured Manufacturing Knowledge
por: Khan, Muhammad Tayyab, et al.
Publicado: (2025)
por: Khan, Muhammad Tayyab, et al.
Publicado: (2025)
RAPTOR: Refined Approach for Product Table Object Recognition
por: Thomas, Eliott, et al.
Publicado: (2025)
por: Thomas, Eliott, et al.
Publicado: (2025)
VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings
por: Giahi, Ramin, et al.
Publicado: (2025)
por: Giahi, Ramin, et al.
Publicado: (2025)
CBM-RAG: Demonstrating Enhanced Interpretability in Radiology Report Generation with Multi-Agent RAG and Concept Bottleneck Models
por: Alam, Hasan Md Tusfiqur, et al.
Publicado: (2025)
por: Alam, Hasan Md Tusfiqur, et al.
Publicado: (2025)
DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval
por: Yang, Yuxin, et al.
Publicado: (2025)
por: Yang, Yuxin, et al.
Publicado: (2025)
AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency
por: Patel, Piyushkumar
Publicado: (2025)
por: Patel, Piyushkumar
Publicado: (2025)
HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
por: Luo, Linyin, et al.
Publicado: (2025)
por: Luo, Linyin, et al.
Publicado: (2025)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
por: Ren, Xubin, et al.
Publicado: (2025)
por: Ren, Xubin, et al.
Publicado: (2025)
Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
por: Mezzi, Emanuele, et al.
Publicado: (2025)
por: Mezzi, Emanuele, et al.
Publicado: (2025)
Sell It Before You Make It: Revolutionizing E-Commerce with Personalized AI-Generated Items
por: Lin, Jianghao, et al.
Publicado: (2025)
por: Lin, Jianghao, et al.
Publicado: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
por: Zhou, Yinan, et al.
Publicado: (2025)
por: Zhou, Yinan, et al.
Publicado: (2025)
FiCo-ITR: bridging fine-grained and coarse-grained image-text retrieval for comparative performance analysis
por: Williams-Lekuona, Mikel, et al.
Publicado: (2024)
por: Williams-Lekuona, Mikel, et al.
Publicado: (2024)
Ejemplares similares
-
Benchmark Granularity and Model Robustness for Image-Text Retrieval
por: Hendriksen, Mariya, et al.
Publicado: (2024) -
Smart Routing for Multimodal Video Retrieval: When to Search What
por: Rosa, Kevin Dela
Publicado: (2025) -
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
por: Lim, Ho Hung, et al.
Publicado: (2026) -
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
por: Cui, Cheng, et al.
Publicado: (2026) -
Attribute-Aware Implicit Modality Alignment for Text Attribute Person Search
por: Wang, Xin, et al.
Publicado: (2024)