Gespeichert in:
| Hauptverfasser: | Xiao, Bin, Simsek, Murat, Kantarci, Burak, Alkheir, Ala Abu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2312.00699 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Multi-Granularity Retrieval Framework for Visually-Rich Documents
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy
von: Kim, Juyeon, et al.
Veröffentlicht: (2025)
von: Kim, Juyeon, et al.
Veröffentlicht: (2025)
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
von: Zhu, Yingjian, et al.
Veröffentlicht: (2026)
von: Zhu, Yingjian, et al.
Veröffentlicht: (2026)
One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image
von: Shereen, Ezzeldin, et al.
Veröffentlicht: (2025)
von: Shereen, Ezzeldin, et al.
Veröffentlicht: (2025)
Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
von: Luo, Tianci, et al.
Veröffentlicht: (2026)
von: Luo, Tianci, et al.
Veröffentlicht: (2026)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
Music Recommendation Based on Facial Emotion Recognition
von: B, Rajesh, et al.
Veröffentlicht: (2024)
von: B, Rajesh, et al.
Veröffentlicht: (2024)
Rethinking Sparse Lexical Representations for Image Retrieval in the Age of Rising Multi-Modal Large Language Models
von: Nakata, Kengo, et al.
Veröffentlicht: (2024)
von: Nakata, Kengo, et al.
Veröffentlicht: (2024)
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
Leveraging Foundation Models for Content-Based Image Retrieval in Radiology
von: Denner, Stefan, et al.
Veröffentlicht: (2024)
von: Denner, Stefan, et al.
Veröffentlicht: (2024)
A Collaborative Jade Recognition System for Mobile Devices Based on Lightweight and Large Models
von: Wang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2025)
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
von: Wang, Junyi, et al.
Veröffentlicht: (2025)
von: Wang, Junyi, et al.
Veröffentlicht: (2025)
ProGEO: Generating Prompts through Image-Text Contrastive Learning for Visual Geo-localization
von: Mao, Chen, et al.
Veröffentlicht: (2024)
von: Mao, Chen, et al.
Veröffentlicht: (2024)
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
DEMO: A Statistical Perspective for Efficient Image-Text Matching
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
von: Guo, Hao, et al.
Veröffentlicht: (2025)
von: Guo, Hao, et al.
Veröffentlicht: (2025)
Attention Grounded Enhancement for Visual Document Retrieval
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
Offline Evaluation of Set-Based Text-to-Image Generation
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification
von: Keita, Mamadou, et al.
Veröffentlicht: (2025)
von: Keita, Mamadou, et al.
Veröffentlicht: (2025)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
Interactive Mars Image Content-Based Search with Interpretable Machine Learning
von: Vasu, Bhavan, et al.
Veröffentlicht: (2024)
von: Vasu, Bhavan, et al.
Veröffentlicht: (2024)
Index Light, Reason Deep: Deferred Visual Ingestion for Visual-Dense Document Question Answering
von: Xu, Tao
Veröffentlicht: (2026)
von: Xu, Tao
Veröffentlicht: (2026)
RAPTOR: Refined Approach for Product Table Object Recognition
von: Thomas, Eliott, et al.
Veröffentlicht: (2025)
von: Thomas, Eliott, et al.
Veröffentlicht: (2025)
Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
von: Zhang, Yao, et al.
Veröffentlicht: (2026)
von: Zhang, Yao, et al.
Veröffentlicht: (2026)
FF-PNet: A Pyramid Network Based on Feature and Field for Brain Image Registration
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
Visual Product Search Benchmark
von: Govindappa, Karthik Sulthanpete
Veröffentlicht: (2026)
von: Govindappa, Karthik Sulthanpete
Veröffentlicht: (2026)
Revisit Anything: Visual Place Recognition via Image Segment Retrieval
von: Garg, Kartik, et al.
Veröffentlicht: (2024)
von: Garg, Kartik, et al.
Veröffentlicht: (2024)
MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2025)
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2025)
ExTTNet: A Deep Learning Algorithm for Extracting Table Texts from Invoice Images
von: Akdoğan, Adem, et al.
Veröffentlicht: (2024)
von: Akdoğan, Adem, et al.
Veröffentlicht: (2024)
Semi-Supervised Image-Based Narrative Extraction: A Case Study with Historical Photographic Records
von: German, Fausto, et al.
Veröffentlicht: (2025)
von: German, Fausto, et al.
Veröffentlicht: (2025)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
Scalable Residual Feature Aggregation Framework with Hybrid Metaheuristic Optimization for Robust Early Pancreatic Neoplasm Detection in Multimodal CT Imaging
von: Thiruvengadam, Janani Annur, et al.
Veröffentlicht: (2025)
von: Thiruvengadam, Janani Annur, et al.
Veröffentlicht: (2025)
VisTopics: A Visual Semantic Unsupervised Approach to Topic Modeling of Video and Image Data
von: Lokmanoglu, Ayse D, et al.
Veröffentlicht: (2025)
von: Lokmanoglu, Ayse D, et al.
Veröffentlicht: (2025)
Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning
von: Kühn, Paul Julius, et al.
Veröffentlicht: (2026)
von: Kühn, Paul Julius, et al.
Veröffentlicht: (2026)
Digitization of Document and Information Extraction using OCR
von: Sinha, Rasha, et al.
Veröffentlicht: (2025)
von: Sinha, Rasha, et al.
Veröffentlicht: (2025)
Entity Image and Mixed-Modal Image Retrieval Datasets
von: Blaga, Cristian-Ioan, et al.
Veröffentlicht: (2025)
von: Blaga, Cristian-Ioan, et al.
Veröffentlicht: (2025)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
von: Ning, Hailong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Multi-Granularity Retrieval Framework for Visually-Rich Documents
von: Xu, Mingjun, et al.
Veröffentlicht: (2025) -
Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy
von: Kim, Juyeon, et al.
Veröffentlicht: (2025) -
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
von: Zhu, Yingjian, et al.
Veröffentlicht: (2026) -
One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image
von: Shereen, Ezzeldin, et al.
Veröffentlicht: (2025) -
Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)