KAP: MLLM-assisted OCR Text Enhancement for Hybrid Retrieval in Chinese Non-Narrative Documents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hsu, Hsin-Ling, Lin, Ping-Sheng, Lin, Jing-Di, Tzeng, Jengnan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DAT: Dynamic Alpha Tuning for Hybrid Retrieval in Retrieval-Augmented Generation
von: Hsu, Hsin-Ling, et al.
Veröffentlicht: (2025)
von: Hsu, Hsin-Ling, et al.
Veröffentlicht: (2025)
Unifying Multimodal Retrieval via Document Screenshot Embedding
von: Ma, Xueguang, et al.
Veröffentlicht: (2024)
von: Ma, Xueguang, et al.
Veröffentlicht: (2024)
SpectraQuery: A Hybrid Retrieval-Augmented Conversational Assistant for Battery Science
von: Vangara, Sreya, et al.
Veröffentlicht: (2026)
von: Vangara, Sreya, et al.
Veröffentlicht: (2026)
AnnoRetrieve: Efficient Structured Retrieval for Unstructured Document Analysis
von: Lin, Teng, et al.
Veröffentlicht: (2026)
von: Lin, Teng, et al.
Veröffentlicht: (2026)
MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
von: Li, Tianyuan, et al.
Veröffentlicht: (2025)
von: Li, Tianyuan, et al.
Veröffentlicht: (2025)
TC-MGC: Text-Conditioned Multi-Grained Contrastive Learning for Text-Video Retrieval
von: Jing, Xiaolun, et al.
Veröffentlicht: (2025)
von: Jing, Xiaolun, et al.
Veröffentlicht: (2025)
Document Screenshot Retrievers are Vulnerable to Pixel Poisoning Attacks
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2025)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2025)
Text-Video Retrieval With Global-Local Contrastive Consistency Learning
von: Jing, Xiaolun, et al.
Veröffentlicht: (2026)
von: Jing, Xiaolun, et al.
Veröffentlicht: (2026)
Ranking Narrative Query Graphs for Biomedical Document Retrieval (Technical Report)
von: Kroll, Hermann, et al.
Veröffentlicht: (2024)
von: Kroll, Hermann, et al.
Veröffentlicht: (2024)
Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models
von: Jin, Yichao, et al.
Veröffentlicht: (2025)
von: Jin, Yichao, et al.
Veröffentlicht: (2025)
HyReC: Exploring Hybrid-based Retriever for Chinese
von: Wang, Zunran, et al.
Veröffentlicht: (2025)
von: Wang, Zunran, et al.
Veröffentlicht: (2025)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
von: Li, Zirui, et al.
Veröffentlicht: (2025)
von: Li, Zirui, et al.
Veröffentlicht: (2025)
Sparse Meets Dense: A Hybrid Approach to Enhance Scientific Document Retrieval
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2024)
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2024)
TCDE: Topic-Centric Dual Expansion of Queries and Documents with Large Language Models for Information Retrieval
von: Yang, Yu, et al.
Veröffentlicht: (2025)
von: Yang, Yu, et al.
Veröffentlicht: (2025)
Towards Cross-Modal Text-Molecule Retrieval with Better Modality Alignment
von: Song, Jia, et al.
Veröffentlicht: (2024)
von: Song, Jia, et al.
Veröffentlicht: (2024)
Aligning Language Models for Versatile Text-based Item Retrieval
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
Causality Enhancement for Cross-Domain Recommendation
von: Wu, Zhibo, et al.
Veröffentlicht: (2025)
von: Wu, Zhibo, et al.
Veröffentlicht: (2025)
Retrieval Feedback Memory Enhancement Large Model Retrieval Generation Method
von: Li, Leqian, et al.
Veröffentlicht: (2025)
von: Li, Leqian, et al.
Veröffentlicht: (2025)
Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval
von: Molina, Adrià, et al.
Veröffentlicht: (2024)
von: Molina, Adrià, et al.
Veröffentlicht: (2024)
Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
A Survey of Generative Information Retrieval
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
Operational Advice for Dense and Sparse Retrievers: HNSW, Flat, or Inverted Indexes?
von: Lin, Jimmy
Veröffentlicht: (2024)
von: Lin, Jimmy
Veröffentlicht: (2024)
PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
von: Tong, Anyang, et al.
Veröffentlicht: (2025)
von: Tong, Anyang, et al.
Veröffentlicht: (2025)
FABULA: Intelligence Report Generation Using Retrieval-Augmented Narrative Construction
von: Ranade, Priyanka, et al.
Veröffentlicht: (2023)
von: Ranade, Priyanka, et al.
Veröffentlicht: (2023)
Synergistic Approach for Simultaneous Optimization of Monolingual, Cross-lingual, and Multilingual Information Retrieval
von: Elmahdy, Adel, et al.
Veröffentlicht: (2024)
von: Elmahdy, Adel, et al.
Veröffentlicht: (2024)
UniRank: End-to-End Domain-Specific Reranking of Hybrid Text-Image Candidates
von: Yang, Yupei, et al.
Veröffentlicht: (2026)
von: Yang, Yupei, et al.
Veröffentlicht: (2026)
Digitization of Document and Information Extraction using OCR
von: Sinha, Rasha, et al.
Veröffentlicht: (2025)
von: Sinha, Rasha, et al.
Veröffentlicht: (2025)
Enhancing Technical Documents Retrieval for RAG
von: Lai, Songjiang, et al.
Veröffentlicht: (2025)
von: Lai, Songjiang, et al.
Veröffentlicht: (2025)
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
von: Chen, Jianlyu, et al.
Veröffentlicht: (2025)
von: Chen, Jianlyu, et al.
Veröffentlicht: (2025)
CAT-ID$^2$: Category-Tree Integrated Document Identifier Learning for Generative Retrieval In E-commerce
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2025)
STARD: A Chinese Statute Retrieval Dataset with Real Queries Issued by Non-professionals
von: Su, Weihang, et al.
Veröffentlicht: (2024)
von: Su, Weihang, et al.
Veröffentlicht: (2024)
Reproducibility Analysis and Enhancements for Multi-Aspect Dense Retriever with Aspect Learning
von: Bi, Keping, et al.
Veröffentlicht: (2024)
von: Bi, Keping, et al.
Veröffentlicht: (2024)
LSTM-based Selective Dense Text Retrieval Guided by Sparse Lexical Retrieval
von: Yang, Yingrui, et al.
Veröffentlicht: (2025)
von: Yang, Yingrui, et al.
Veröffentlicht: (2025)
Retrieval Augmented Zero-Shot Text Classification
von: Abdullahi, Tassallah, et al.
Veröffentlicht: (2024)
von: Abdullahi, Tassallah, et al.
Veröffentlicht: (2024)
Benchmarking and Enabling Efficient Chinese Medical Retrieval via Asymmetric Encoders
von: Jiang, Angqing, et al.
Veröffentlicht: (2026)
von: Jiang, Angqing, et al.
Veröffentlicht: (2026)
A Hybrid Retrieval and Reranking Framework for Evidence-Grounded Retrieval-Augmented Generation
von: Irany, Fariba Afrin, et al.
Veröffentlicht: (2026)
von: Irany, Fariba Afrin, et al.
Veröffentlicht: (2026)
Efficient and Effective Retrieval of Dense-Sparse Hybrid Vectors using Graph-based Approximate Nearest Neighbor Search
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
von: Hsu, Tz-Huan, et al.
Veröffentlicht: (2026)
von: Hsu, Tz-Huan, et al.
Veröffentlicht: (2026)
DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
von: Yin, Kai, et al.
Veröffentlicht: (2025)
von: Yin, Kai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DAT: Dynamic Alpha Tuning for Hybrid Retrieval in Retrieval-Augmented Generation
von: Hsu, Hsin-Ling, et al.
Veröffentlicht: (2025) -
Unifying Multimodal Retrieval via Document Screenshot Embedding
von: Ma, Xueguang, et al.
Veröffentlicht: (2024) -
SpectraQuery: A Hybrid Retrieval-Augmented Conversational Assistant for Battery Science
von: Vangara, Sreya, et al.
Veröffentlicht: (2026) -
AnnoRetrieve: Efficient Structured Retrieval for Unstructured Document Analysis
von: Lin, Teng, et al.
Veröffentlicht: (2026) -
MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
von: Li, Tianyuan, et al.
Veröffentlicht: (2025)