Evaluation of Embedding-Based and Generative Methods for LLM-Driven Document Classification: Opportunities and Challenges
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Rong, Liu, Hao, Hou, Song |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
von: Zhang, Longxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Longxiang, et al.
Veröffentlicht: (2026)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024)
von: Yu, Shi, et al.
Veröffentlicht: (2024)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
Smart Multi-Modal Search: Contextual Sparse and Dense Embedding Integration in Adobe Express
von: Aroraa, Cherag, et al.
Veröffentlicht: (2024)
von: Aroraa, Cherag, et al.
Veröffentlicht: (2024)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
von: Guo, Hao, et al.
Veröffentlicht: (2025)
von: Guo, Hao, et al.
Veröffentlicht: (2025)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
E5-V: Universal Embeddings with Multimodal Large Language Models
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
von: Shen, Wenxuan, et al.
Veröffentlicht: (2025)
von: Shen, Wenxuan, et al.
Veröffentlicht: (2025)
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
Capability-aware Prompt Reformulation Learning for Text-to-Image Generation
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
von: Zhan, Jingtao, et al.
Veröffentlicht: (2024)
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
von: Zhu, Jing, et al.
Veröffentlicht: (2025)
von: Zhu, Jing, et al.
Veröffentlicht: (2025)
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
von: Yang, Mengzheng, et al.
Veröffentlicht: (2025)
von: Yang, Mengzheng, et al.
Veröffentlicht: (2025)
ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding
von: Rizk, Basem, et al.
Veröffentlicht: (2025)
von: Rizk, Basem, et al.
Veröffentlicht: (2025)
Self Knowledge Re-expression: A Fully Local Method for Adapting LLMs to Tasks Using Intrinsic Knowledge
von: Wang, Mengyu, et al.
Veröffentlicht: (2026)
von: Wang, Mengyu, et al.
Veröffentlicht: (2026)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
ITEm: Unsupervised Image-Text Embedding Learning for eCommerce
von: Liao, Baohao, et al.
Veröffentlicht: (2023)
von: Liao, Baohao, et al.
Veröffentlicht: (2023)
Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhuchenyang, et al.
Veröffentlicht: (2026)
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
von: Shohan, Faisal Tareque, et al.
Veröffentlicht: (2024)
von: Shohan, Faisal Tareque, et al.
Veröffentlicht: (2024)
ReinPool: Reinforcement Learning Pooling Multi-Vector Embeddings for Retrieval System
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
Attention Grounded Enhancement for Visual Document Retrieval
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
PathoScribe: Transforming Pathology Data into a Living Library with a Unified LLM-Driven Framework for Semantic Retrieval and Clinical Integration
von: Akbar, Abdul Rehman, et al.
Veröffentlicht: (2026)
von: Akbar, Abdul Rehman, et al.
Veröffentlicht: (2026)
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
von: Guo, Minghao, et al.
Veröffentlicht: (2026)
von: Guo, Minghao, et al.
Veröffentlicht: (2026)
TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations
von: Si, Jacob, et al.
Veröffentlicht: (2025)
von: Si, Jacob, et al.
Veröffentlicht: (2025)
ColPali: Efficient Document Retrieval with Vision Language Models
von: Faysse, Manuel, et al.
Veröffentlicht: (2024)
von: Faysse, Manuel, et al.
Veröffentlicht: (2024)
Improving Applicability of Deep Learning based Token Classification models during Training
von: Mehra, Anket, et al.
Veröffentlicht: (2025)
von: Mehra, Anket, et al.
Veröffentlicht: (2025)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
von: Zhang, Longxiang, et al.
Veröffentlicht: (2026) -
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024) -
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025) -
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025) -
Smart Multi-Modal Search: Contextual Sparse and Dense Embedding Integration in Adobe Express
von: Aroraa, Cherag, et al.
Veröffentlicht: (2024)