MMORE: Massive Multimodal Open RAG & Extraction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sallinen, Alexandre, Krsteski, Stefan, Teiletche, Paul, Allard, Marc-Antoine, Lecoeur, Baptiste, Zhang, Michael, Nemo, Fabrice, Kalajdzic, David, Meyer, Matthias, Hartley, Mary-Anne
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916950073933824
author Sallinen, Alexandre
Krsteski, Stefan
Teiletche, Paul
Allard, Marc-Antoine
Lecoeur, Baptiste
Zhang, Michael
Nemo, Fabrice
Kalajdzic, David
Meyer, Matthias
Hartley, Mary-Anne
author_facet Sallinen, Alexandre
Krsteski, Stefan
Teiletche, Paul
Allard, Marc-Antoine
Lecoeur, Baptiste
Zhang, Michael
Nemo, Fabrice
Kalajdzic, David
Meyer, Matthias
Hartley, Mary-Anne
contents We introduce MMORE, an open-source pipeline for Massive Multimodal Open RetrievalAugmented Generation and Extraction, designed to ingest, transform, and retrieve knowledge from heterogeneous document formats at scale. MMORE supports more than fifteen file types, including text, tables, images, emails, audio, and video, and processes them into a unified format to enable downstream applications for LLMs. The architecture offers modular, distributed processing, enabling scalable parallelization across CPUs and GPUs. On processing benchmarks, MMORE demonstrates a 3.8-fold speedup over single-node baselines and 40% higher accuracy than Docling on scanned PDFs. The pipeline integrates hybrid dense-sparse retrieval and supports both interactive APIs and batch RAG endpoints. Evaluated on PubMedQA, MMORE-augmented medical LLMs improve biomedical QA accuracy with increasing retrieval depth. MMORE provides a robust, extensible foundation for deploying task-agnostic RAG systems on diverse, real-world multimodal data. The codebase is available at https://github.com/swiss-ai/mmore.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11937
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMORE: Massive Multimodal Open RAG & Extraction
Sallinen, Alexandre
Krsteski, Stefan
Teiletche, Paul
Allard, Marc-Antoine
Lecoeur, Baptiste
Zhang, Michael
Nemo, Fabrice
Kalajdzic, David
Meyer, Matthias
Hartley, Mary-Anne
Software Engineering
Artificial Intelligence
D.2.0; E.m
We introduce MMORE, an open-source pipeline for Massive Multimodal Open RetrievalAugmented Generation and Extraction, designed to ingest, transform, and retrieve knowledge from heterogeneous document formats at scale. MMORE supports more than fifteen file types, including text, tables, images, emails, audio, and video, and processes them into a unified format to enable downstream applications for LLMs. The architecture offers modular, distributed processing, enabling scalable parallelization across CPUs and GPUs. On processing benchmarks, MMORE demonstrates a 3.8-fold speedup over single-node baselines and 40% higher accuracy than Docling on scanned PDFs. The pipeline integrates hybrid dense-sparse retrieval and supports both interactive APIs and batch RAG endpoints. Evaluated on PubMedQA, MMORE-augmented medical LLMs improve biomedical QA accuracy with increasing retrieval depth. MMORE provides a robust, extensible foundation for deploying task-agnostic RAG systems on diverse, real-world multimodal data. The codebase is available at https://github.com/swiss-ai/mmore.
title MMORE: Massive Multimodal Open RAG & Extraction
topic Software Engineering
Artificial Intelligence
D.2.0; E.m
url https://arxiv.org/abs/2509.11937