The Science Data Lake: A Unified Open Infrastructure Integrating 293 Million Papers Across Eight Scholarly Sources with Embedding-Based Ontology Alignment
Fuente:
arXiv
Saved in:
| Main Author: | Wilinski, Jonas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DatAasee -- A Metadata-Lake as Metadata Catalog for a Virtual Data-Lake
by: Himpe, Christian
Published: (2024)
by: Himpe, Christian
Published: (2024)
STEP: Stepwise Curriculum Learning for Context-Knowledge Fusion in Conversational Recommendation
by: Yang, Zhenye, et al.
Published: (2025)
by: Yang, Zhenye, et al.
Published: (2025)
Reviewing the Reviewer: Graph-Enhanced LLMs for E-commerce Appeal Adjudication
by: Du, Yuchen, et al.
Published: (2026)
by: Du, Yuchen, et al.
Published: (2026)
DataFactory: Collaborative Multi-Agent Framework for Advanced Table Question Answering
by: Wang, Tong, et al.
Published: (2026)
by: Wang, Tong, et al.
Published: (2026)
Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"?
by: Eizaguirre, Germán T., et al.
Published: (2026)
by: Eizaguirre, Germán T., et al.
Published: (2026)
Free Access to World News: Reconstructing Full-Text Articles from GDELT
by: Colladon, A. Fronzetti, et al.
Published: (2025)
by: Colladon, A. Fronzetti, et al.
Published: (2025)
SotA Lens: A Network-Augmented Methodology and Tool for Exploratory State-of-the-Art Reviews
by: Cordeiro, Diogo Peralta
Published: (2026)
by: Cordeiro, Diogo Peralta
Published: (2026)
WisPaper: Your AI Scholar Search Engine
by: Ju, Li, et al.
Published: (2025)
by: Ju, Li, et al.
Published: (2025)
BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study
by: Mutsaddi, Atharva, et al.
Published: (2025)
by: Mutsaddi, Atharva, et al.
Published: (2025)
Session Context Embedding for Intent Understanding in Product Search
by: Mehrdad, Navid, et al.
Published: (2024)
by: Mehrdad, Navid, et al.
Published: (2024)
HySemRAG: A Hybrid Semantic Retrieval-Augmented Generation Framework for Automated Literature Synthesis and Methodological Gap Analysis
by: Godinez, Alejandro
Published: (2025)
by: Godinez, Alejandro
Published: (2025)
Doctoral Theses in France (1985-2025): A Linked Dataset of PhDs, Academic Networks, and Institutions
by: Aboucaya, William, et al.
Published: (2026)
by: Aboucaya, William, et al.
Published: (2026)
VulCPE: Context-Aware Cybersecurity Vulnerability Retrieval and Management
by: Jiang, Yuning, et al.
Published: (2025)
by: Jiang, Yuning, et al.
Published: (2025)
Falkor-IRAC: Graph-Constrained Generation for Verified Legal Reasoning in Indian Judicial AI
by: Bose, Joy
Published: (2026)
by: Bose, Joy
Published: (2026)
From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms
by: Kai, Zhang, et al.
Published: (2026)
by: Kai, Zhang, et al.
Published: (2026)
AI-Friendly LaTeX: Using LaTeX Code as a Knowledge Source for Retrieval-Augmented Generation
by: Verhoeff, Tom
Published: (2026)
by: Verhoeff, Tom
Published: (2026)
Topic Is Not Agenda: A Citation-Community Audit of Text Embeddings
by: Yoo, Junseon
Published: (2026)
by: Yoo, Junseon
Published: (2026)
VOGUE: A Multimodal Dataset for Conversational Recommendation in Fashion
by: Guo, David, et al.
Published: (2025)
by: Guo, David, et al.
Published: (2025)
SoccerRAG: Multimodal Soccer Information Retrieval via Natural Queries
by: Strand, Aleksander Theo, et al.
Published: (2024)
by: Strand, Aleksander Theo, et al.
Published: (2024)
Demo: Soccer Information Retrieval via Natural Queries using SoccerRAG
by: Strand, Aleksander Theo, et al.
Published: (2024)
by: Strand, Aleksander Theo, et al.
Published: (2024)
Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction
by: Litvak, Ivan Leonidovich, et al.
Published: (2025)
by: Litvak, Ivan Leonidovich, et al.
Published: (2025)
MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
BridgeRAG: Training-Free Bridge-Conditioned Retrieval for Multi-Hop Question Answering
by: Bacellar, Andre
Published: (2026)
by: Bacellar, Andre
Published: (2026)
Train Once, Use Flexibly: A Modular Framework for Multi-Aspect Neural News Recommendation
by: Iana, Andreea, et al.
Published: (2023)
by: Iana, Andreea, et al.
Published: (2023)
Scaling Multilingual Semantic Search in Uber Eats Delivery
by: Ling, Bo, et al.
Published: (2026)
by: Ling, Bo, et al.
Published: (2026)
Does UMBRELA Work on Other LLMs?
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Algorithmic Trust and Compliance: Benchmarking Brand Notability for UK iGaming Entities in Generative Search Engines
by: Oruesagasti, Julen
Published: (2026)
by: Oruesagasti, Julen
Published: (2026)
Criteria-Based LLM Relevance Judgments
by: Farzi, Naghmeh, et al.
Published: (2025)
by: Farzi, Naghmeh, et al.
Published: (2025)
Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs
by: Koutsiaris, Christos
Published: (2026)
by: Koutsiaris, Christos
Published: (2026)
Graph-GRPO: Dependency-Aware Credit Assignment for Generative E-commerce Search Relevance
by: Che, Jiarui, et al.
Published: (2026)
by: Che, Jiarui, et al.
Published: (2026)
Optimizing open-domain question answering with graph-based retrieval augmented generation
by: Cahoon, Joyce, et al.
Published: (2025)
by: Cahoon, Joyce, et al.
Published: (2025)
Architecture Matters More Than Scale: A Comparative Study of Retrieval and Memory Augmentation for Financial QA Under SME Compute Constraints
by: Liu, Jianan, et al.
Published: (2026)
by: Liu, Jianan, et al.
Published: (2026)
HiPS: Hierarchical PDF Segmentation of Textbooks
by: Wehnert, Sabine, et al.
Published: (2025)
by: Wehnert, Sabine, et al.
Published: (2025)
BatchBench: Toward a Workload-Aware Benchmark for Autoscaling Policies in Big Data Batch Processing -- A Proposed Framework
by: Budigi, Venkata Krishna Prasanth, et al.
Published: (2026)
by: Budigi, Venkata Krishna Prasanth, et al.
Published: (2026)
HiFi-RAG: Hierarchical Content Filtering and Two-Pass Generation for Open-Domain RAG
by: Nuengsigkapian, Cattalyya
Published: (2025)
by: Nuengsigkapian, Cattalyya
Published: (2025)
Un análisis bibliométrico de la producción científica acerca del agrupamiento de trayectorias GPS
by: Reyes, Gary, et al.
Published: (2024)
by: Reyes, Gary, et al.
Published: (2024)
Context-aware Privacy Bounds for Linear Queries
by: Zhao, Heng, et al.
Published: (2026)
by: Zhao, Heng, et al.
Published: (2026)
Diversification as Risk Minimization
by: Takehi, Rikiya, et al.
Published: (2025)
by: Takehi, Rikiya, et al.
Published: (2025)
AuthorityBench: Benchmarking LLM Authority Perception for Reliable Retrieval-Augmented Generation
by: Yao, Zhihui, et al.
Published: (2026)
by: Yao, Zhihui, et al.
Published: (2026)
MeVer at CheckThat! 2026: Cluster-Aware Hard-Negative Mining for Multilingual Scientific-Source Retrieval
by: Bakagianni, Juli, et al.
Published: (2026)
by: Bakagianni, Juli, et al.
Published: (2026)
Similar Items
-
DatAasee -- A Metadata-Lake as Metadata Catalog for a Virtual Data-Lake
by: Himpe, Christian
Published: (2024) -
STEP: Stepwise Curriculum Learning for Context-Knowledge Fusion in Conversational Recommendation
by: Yang, Zhenye, et al.
Published: (2025) -
Reviewing the Reviewer: Graph-Enhanced LLMs for E-commerce Appeal Adjudication
by: Du, Yuchen, et al.
Published: (2026) -
DataFactory: Collaborative Multi-Agent Framework for Advanced Table Question Answering
by: Wang, Tong, et al.
Published: (2026) -
Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"?
by: Eizaguirre, Germán T., et al.
Published: (2026)