Saved in:
Bibliographic Details
Main Authors: Patel, Devendra, Jain, Aaditya, Verma, Jayant, Rajput, Divyansh, Mahala, Sunil, Khapare, Ketki Suresh, Kalla, Jayateja
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.03329
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916825625788416
author Patel, Devendra
Jain, Aaditya
Verma, Jayant
Rajput, Divyansh
Mahala, Sunil
Khapare, Ketki Suresh
Kalla, Jayateja
author_facet Patel, Devendra
Jain, Aaditya
Verma, Jayant
Rajput, Divyansh
Mahala, Sunil
Khapare, Ketki Suresh
Kalla, Jayateja
contents We present NDAI-NeuroMAP, the first neuroscience-domain-specific dense vector embedding model engineered for high-precision information retrieval tasks. Our methodology encompasses the curation of an extensive domain-specific training corpus comprising 500,000 carefully constructed triplets (query-positive-negative configurations), augmented with 250,000 neuroscience-specific definitional entries and 250,000 structured knowledge-graph triplets derived from authoritative neurological ontologies. We employ a sophisticated fine-tuning approach utilizing the FremyCompany/BioLORD-2023 foundation model, implementing a multi-objective optimization framework combining contrastive learning with triplet-based metric learning paradigms. Comprehensive evaluation on a held-out test dataset comprising approximately 24,000 neuroscience-specific queries demonstrates substantial performance improvements over state-of-the-art general-purpose and biomedical embedding models. These empirical findings underscore the critical importance of domain-specific embedding architectures for neuroscience-oriented RAG systems and related clinical natural language processing applications.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03329
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NDAI-NeuroMAP: A Neuroscience-Specific Embedding Model for Domain-Specific Retrieval
Patel, Devendra
Jain, Aaditya
Verma, Jayant
Rajput, Divyansh
Mahala, Sunil
Khapare, Ketki Suresh
Kalla, Jayateja
Artificial Intelligence
We present NDAI-NeuroMAP, the first neuroscience-domain-specific dense vector embedding model engineered for high-precision information retrieval tasks. Our methodology encompasses the curation of an extensive domain-specific training corpus comprising 500,000 carefully constructed triplets (query-positive-negative configurations), augmented with 250,000 neuroscience-specific definitional entries and 250,000 structured knowledge-graph triplets derived from authoritative neurological ontologies. We employ a sophisticated fine-tuning approach utilizing the FremyCompany/BioLORD-2023 foundation model, implementing a multi-objective optimization framework combining contrastive learning with triplet-based metric learning paradigms. Comprehensive evaluation on a held-out test dataset comprising approximately 24,000 neuroscience-specific queries demonstrates substantial performance improvements over state-of-the-art general-purpose and biomedical embedding models. These empirical findings underscore the critical importance of domain-specific embedding architectures for neuroscience-oriented RAG systems and related clinical natural language processing applications.
title NDAI-NeuroMAP: A Neuroscience-Specific Embedding Model for Domain-Specific Retrieval
topic Artificial Intelligence
url https://arxiv.org/abs/2507.03329