_version_ 1866913969119166464
author Upreti, Nitish
Simhadri, Harsha Vardhan
Sundar, Hari Sudan
Sundaram, Krishnan
Boshra, Samer
Perumalswamy, Balachandar
Atri, Shivam
Chisholm, Martin
Singh, Revti Raman
Yang, Greg
Hass, Tamara
Dudhey, Nitesh
Pattipaka, Subramanyam
Hildebrand, Mark
Manohar, Magdalen
Moffitt, Jack
Xu, Haiyang
Datha, Naren
Gupta, Suryansh
Krishnaswamy, Ravishankar
Gupta, Prashant
Sahu, Abhishek
Varada, Hemeswari
Barthwal, Sudhanshu
Mor, Ritika
Codella, James
Cooper, Shaun
Pilch, Kevin
Moreno, Simon
Kataria, Aayush
Kulkarni, Santosh
Deshpande, Neil
Sagare, Amar
Billa, Dinesh
Fu, Zishan
Vishal, Vipul
author_facet Upreti, Nitish
Simhadri, Harsha Vardhan
Sundar, Hari Sudan
Sundaram, Krishnan
Boshra, Samer
Perumalswamy, Balachandar
Atri, Shivam
Chisholm, Martin
Singh, Revti Raman
Yang, Greg
Hass, Tamara
Dudhey, Nitesh
Pattipaka, Subramanyam
Hildebrand, Mark
Manohar, Magdalen
Moffitt, Jack
Xu, Haiyang
Datha, Naren
Gupta, Suryansh
Krishnaswamy, Ravishankar
Gupta, Prashant
Sahu, Abhishek
Varada, Hemeswari
Barthwal, Sudhanshu
Mor, Ritika
Codella, James
Cooper, Shaun
Pilch, Kevin
Moreno, Simon
Kataria, Aayush
Kulkarni, Santosh
Deshpande, Neil
Sagare, Amar
Billa, Dinesh
Fu, Zishan
Vishal, Vipul
contents Vector indexing enables semantic search over diverse corpora and has become an important interface to databases for both users and AI agents. Efficient vector search requires deep optimizations in database systems. This has motivated a new class of specialized vector databases that optimize for vector search quality and cost. Instead, we argue that a scalable, high-performance, and cost-efficient vector search system can be built inside a cloud-native operational database like Azure Cosmos DB while leveraging the benefits of a distributed database such as high availability, durability, and scale. We do this by deeply integrating DiskANN, a state-of-the-art vector indexing library, inside Azure Cosmos DB NoSQL. This system uses a single vector index per partition stored in existing index trees, and kept in sync with underlying data. It supports < 20ms query latency over an index spanning 10 million vectors, has stable recall over updates, and offers approximately 43x and 12x lower query cost compared to Pinecone and Zilliz serverless enterprise products. It also scales out to billions of vectors via automatic partitioning. This convergent design presents a point in favor of integrating vector indices into operational databases in the context of recent debates on specialized vector databases, and offers a template for vector indexing in other databases.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05885
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cost-Effective, Low Latency Vector Search with Azure Cosmos DB
Upreti, Nitish
Simhadri, Harsha Vardhan
Sundar, Hari Sudan
Sundaram, Krishnan
Boshra, Samer
Perumalswamy, Balachandar
Atri, Shivam
Chisholm, Martin
Singh, Revti Raman
Yang, Greg
Hass, Tamara
Dudhey, Nitesh
Pattipaka, Subramanyam
Hildebrand, Mark
Manohar, Magdalen
Moffitt, Jack
Xu, Haiyang
Datha, Naren
Gupta, Suryansh
Krishnaswamy, Ravishankar
Gupta, Prashant
Sahu, Abhishek
Varada, Hemeswari
Barthwal, Sudhanshu
Mor, Ritika
Codella, James
Cooper, Shaun
Pilch, Kevin
Moreno, Simon
Kataria, Aayush
Kulkarni, Santosh
Deshpande, Neil
Sagare, Amar
Billa, Dinesh
Fu, Zishan
Vishal, Vipul
Databases
Information Retrieval
H.3.3
Vector indexing enables semantic search over diverse corpora and has become an important interface to databases for both users and AI agents. Efficient vector search requires deep optimizations in database systems. This has motivated a new class of specialized vector databases that optimize for vector search quality and cost. Instead, we argue that a scalable, high-performance, and cost-efficient vector search system can be built inside a cloud-native operational database like Azure Cosmos DB while leveraging the benefits of a distributed database such as high availability, durability, and scale. We do this by deeply integrating DiskANN, a state-of-the-art vector indexing library, inside Azure Cosmos DB NoSQL. This system uses a single vector index per partition stored in existing index trees, and kept in sync with underlying data. It supports < 20ms query latency over an index spanning 10 million vectors, has stable recall over updates, and offers approximately 43x and 12x lower query cost compared to Pinecone and Zilliz serverless enterprise products. It also scales out to billions of vectors via automatic partitioning. This convergent design presents a point in favor of integrating vector indices into operational databases in the context of recent debates on specialized vector databases, and offers a template for vector indexing in other databases.
title Cost-Effective, Low Latency Vector Search with Azure Cosmos DB
topic Databases
Information Retrieval
H.3.3
url https://arxiv.org/abs/2505.05885