Cost-Effective, Low Latency Vector Search with Azure Cosmos DB
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913969119166464 |
|---|---|
| author | Upreti, Nitish Simhadri, Harsha Vardhan Sundar, Hari Sudan Sundaram, Krishnan Boshra, Samer Perumalswamy, Balachandar Atri, Shivam Chisholm, Martin Singh, Revti Raman Yang, Greg Hass, Tamara Dudhey, Nitesh Pattipaka, Subramanyam Hildebrand, Mark Manohar, Magdalen Moffitt, Jack Xu, Haiyang Datha, Naren Gupta, Suryansh Krishnaswamy, Ravishankar Gupta, Prashant Sahu, Abhishek Varada, Hemeswari Barthwal, Sudhanshu Mor, Ritika Codella, James Cooper, Shaun Pilch, Kevin Moreno, Simon Kataria, Aayush Kulkarni, Santosh Deshpande, Neil Sagare, Amar Billa, Dinesh Fu, Zishan Vishal, Vipul |
| author_facet | Upreti, Nitish Simhadri, Harsha Vardhan Sundar, Hari Sudan Sundaram, Krishnan Boshra, Samer Perumalswamy, Balachandar Atri, Shivam Chisholm, Martin Singh, Revti Raman Yang, Greg Hass, Tamara Dudhey, Nitesh Pattipaka, Subramanyam Hildebrand, Mark Manohar, Magdalen Moffitt, Jack Xu, Haiyang Datha, Naren Gupta, Suryansh Krishnaswamy, Ravishankar Gupta, Prashant Sahu, Abhishek Varada, Hemeswari Barthwal, Sudhanshu Mor, Ritika Codella, James Cooper, Shaun Pilch, Kevin Moreno, Simon Kataria, Aayush Kulkarni, Santosh Deshpande, Neil Sagare, Amar Billa, Dinesh Fu, Zishan Vishal, Vipul |
| contents | Vector indexing enables semantic search over diverse corpora and has become an important interface to databases for both users and AI agents. Efficient vector search requires deep optimizations in database systems. This has motivated a new class of specialized vector databases that optimize for vector search quality and cost. Instead, we argue that a scalable, high-performance, and cost-efficient vector search system can be built inside a cloud-native operational database like Azure Cosmos DB while leveraging the benefits of a distributed database such as high availability, durability, and scale. We do this by deeply integrating DiskANN, a state-of-the-art vector indexing library, inside Azure Cosmos DB NoSQL. This system uses a single vector index per partition stored in existing index trees, and kept in sync with underlying data. It supports < 20ms query latency over an index spanning 10 million vectors, has stable recall over updates, and offers approximately 43x and 12x lower query cost compared to Pinecone and Zilliz serverless enterprise products. It also scales out to billions of vectors via automatic partitioning. This convergent design presents a point in favor of integrating vector indices into operational databases in the context of recent debates on specialized vector databases, and offers a template for vector indexing in other databases. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_05885 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Cost-Effective, Low Latency Vector Search with Azure Cosmos DB Upreti, Nitish Simhadri, Harsha Vardhan Sundar, Hari Sudan Sundaram, Krishnan Boshra, Samer Perumalswamy, Balachandar Atri, Shivam Chisholm, Martin Singh, Revti Raman Yang, Greg Hass, Tamara Dudhey, Nitesh Pattipaka, Subramanyam Hildebrand, Mark Manohar, Magdalen Moffitt, Jack Xu, Haiyang Datha, Naren Gupta, Suryansh Krishnaswamy, Ravishankar Gupta, Prashant Sahu, Abhishek Varada, Hemeswari Barthwal, Sudhanshu Mor, Ritika Codella, James Cooper, Shaun Pilch, Kevin Moreno, Simon Kataria, Aayush Kulkarni, Santosh Deshpande, Neil Sagare, Amar Billa, Dinesh Fu, Zishan Vishal, Vipul Databases Information Retrieval H.3.3 Vector indexing enables semantic search over diverse corpora and has become an important interface to databases for both users and AI agents. Efficient vector search requires deep optimizations in database systems. This has motivated a new class of specialized vector databases that optimize for vector search quality and cost. Instead, we argue that a scalable, high-performance, and cost-efficient vector search system can be built inside a cloud-native operational database like Azure Cosmos DB while leveraging the benefits of a distributed database such as high availability, durability, and scale. We do this by deeply integrating DiskANN, a state-of-the-art vector indexing library, inside Azure Cosmos DB NoSQL. This system uses a single vector index per partition stored in existing index trees, and kept in sync with underlying data. It supports < 20ms query latency over an index spanning 10 million vectors, has stable recall over updates, and offers approximately 43x and 12x lower query cost compared to Pinecone and Zilliz serverless enterprise products. It also scales out to billions of vectors via automatic partitioning. This convergent design presents a point in favor of integrating vector indices into operational databases in the context of recent debates on specialized vector databases, and offers a template for vector indexing in other databases. |
| title | Cost-Effective, Low Latency Vector Search with Azure Cosmos DB |
| topic | Databases Information Retrieval H.3.3 |
| url | https://arxiv.org/abs/2505.05885 |