Which vector database is best for scaling an NLP application with millions of embeddings?
We are transitioning our semantic search tool from a small prototype to a production environment. We currently have about 5 million document embeddings. I’m torn between Pinecone, Weaviate, and Milvus. We need low latency for k-nearest neighbor searches, but we also need to keep costs manageable as our dataset grows. What has your experience been with "cold start" latency and filtering?
2025-07-05 in Data Science by Douglas Wright
| 12921 Views
All answers to this question.
If you want a "set it and forget it" managed service, Pinecone is excellent, but the costs can scale aggressively. For 5 million vectors, you should look into Milvus or Weaviate if you have the DevOps capacity to host them on-prem or in your VPC. Milvus is particularly good at handling massive scale due to its distributed architecture, though the setup is complex. Weaviate’s GraphQL interface is very developer-friendly. In 2024, I’ve seen a shift toward pgvector (PostgreSQL) for teams who already use SQL and don't want to manage a separate DB for a few million vectors. It’s surprisingly performant now.
Answered 2025-07-08 by Theresa Mendoza
Have you looked at the HNSW (Hierarchical Navigable Small World) index parameters? Sometimes the latency isn't the DB's fault, but rather how the index is tuned.
Answered 2025-07-10 by Raymond Grant
-
Raymond, you hit the nail on the head. We were struggling with Weaviate latency until we adjusted the 'efConstruction' and 'maxConnections' parameters. It’s a balancing act between search speed and recall accuracy. If you set 'ef' too low, you get results in 10ms, but they might not be the actual nearest neighbors. Testing these hyperparameters on a representative subset of your data is a mandatory step before going live.
Commented 2025-07-13 by Arthur Boyd
Don't ignore Qdrant. It’s written in Rust, extremely fast, and their filtering capabilities for metadata are some of the most intuitive I’ve used in any vector project.
Answered 2025-07-15 by Pamela Duncan
-
I’ve heard great things about Qdrant’s performance. I think I’ll set up a benchmark test between Qdrant and Milvus this week to see which one handles our specific metadata filters better.
Commented 2025-07-17 by Douglas Wright
Write a Comment
Your email address will not be published. Required fields are marked (*)

