Building RAG with Milvus Vector Database
When Millions of Chunks No Longer Fit in RAM
Relational databases with cosine distance fail at volumes >100K vectors. We encountered this in a fintech project: 2.5M documentation chunks, 6 languages, requirement P99 latency <500 ms. We chose Milvus — an open-source vector DB with HNSW indexes and GPU acceleration. We'll share how we tuned hybrid search (dense + sparse) in production.
A typical problem is "phantom" relevant results: the model returns semantically similar but contextually irrelevant items. We solved this with hybrid search using RRF reranking. Our experience shows that proper index configuration reduces latency by 40%.
Problems We Solve
Scale: 2.5M chunks is not the limit. Milvus scales horizontally easily: we deployed clusters up to 10 nodes supporting billions of vectors. Choosing the right index is key: for our client we used HNSW (M=32, efConstruction=400) for dense and SPARSE_INVERTED_INDEX for sparse.
Hybrid search: dense vs sparse. Dense vectors capture semantics well but miss exact keyword matches. We combined both approaches with RRF reranking. Relevance increased by 25%.
Multi-tenancy without headaches. Partitioning isolates client data without creating separate collections. We implemented a dynamic partition scheme — easy and secure.
Why Milvus Is the Best Choice for Enterprise RAG?
Milvus outperforms other vector DBs in speed and cost under high load. Compare indexes:
| Index | Speed (QPS) | Recall@10 | Memory/vector |
|---|---|---|---|
| HNSW | 900 | 98% | 2-4 MB |
| IVF_FLAT | 600 | 95% | 0.5-1 MB |
| IVF_SQ8 | 800 | 92% | 0.2-0.5 MB |
| DISKANN | 400 | 90% | ~0 MB (disk) |
We use HNSW as a universal choice for production, and DISKANN for archival data on SSD.
Why Milvus Is Better Than Pinecone for High Loads?
Milvus is cheaper at volumes >1M vectors since it doesn't charge per query. Compare:
| Feature | Milvus (self-hosted) | Pinecone (managed) |
|---|---|---|
| Cost | Hardware + support | $0.10/1000 vectors/day |
| GPU acceleration | Yes (NVIDIA CUDA) | No |
| Hybrid search | Built-in | Via plugins |
| Control | Full | Limited |
For a typical 10M vector cluster, Milvus is 3-5 times cheaper than Pinecone with the same performance.
How to Configure HNSW Indexes for Minimal Latency?
Key parameters: M (16-64) and efConstruction (200-500). Higher efConstruction means better accuracy but slower building. In production we use M=32, efConstruction=400, and for search ef=100-200. This gives P99 latency <400 ms on 2.5M chunks.
How We Configured Hybrid Search for High Relevance?
Here's the code we applied for our client (fintech, 2.5M chunks):
from pymilvus import connections, Collection, FieldSchema, CollectionSchema, DataType, utility # Connect to Milvus connections.connect( alias="default", host="localhost", port="19530" ) # Or via URI (Milvus Lite for local development) from pymilvus import MilvusClient client = MilvusClient("./milvus_local.db") # SQLite-like file fields = [ FieldSchema(name="id", dtype=DataType.INT64, is_primary=True, auto_id=True), FieldSchema(name="text", dtype=DataType.VARCHAR, max_length=4096), FieldSchema(name="source", dtype=DataType.VARCHAR, max_length=512), FieldSchema(name="doc_type", dtype=DataType.VARCHAR, max_length=64), FieldSchema(name="page", dtype=DataType.INT32), FieldSchema( name="dense_vector", dtype=DataType.FLOAT_VECTOR, dim=1536 # text-embedding-3-small ), FieldSchema( name="sparse_vector", dtype=DataType.SPARSE_FLOAT_VECTOR # BM25 ), ] schema = CollectionSchema(fields=fields, description="Corporate Knowledge Base") collection = Collection(name="knowledge_base", schema=schema) # Indexes for vector fields collection.create_index( field_name="dense_vector", index_params={"metric_type": "COSINE", "index_type": "HNSW", "params": {"M": 16, "efConstruction": 200}} ) collection.create_index( field_name="sparse_vector", index_params={"metric_type": "IP", "index_type": "SPARSE_INVERTED_INDEX"} ) collection.load() from pymilvus import AnnSearchRequest, RRFRanker def milvus_hybrid_search(query: str, top_k: int = 5) -> list: # Dense vector dense_vec = dense_embedder.embed_query(query) # Sparse vector (via built-in BM25Encoder) sparse_vec = sparse_encoder.encode_queries([query]) # Two requests for RRF dense_req = AnnSearchRequest( data=[dense_vec], anns_field="dense_vector", param={"metric_type": "COSINE", "params": {"ef": 100}}, limit=30, ) sparse_req = AnnSearchRequest( data=sparse_vec, anns_field="sparse_vector", param={"metric_type": "IP"}, limit=30, ) # RRF fusion results = collection.hybrid_search( reqs=[dense_req, sparse_req], rerank=RRFRanker(k=60), limit=top_k, output_fields=["text", "source", "doc_type"], ) return results Result: 850 QPS with P99 latency <400ms on a 3-node cluster (8 vCPU, 32GB RAM each).
What's Included in the Work
- Data audit and collection schema design
- Milvus cluster deployment (Kubernetes / bare-metal) with GPU acceleration
- Index and parameter tuning (efConstruction, M, ef) for load
- Ingestion pipeline (PySpark / Kafka / direct)
- RAG pipeline with LangChain or LlamaIndex
- Documentation and team training
- 2 weeks of post-launch support
Process
- Analysis: gather requirements, assess volume and query frequency.
- Design: select schema, indexes, cluster topology.
- Implementation: deploy Milvus, write pipelines, integrate with LLM.
- Testing: load testing, parameter tweaks.
- Deployment: CI/CD, monitoring (Prometheus + Grafana).
Timeline
- Milvus cluster setup + schema: 3–5 days
- Ingestion pipeline with hybrid indexing: 5–10 days
- RAG pipeline and evaluation: 1–2 weeks
- Total: 3–5 weeks
Company Metrics
Over 7 years working with vector databases, 30+ Milvus production deployments. We guarantee 99.9% uptime on the cluster and expert support. We'll evaluate your project in 2 days — contact us. Get a consultation on RAG architecture: we'll estimate cost and timeline for your data volume.







