Building RAG with Milvus Vector Database

Building RAG with Milvus Vector Database

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1284
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1240
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_logo-aider_0.webp
    AIDER company logo development
    917
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1031

Building RAG with Milvus Vector Database

When Millions of Chunks No Longer Fit in RAM

Relational databases with cosine distance fail at volumes >100K vectors. We encountered this in a fintech project: 2.5M documentation chunks, 6 languages, requirement P99 latency <500 ms. We chose Milvus — an open-source vector DB with HNSW indexes and GPU acceleration. We'll share how we tuned hybrid search (dense + sparse) in production.

A typical problem is "phantom" relevant results: the model returns semantically similar but contextually irrelevant items. We solved this with hybrid search using RRF reranking. Our experience shows that proper index configuration reduces latency by 40%.

Problems We Solve

Scale: 2.5M chunks is not the limit. Milvus scales horizontally easily: we deployed clusters up to 10 nodes supporting billions of vectors. Choosing the right index is key: for our client we used HNSW (M=32, efConstruction=400) for dense and SPARSE_INVERTED_INDEX for sparse.

Hybrid search: dense vs sparse. Dense vectors capture semantics well but miss exact keyword matches. We combined both approaches with RRF reranking. Relevance increased by 25%.

Multi-tenancy without headaches. Partitioning isolates client data without creating separate collections. We implemented a dynamic partition scheme — easy and secure.

Why Milvus Is the Best Choice for Enterprise RAG?

Milvus outperforms other vector DBs in speed and cost under high load. Compare indexes:

Index Speed (QPS) Recall@10 Memory/vector
HNSW 900 98% 2-4 MB
IVF_FLAT 600 95% 0.5-1 MB
IVF_SQ8 800 92% 0.2-0.5 MB
DISKANN 400 90% ~0 MB (disk)

We use HNSW as a universal choice for production, and DISKANN for archival data on SSD.

Why Milvus Is Better Than Pinecone for High Loads?

Milvus is cheaper at volumes >1M vectors since it doesn't charge per query. Compare:

Feature Milvus (self-hosted) Pinecone (managed)
Cost Hardware + support $0.10/1000 vectors/day
GPU acceleration Yes (NVIDIA CUDA) No
Hybrid search Built-in Via plugins
Control Full Limited

For a typical 10M vector cluster, Milvus is 3-5 times cheaper than Pinecone with the same performance.

How to Configure HNSW Indexes for Minimal Latency?

Key parameters: M (16-64) and efConstruction (200-500). Higher efConstruction means better accuracy but slower building. In production we use M=32, efConstruction=400, and for search ef=100-200. This gives P99 latency <400 ms on 2.5M chunks.

How We Configured Hybrid Search for High Relevance?

Here's the code we applied for our client (fintech, 2.5M chunks):

from pymilvus import connections, Collection, FieldSchema, CollectionSchema, DataType, utility # Connect to Milvus connections.connect( alias="default", host="localhost", port="19530" ) # Or via URI (Milvus Lite for local development) from pymilvus import MilvusClient client = MilvusClient("./milvus_local.db") # SQLite-like file 
fields = [ FieldSchema(name="id", dtype=DataType.INT64, is_primary=True, auto_id=True), FieldSchema(name="text", dtype=DataType.VARCHAR, max_length=4096), FieldSchema(name="source", dtype=DataType.VARCHAR, max_length=512), FieldSchema(name="doc_type", dtype=DataType.VARCHAR, max_length=64), FieldSchema(name="page", dtype=DataType.INT32), FieldSchema( name="dense_vector", dtype=DataType.FLOAT_VECTOR, dim=1536 # text-embedding-3-small ), FieldSchema( name="sparse_vector", dtype=DataType.SPARSE_FLOAT_VECTOR # BM25 ), ] schema = CollectionSchema(fields=fields, description="Corporate Knowledge Base") collection = Collection(name="knowledge_base", schema=schema) # Indexes for vector fields collection.create_index( field_name="dense_vector", index_params={"metric_type": "COSINE", "index_type": "HNSW", "params": {"M": 16, "efConstruction": 200}} ) collection.create_index( field_name="sparse_vector", index_params={"metric_type": "IP", "index_type": "SPARSE_INVERTED_INDEX"} ) collection.load() 
from pymilvus import AnnSearchRequest, RRFRanker def milvus_hybrid_search(query: str, top_k: int = 5) -> list: # Dense vector dense_vec = dense_embedder.embed_query(query) # Sparse vector (via built-in BM25Encoder) sparse_vec = sparse_encoder.encode_queries([query]) # Two requests for RRF dense_req = AnnSearchRequest( data=[dense_vec], anns_field="dense_vector", param={"metric_type": "COSINE", "params": {"ef": 100}}, limit=30, ) sparse_req = AnnSearchRequest( data=sparse_vec, anns_field="sparse_vector", param={"metric_type": "IP"}, limit=30, ) # RRF fusion results = collection.hybrid_search( reqs=[dense_req, sparse_req], rerank=RRFRanker(k=60), limit=top_k, output_fields=["text", "source", "doc_type"], ) return results 

Result: 850 QPS with P99 latency <400ms on a 3-node cluster (8 vCPU, 32GB RAM each).

What's Included in the Work

  • Data audit and collection schema design
  • Milvus cluster deployment (Kubernetes / bare-metal) with GPU acceleration
  • Index and parameter tuning (efConstruction, M, ef) for load
  • Ingestion pipeline (PySpark / Kafka / direct)
  • RAG pipeline with LangChain or LlamaIndex
  • Documentation and team training
  • 2 weeks of post-launch support

Process

  1. Analysis: gather requirements, assess volume and query frequency.
  2. Design: select schema, indexes, cluster topology.
  3. Implementation: deploy Milvus, write pipelines, integrate with LLM.
  4. Testing: load testing, parameter tweaks.
  5. Deployment: CI/CD, monitoring (Prometheus + Grafana).

Timeline

  • Milvus cluster setup + schema: 3–5 days
  • Ingestion pipeline with hybrid indexing: 5–10 days
  • RAG pipeline and evaluation: 1–2 weeks
  • Total: 3–5 weeks

Company Metrics

Over 7 years working with vector databases, 30+ Milvus production deployments. We guarantee 99.9% uptime on the cluster and expert support. We'll evaluate your project in 2 days — contact us. Get a consultation on RAG architecture: we'll estimate cost and timeline for your data volume.