Building RAG with Milvus Vector Database

When data volumes no longer fit into memory and relational databases struggle with searching millions of vectors, RAG on Milvus comes to the rescue. We develop and implement solutions with the Milvus vector database, configuring hybrid search and indexes for fast and accurate responses. Our team delivers turnkey projects—from audit to support—ensuring reliable operation and scaling alongside your business.

AI Development Areas

Frequently Asked Questions

Latest works

  • Development of a web application for FEEDME
    Development of a web application for FEEDME
    1344
  • Development of an online store for the company FURNORO
    Development of an online store for the company FURNORO
    1306
  • B2B Advance company logo design
    B2B Advance company logo design
    753
  • Development of a web application for Enviok
    Development of a web application for Enviok
    1049
  • AIDER company logo development
    AIDER company logo development
    992
  • CRM development for Chasseurs
    CRM development for Chasseurs
    1097

Building RAG with Milvus Vector Database

When Millions of Chunks No Longer Fit in RAM

Relational databases with cosine distance fail at volumes >100K vectors. We encountered this in a fintech project: 2.5M documentation chunks, 6 languages, requirement P99 latency <500 ms. We chose Milvus — an open-source vector DB with HNSW indexes and GPU acceleration. We'll share how we tuned hybrid search (dense + sparse) in production.

A typical problem is "phantom" relevant results: the model returns semantically similar but contextually irrelevant items. We solved this with hybrid search using RRF reranking. Our experience shows that proper index configuration reduces latency by 40%.

Problems We Solve

Scale: 2.5M chunks is not the limit. Milvus scales horizontally easily: we deployed clusters up to 10 nodes supporting billions of vectors. Choosing the right index is key: for our client we used HNSW (M=32, efConstruction=400) for dense and SPARSE_INVERTED_INDEX for sparse.

Hybrid search: dense vs sparse. Dense vectors capture semantics well but miss exact keyword matches. We combined both approaches with RRF reranking. Relevance increased by 25%.

Multi-tenancy without headaches. Partitioning isolates client data without creating separate collections. We implemented a dynamic partition scheme — easy and secure.

Why Milvus Is the Best Choice for Enterprise RAG?

Milvus outperforms other vector DBs in speed and cost under high load. Compare indexes:

Index Speed (QPS) Recall@10 Memory/vector
HNSW 900 98% 2-4 MB
IVF_FLAT 600 95% 0.5-1 MB
IVF_SQ8 800 92% 0.2-0.5 MB
DISKANN 400 90% ~0 MB (disk)

We use HNSW as a universal choice for production, and DISKANN for archival data on SSD.

Why Milvus Is Better Than Pinecone for High Loads?

Milvus is cheaper at volumes >1M vectors since it doesn't charge per query. Compare:

Feature Milvus (self-hosted) Pinecone (managed)
Cost Hardware + support $0.10/1000 vectors/day
GPU acceleration Yes (NVIDIA CUDA) No
Hybrid search Built-in Via plugins
Control Full Limited

For a typical 10M vector cluster, Milvus is 3-5 times cheaper than Pinecone with the same performance.

How to Configure HNSW Indexes for Minimal Latency?

Key parameters: M (16-64) and efConstruction (200-500). Higher efConstruction means better accuracy but slower building. In production we use M=32, efConstruction=400, and for search ef=100-200. This gives P99 latency <400 ms on 2.5M chunks.

How We Configured Hybrid Search for High Relevance?

Here's the code we applied for our client (fintech, 2.5M chunks):

from pymilvus import connections, Collection, FieldSchema, CollectionSchema, DataType, utility

# Connect to Milvus
connections.connect(
    alias="default",
    host="localhost",
    port="19530"
)

# Or via URI (Milvus Lite for local development)
from pymilvus import MilvusClient
client = MilvusClient("./milvus_local.db")  # SQLite-like file
fields = [
    FieldSchema(name="id", dtype=DataType.INT64, is_primary=True, auto_id=True),
    FieldSchema(name="text", dtype=DataType.VARCHAR, max_length=4096),
    FieldSchema(name="source", dtype=DataType.VARCHAR, max_length=512),
    FieldSchema(name="doc_type", dtype=DataType.VARCHAR, max_length=64),
    FieldSchema(name="page", dtype=DataType.INT32),
    FieldSchema(
        name="dense_vector",
        dtype=DataType.FLOAT_VECTOR,
        dim=1536  # text-embedding-3-small
    ),
    FieldSchema(
        name="sparse_vector",
        dtype=DataType.SPARSE_FLOAT_VECTOR  # BM25
    ),
]
schema = CollectionSchema(fields=fields, description="Corporate Knowledge Base")
collection = Collection(name="knowledge_base", schema=schema)
# Indexes for vector fields
collection.create_index(
    field_name="dense_vector",
    index_params={"metric_type": "COSINE", "index_type": "HNSW", "params": {"M": 16, "efConstruction": 200}}
)
collection.create_index(
    field_name="sparse_vector",
    index_params={"metric_type": "IP", "index_type": "SPARSE_INVERTED_INDEX"}
)
collection.load()
from pymilvus import AnnSearchRequest, RRFRanker

def milvus_hybrid_search(query: str, top_k: int = 5) -> list:
    # Dense vector
    dense_vec = dense_embedder.embed_query(query)
    # Sparse vector (via built-in BM25Encoder)
    sparse_vec = sparse_encoder.encode_queries([query])

    # Two requests for RRF
    dense_req = AnnSearchRequest(
        data=[dense_vec],
        anns_field="dense_vector",
        param={"metric_type": "COSINE", "params": {"ef": 100}},
        limit=30,
    )
    sparse_req = AnnSearchRequest(
        data=sparse_vec,
        anns_field="sparse_vector",
        param={"metric_type": "IP"},
        limit=30,
    )

    # RRF fusion
    results = collection.hybrid_search(
        reqs=[dense_req, sparse_req],
        rerank=RRFRanker(k=60),
        limit=top_k,
        output_fields=["text", "source", "doc_type"],
    )
    return results

Result: 850 QPS with P99 latency <400ms on a 3-node cluster (8 vCPU, 32GB RAM each).

What's Included in the Work

  • Data audit and collection schema design
  • Milvus cluster deployment (Kubernetes / bare-metal) with GPU acceleration
  • Index and parameter tuning (efConstruction, M, ef) for load
  • Ingestion pipeline (PySpark / Kafka / direct)
  • RAG pipeline with LangChain or LlamaIndex
  • Documentation and team training
  • 2 weeks of post-launch support

Process

  1. Analysis: gather requirements, assess volume and query frequency.
  2. Design: select schema, indexes, cluster topology.
  3. Implementation: deploy Milvus, write pipelines, integrate with LLM.
  4. Testing: load testing, parameter tweaks.
  5. Deployment: CI/CD, monitoring (Prometheus + Grafana).

Timeline

  • Milvus cluster setup + schema: 3–5 days
  • Ingestion pipeline with hybrid indexing: 5–10 days
  • RAG pipeline and evaluation: 1–2 weeks
  • Total: 3–5 weeks

Company Metrics

Over 7 years working with vector databases, 30+ Milvus production deployments. We guarantee 99.9% uptime on the cluster and expert support. We'll evaluate your project in 2 days — contact us. Get a consultation on RAG architecture: we'll estimate cost and timeline for your data volume.