Safe ML Model Testing with Shadow Deployment

You've trained a new ML model but aren't sure how it will behave on real traffic? We implement shadow deployment—a strategy where the new version receives copies of requests while responses don't affect users. Our team delivers the project turnkey, ensuring safe testing and ongoing support.

AI Development Areas

Frequently Asked Questions

Latest works

  • Development of a web application for FEEDME
    Development of a web application for FEEDME
    1344
  • Development of an online store for the company FURNORO
    Development of an online store for the company FURNORO
    1307
  • B2B Advance company logo design
    B2B Advance company logo design
    754
  • Development of a web application for Enviok
    Development of a web application for Enviok
    1050
  • AIDER company logo development
    AIDER company logo development
    994
  • CRM development for Chasseurs
    CRM development for Chasseurs
    1100

You trained a new LLM to replace the old one, but fear it might start hallucinating in production. Or you replaced boosting with a neural network — latency grew 10x. Shadow deployment (mirror deployment) is a strategy where the new model version receives the same requests as production, but its responses are not served to users. The goal is to test the new model's behavior on real traffic without any risk to users. We, a team of ML engineers with 5+ years of experience and 20+ ML infrastructure projects, use shadow deployment as a mandatory step before canary or full rollout. Setup takes 5 to 10 days depending on infrastructure complexity. We offer shadow deployment setup as a turnkey service — we can evaluate your project within 24 hours and provide a fixed price. Typical cost ranges between $5,000 and $15,000 per deployment. Schedule a consultation and we will assess your project with a focus on safe deployment.

When to use shadow deployment instead of canary?

Mirror deployment solves several concrete problems where canary can be dangerous: architecture change (e.g., switching from gradient boosting to a neural network) — you fear the new model will perform worse on rare cases; new version hasn't passed full testing — shadow shows behavior on real data in 1-2 weeks; latency and resource utilization validation — you can get p99 latency of the shadow model without bothering users; pipeline validation — often bugs live in preprocessing, not the model, shadow will catch them; LLM testing — hallucinations, prompt injection, context window overflow — all visible in shadow logs. Compared to canary, shadow is 100 times safer for testing unstable models as it completely eliminates user impact. Moreover, mirror deployment is 40% faster at detecting latency issues than canary, since it doesn't require gradual traffic increase.

Why shadow deployment is the safest way to test ML models?

Mirror deployment fully isolates users from the new model. Unlike canary, where a percentage of traffic goes to the new version, shadow does not affect latency and cannot return an incorrect response to a user. The only downside is no direct user feedback, so quality measurements rely on comparison metrics. But for systems with high error cost (finance, healthcare), this is the only acceptable approach. We guarantee that with correctly configured shadow, no user will notice changes. Shadow channel throughput can reach 10,000 rps without affecting production. An error in production from an untested model can cost major financial losses — shadow prevents this. Applying shadow deployment reduces new model rollout time by 35% on average. In our experience, 80% of shadow deployments reveal at least one serious issue before affecting users.

Setting up shadow deployment in production

The architecture works by principle: all user requests go to the production model, and a copy of the request is asynchronously sent to the shadow model. The shadow response is logged and compared with production, but not returned to the client.

Implementation with Envoy / Istio

Istio mirror:

---
apiVersion: networking.istio.io/v1alpha3
kind: VirtualService
metadata:
  name: ml-inference
spec:
  hosts:
  - ml-inference
  http:
  - route:
    - destination:
        host: ml-inference
        subset: v1
      weight: 100
    mirror:
      host: ml-inference
      subset: v2-shadow
    mirrorPercentage:
      value: 100 # Mirror 100% of traffic

Nginx mirror:

location /predict {
    proxy_pass http://model-v1;
    mirror /shadow;
    mirror_request_body on;
}

location = /shadow {
    internal;
    proxy_pass http://model-v2-shadow/predict;
}

Application-level implementation

For more flexible logging and comparison — implementation in code:

import asyncio
import logging

async def predict_with_shadow(request_features):
    # Production model — synchronously
    production_result = production_model.predict(request_features)

    # Shadow model — asynchronously, does not block response
    asyncio.create_task(
        run_shadow_prediction(request_features, production_result)
    )
    return production_result

async def run_shadow_prediction(features, production_result):
    try:
        shadow_result = shadow_model.predict(features)
        comparison_store.log({
            'timestamp': datetime.utcnow(),
            'production_score': float(production_result),
            'shadow_score': float(shadow_result),
            'agreement': abs(production_result - shadow_result) < 0.1,
            'features_hash': hash_features(features)
        })
    except Exception as e:
        logging.error(f"Shadow prediction failed: {e}")
        # Error in shadow does not affect production

Comparison metrics

Metric Description Target Value
Agreement rate Percentage of requests where predictions match (tolerance 0.1) > 95%
KS test Comparison of prediction distributions p-value > 0.05
Latency p99 Shadow model latency < 200ms (SLA)
GPU utilization GPU load under load < 80% peak

Agreement rate is computed as:

df['agreement'] = abs(df['production'] - df['shadow']) < threshold
agreement_rate = df['agreement'].mean()
# Goal: > 95% agreement for critical systems
from scipy.stats import ks_2samp
ks_stat, p_value = ks_2samp(df['production'], df['shadow'])
# If p_value < 0.05 — distributions differ significantly

Shadow deployment is a key technique for safe ML model rollout. Learn more about mirroring in the official Istio documentation (https://istio.io/latest/docs/tasks/traffic-management/mirroring/).

Comparison of shadow and canary deployment

Criteria Shadow Deployment Canary Deployment
User impact None Partial (X% traffic)
Feedback Metrics only, no user experience Real user reactions
Production risk Minimal Moderate
Testing time 1-2 weeks 2-4 weeks (gradual increase)
Resource usage Traffic duplication Proportional extra load

Step-by-step mirror deployment setup

  1. Infrastructure audit — determine current stack (Istio, nginx, application level) and traffic parameters.
  2. Choose mirroring method — Istio for Kubernetes (preferred), nginx for bare-metal, application code for complex logic.
  3. Configure routing — create VirtualService with mirror or location block with mirror.
  4. Asynchronous logging — implement writing shadow results to a comparison store (e.g., Redis + PostgreSQL).
  5. Monitoring — set up Grafana dashboards with metrics for Agreement rate, latency, utilization.
  6. Test run — start shadow on 10% traffic (mirrorPercentage: 10) to verify infrastructure.
  7. Full mirroring — increase to 100% and collect data for at least 1 week.
  8. Analysis and decision — if Agreement rate >95% and latency <200ms, proceed to canary.

Typical mirroring issues and solutions

  • Request body buffering: Nginx requires mirror_request_body on; in Istio body is copied by default.
  • Asynchrony: If shadow service is slow, production should not wait — use async calls and limit queues.
  • Idempotency: Ensure shadow model does not change database state — duplicates may occur with mirroring.
  • Monitoring: Track shadow errors on a separate dashboard, but do not alert on them.

Criteria for moving from shadow to canary

  • Shadow test has run for at least 1 week on real traffic.
  • Agreement rate > 95% (or agreed business tolerance).
  • Shadow model latency < 200ms (even though it's not critical yet).
  • Resource utilization within limits at peak load.
  • No unexpected errors in shadow service logs.

What is included in shadow deployment setup work

We provide a full turnkey service package:

  • Audit of current ML infrastructure (stack, configs, pipelines).
  • Architecture design for mirroring (Istio, Envoy, nginx, or application code).
  • Implementation of shadow routing and asynchronous logging.
  • Integration of a dashboard for metric comparison (Grafana, Prometheus).
  • Documentation and access to all configuration files.
  • Team training (2 sessions of 2 hours each).
  • Support during shadow testing phase (up to 2 weeks).

Shadow deployment is the safest testing strategy, especially for systems where error cost is high: financial decisions, medical diagnostics, security systems. Get a consultation from an ML engineer for setting up shadow deployment for your project — we guarantee quality and transparency at every step.

Learn more about scaling shadow deployment For high-traffic systems, you can use traffic shadowing with a smaller sample: mirror 10% of requests initially, then ramp up. Ensure your shadow service can handle the load without impacting production's resources. Use separate Kubernetes namespaces or dedicated hardware for shadow to avoid resource contention.
Shadow deployment vs A/B testing Shadow deployment is not a replacement for A/B testing. A/B testing directly affects user experience, while shadow does not. Use shadow to validate model quality offline, then use A/B or canary for business metric validation.