You trained a new LLM to replace the old one, but fear it might start hallucinating in production. Or you replaced boosting with a neural network — latency grew 10x. Shadow deployment (mirror deployment) is a strategy where the new model version receives the same requests as production, but its responses are not served to users. The goal is to test the new model's behavior on real traffic without any risk to users. We, a team of ML engineers with 5+ years of experience and 20+ ML infrastructure projects, use shadow deployment as a mandatory step before canary or full rollout. Setup takes 5 to 10 days depending on infrastructure complexity. We offer shadow deployment setup as a turnkey service — we can evaluate your project within 24 hours and provide a fixed price. Typical cost ranges between $5,000 and $15,000 per deployment. Schedule a consultation and we will assess your project with a focus on safe deployment.
When to use shadow deployment instead of canary?
Mirror deployment solves several concrete problems where canary can be dangerous: architecture change (e.g., switching from gradient boosting to a neural network) — you fear the new model will perform worse on rare cases; new version hasn't passed full testing — shadow shows behavior on real data in 1-2 weeks; latency and resource utilization validation — you can get p99 latency of the shadow model without bothering users; pipeline validation — often bugs live in preprocessing, not the model, shadow will catch them; LLM testing — hallucinations, prompt injection, context window overflow — all visible in shadow logs. Compared to canary, shadow is 100 times safer for testing unstable models as it completely eliminates user impact. Moreover, mirror deployment is 40% faster at detecting latency issues than canary, since it doesn't require gradual traffic increase.
Why shadow deployment is the safest way to test ML models?
Mirror deployment fully isolates users from the new model. Unlike canary, where a percentage of traffic goes to the new version, shadow does not affect latency and cannot return an incorrect response to a user. The only downside is no direct user feedback, so quality measurements rely on comparison metrics. But for systems with high error cost (finance, healthcare), this is the only acceptable approach. We guarantee that with correctly configured shadow, no user will notice changes. Shadow channel throughput can reach 10,000 rps without affecting production. An error in production from an untested model can cost major financial losses — shadow prevents this. Applying shadow deployment reduces new model rollout time by 35% on average. In our experience, 80% of shadow deployments reveal at least one serious issue before affecting users.
Setting up shadow deployment in production
The architecture works by principle: all user requests go to the production model, and a copy of the request is asynchronously sent to the shadow model. The shadow response is logged and compared with production, but not returned to the client.
Implementation with Envoy / Istio
Istio mirror:
apiVersion: networking.istio.io/v1alpha3 kind: VirtualService metadata: name: ml-inference spec: hosts: - ml-inference http: - route: - destination: host: ml-inference subset: v1 weight: 100 mirror: host: ml-inference subset: v2-shadow mirrorPercentage: value: 100 # Mirror 100% of traffic Nginx mirror:
location /predict { proxy_pass http://model-v1; mirror /shadow; mirror_request_body on; } location = /shadow { internal; proxy_pass http://model-v2-shadow/predict; } Application-level implementation
For more flexible logging and comparison — implementation in code:
import asyncio import logging async def predict_with_shadow(request_features): # Production model — synchronously production_result = production_model.predict(request_features) # Shadow model — asynchronously, does not block response asyncio.create_task( run_shadow_prediction(request_features, production_result) ) return production_result async def run_shadow_prediction(features, production_result): try: shadow_result = shadow_model.predict(features) comparison_store.log({ 'timestamp': datetime.utcnow(), 'production_score': float(production_result), 'shadow_score': float(shadow_result), 'agreement': abs(production_result - shadow_result) < 0.1, 'features_hash': hash_features(features) }) except Exception as e: logging.error(f"Shadow prediction failed: {e}") # Error in shadow does not affect production Comparison metrics
| Metric | Description | Target Value |
|---|---|---|
| Agreement rate | Percentage of requests where predictions match (tolerance 0.1) | > 95% |
| KS test | Comparison of prediction distributions | p-value > 0.05 |
| Latency p99 | Shadow model latency | < 200ms (SLA) |
| GPU utilization | GPU load under load | < 80% peak |
Agreement rate is computed as:
df['agreement'] = abs(df['production'] - df['shadow']) < threshold agreement_rate = df['agreement'].mean() # Goal: > 95% agreement for critical systems from scipy.stats import ks_2samp ks_stat, p_value = ks_2samp(df['production'], df['shadow']) # If p_value < 0.05 — distributions differ significantly Shadow deployment is a key technique for safe ML model rollout. Learn more about mirroring in the official Istio documentation (https://istio.io/latest/docs/tasks/traffic-management/mirroring/).
Comparison of shadow and canary deployment
| Criteria | Shadow Deployment | Canary Deployment |
|---|---|---|
| User impact | None | Partial (X% traffic) |
| Feedback | Metrics only, no user experience | Real user reactions |
| Production risk | Minimal | Moderate |
| Testing time | 1-2 weeks | 2-4 weeks (gradual increase) |
| Resource usage | Traffic duplication | Proportional extra load |
Step-by-step mirror deployment setup
- Infrastructure audit — determine current stack (Istio, nginx, application level) and traffic parameters.
- Choose mirroring method — Istio for Kubernetes (preferred), nginx for bare-metal, application code for complex logic.
- Configure routing — create VirtualService with mirror or location block with mirror.
- Asynchronous logging — implement writing shadow results to a comparison store (e.g., Redis + PostgreSQL).
- Monitoring — set up Grafana dashboards with metrics for Agreement rate, latency, utilization.
- Test run — start shadow on 10% traffic (mirrorPercentage: 10) to verify infrastructure.
- Full mirroring — increase to 100% and collect data for at least 1 week.
- Analysis and decision — if Agreement rate >95% and latency <200ms, proceed to canary.
Typical mirroring issues and solutions
- Request body buffering: Nginx requires
mirror_request_body on; in Istio body is copied by default. - Asynchrony: If shadow service is slow, production should not wait — use async calls and limit queues.
- Idempotency: Ensure shadow model does not change database state — duplicates may occur with mirroring.
- Monitoring: Track shadow errors on a separate dashboard, but do not alert on them.
Criteria for moving from shadow to canary
- Shadow test has run for at least 1 week on real traffic.
- Agreement rate > 95% (or agreed business tolerance).
- Shadow model latency < 200ms (even though it's not critical yet).
- Resource utilization within limits at peak load.
- No unexpected errors in shadow service logs.
What is included in shadow deployment setup work
We provide a full turnkey service package:
- Audit of current ML infrastructure (stack, configs, pipelines).
- Architecture design for mirroring (Istio, Envoy, nginx, or application code).
- Implementation of shadow routing and asynchronous logging.
- Integration of a dashboard for metric comparison (Grafana, Prometheus).
- Documentation and access to all configuration files.
- Team training (2 sessions of 2 hours each).
- Support during shadow testing phase (up to 2 weeks).
Shadow deployment is the safest testing strategy, especially for systems where error cost is high: financial decisions, medical diagnostics, security systems. Get a consultation from an ML engineer for setting up shadow deployment for your project — we guarantee quality and transparency at every step.







