Imagine: your e-commerce site on a Friday evening under load — CPU at 90%, latency rising, and you can't add servers manually in time. Or the opposite: on weekdays, servers idle while you pay for 80% unused capacity. Autoscaling solves both problems. We've set up horizontal scaling for 20+ projects — from startups to enterprise. Below — real code, configs, and common mistakes.
Problems We Solve
Overspending on Resources in Quiet Times
Average web server CPU utilization is 15-20%. Holding capacity for peak means paying for 80% unused power. Autoscaling keeps minimum instances and adds new ones as load grows. For example, an e-commerce site with peak load 20k RPS normally runs 10 servers, though average load is 2k RPS. Autoscaling allows running 2 servers and adding as needed. For infrastructure cost optimization, savings amount to 30-40% of the budget. For one client (e-commerce with 20k RPS peak), implementation cut monthly expenses by $4,500. Cost savings: $4,500 per month.
Downtime During Sudden Spikes
A DDoS or viral post can multiply load tenfold in minutes. Without scaling, the site goes down. Our solution reacts within seconds using Target Tracking and a 60s scale-out cooldown.
Complexity of Manual Management
Even an experienced admin can't manually add servers in time. Automation eliminates human error.
How We Do It: Stack and Case Study
AWS Auto Scaling Group with Target Tracking
We use Terraform for infrastructure as code. Example configuration:
resource "aws_autoscaling_group" "app" { name = "app-asg" min_size = 2 max_size = 20 desired_capacity = 3 vpc_zone_identifier = var.private_subnet_ids launch_template { id = aws_launch_template.app.id version = "$Latest" } health_check_type = "ELB" health_check_grace_period = 60 target_group_arns = [aws_lb_target_group.app.arn] } # Target Tracking: keep CPU at 60% resource "aws_autoscaling_policy" "cpu_tracking" { name = "cpu-tracking" autoscaling_group_name = aws_autoscaling_group.app.name policy_type = "TargetTrackingScaling" target_tracking_configuration { predefined_metric_specification { predefined_metric_type = "ASGAverageCPUUtilization" } target_value = 60.0 scale_in_cooldown = 300 scale_out_cooldown = 60 } } Scale-out cooldown (60s) is shorter than scale-in (300s) — reacting quickly to growth, slowly removing resources.
Kubernetes HPA with Custom Metrics
Horizontal Pod Autoscaler combined with Prometheus Adapter allows scaling by custom metrics:
apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: app-hpa spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: app minReplicas: 2 maxReplicas: 50 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 60 - type: Pods pods: metric: name: http_requests_per_second target: type: AverageValue averageValue: "100" The http_requests_per_second metric comes from Prometheus via kube-state-metrics and Prometheus Adapter.
KEDA: Scaling by External Sources
KEDA (Kubernetes Event-Driven Autoscaling) scales pods by queue length from RabbitMQ, Kafka, SQS:
apiVersion: keda.sh/v1alpha1 kind: ScaledObject metadata: name: queue-processor spec: scaleTargetRef: name: worker-deployment minReplicaCount: 1 maxReplicaCount: 30 triggers: - type: rabbitmq metadata: host: amqp://rabbitmq:5672/ queueName: tasks queueLength: "50" Scaling to zero when the queue is empty — saving resources.
Predictive Scaling
AWS Predictive Scaling forecasts load based on historical data (minimum 14 days) and adds resources proactively:
resource "aws_autoscaling_policy" "predictive" { name = "predictive" autoscaling_group_name = aws_autoscaling_group.app.name policy_type = "PredictiveScaling" predictive_scaling_configuration { mode = "ForecastAndScale" scheduling_buffer_time = 300 max_capacity_breach_behavior = "IncreaseMaxCapacity" metric_specification { target_value = 60 predefined_scaling_metric_specification { predefined_metric_type = "ASGAverageCPUUtilization" } predefined_load_metric_specification { predefined_metric_type = "ASGTotalNetworkIn" } } } } Scaling Method Comparison
| Method | Metrics | Reaction Speed | Complexity |
|---|---|---|---|
| AWS ASG + Target Tracking | CPU, Network, Request Count | 1-2 minutes | Low |
| Kubernetes HPA | CPU, Memory, Custom | 30-60 seconds | Medium |
| KEDA | Queue Length, External | 10-30 seconds | Medium |
| Predictive Scaling | Historical Trends | Proactive | High |
How to Choose the Right Metric?
The metric is key to effective scaling. CPU lags, Request Rate needs a baseline, P95 is the most accurate but complex. In practice we use a combination of CPU + Request Rate. For async workers — Queue Depth. According to AWS Auto Scaling documentation (Amazon EC2 Auto Scaling) supports multiple metrics in one policy. Infrastructure costs drop by 35% with properly selected metrics.
What If Scaling Doesn't Trigger on Time?
Check cooldown (scale-out faster than scale-in), health check grace period. Sometimes the metric hasn't updated. We recommend a load test with k6: k6 run --vus 1000 --duration 10m script.js.
Process
- Analytics: analyze load profile, historical data, select metrics.
- Design: determine scaling type, tools (AWS ASG, K8s HPA, KEDA).
- Implementation: write IaC (Terraform/Pulumi), set up monitoring (CloudWatch, Prometheus).
- Testing: load testing, verify response time and downtime.
- Deploy: production rollout, alerts (SNS, PagerDuty).
- Support: monitor effectiveness, adjust thresholds.
What's Included
- Architectural documentation
- Infrastructure code (Terraform/Helm)
- Monitoring and alerts
- Load test case
- Team training (1 session)
Implementation Timelines
| Component | Timeframe |
|---|---|
| ASG + Target Tracking (AWS) | 2-3 days |
| HPA + Prometheus Adapter (K8s) | 3-5 days |
| KEDA for queue-based workloads | 2-3 days |
| Predictive Scaling | 1-2 days (after 14 days of data) |
| Load testing + tuning | 2-3 days |
Cost is calculated individually after load audit. To get an accurate estimate and timeline, contact us.
Common Mistakes (click to expand)
- Incorrect cooldown settings (flapping)
- Scaling only by CPU (ignoring memory/network)
- No health check during scale-in (dropping active sessions)
- Too wide min/max range (risk of infinite scaling)
Our Experience
Over 5 years in high-load infrastructure. We've implemented horizontal scaling for 20+ projects — from startups to enterprise (e-commerce, fintech). We use proven solutions with SLA guarantees. Want to set up autoscaling? Get a consultation — we'll evaluate your project and propose the optimal solution.







