A three-person startup deployed Llama-3-70B on a single A100 80GB — and the model crashed with CUDA OOM on every second request. The issue wasn't the GPU but the configuration: vLLM allocated 0.95 of GPU memory to the KV-cache, leaving no headroom for context length variations. We switched gpu-memory-utilization to 0.85 and enabled CPU swapping — the crashes stopped. Such situations are our daily work. Over the last few years, we've deployed more than 30 production inference systems for Large language model on dedicated GPU servers (on-premise or bare metal) and guarantee stable operation even under peak loads. Our team has 5+ years of ML infrastructure experience and has successfully deployed 30+ production inference systems with 99.9% uptime.
A dedicated server delivers predictable performance, no cold starts, and full data control. It is the optimal choice for high-load scenarios with data residency requirements or custom pipelines.
Problems we solve
VRAM miscalculation. For example, Llama-3-70B in BF16 requires 140 GB — not every server can handle that. We use quantization (AWQ/GPTQ) and tensor parallelism to optimally utilize resources. 4-bit quantization reduces VRAM requirement by 75% compared to BF16, allowing a 70B model to fit on 2×A100 and saving up to $15,000 in hardware costs. Mixtral-8x7B (MoE) activates only 13B parameters but needs 90 GB VRAM — easily fits on 2×A100 80GB after 4-bit quantization.
Service crash on OOM. On long contexts (>16k tokens), vLLM may crash. The solution is configuring max-model-len and gpu-memory-utilization paired with a systemd service that automatically restarts the process on failure. Additionally, we set up a watchdog script that checks the health endpoint.
Downtime during model updates. Without blue-green deployment, every update means downtime. We start a new version on a different port, test it, then switch the nginx upstream — no downtime.
How to select the right GPU?
Choosing a GPU depends on model size and required latency. Here's a recommendation table for popular models, with cost considerations: using 4-bit quantization can cut hardware expenses by 75%.
| Model | BF16 VRAM | 4-bit VRAM | Recommended GPUs |
|---|---|---|---|
| 7B | 16 GB | 6-8 GB | RTX 4080, A10G, L4 |
| 13B | 28 GB | 8-10 GB | A30, RTX 4090 (INT8) |
| 70B | 140 GB | 40 GB | 2×A100 80GB, 4×A40 48GB |
| Mixtral 8x7B | 90 GB | 30 GB | 2×A100 80GB |
Quantization to INT4/INT8 reduces VRAM requirements 3–4 times, allowing a 70B model to fit on 2×A100. For a typical 70B deployment, this saves $15,000 in GPU hardware.
vLLM vs TGI: performance comparison
vLLM uses PagedAttention — efficient KV-cache management, giving 50% higher throughput at small batch sizes compared to TGI. In our benchmark of Llama-3-8B with AWQ quantization: vLLM delivered 1200 tokens/sec vs 800 for TGI (batch=32). vLLM outperforms TGI by 1.5x in throughput for batch sizes up to 32. However, TGI handles long context (≥32k tokens) better due to more aggressive tensor parallelism.
| Characteristic | vLLM | TGI |
|---|---|---|
| Throughput (batch=32) | 1200 tok/s | 800 tok/s |
| Long context handling | Good | Excellent |
| Streaming responses | Yes | Yes |
| Tensor parallelism | Yes | Yes, more aggressive |
If your scenario is an online chat with short queries, choose vLLM. If you need to process multi-page documents, go with TGI.
How to prevent OOM errors?
On OOM, first check gpu-memory-utilization — it is often set too high. The recommended value is 0.85–0.90, leaving the rest for CPU swap. Also configure max-model-len proportional to average context length. In systemd, add Restart=always and a health-check script — the service will restart automatically.
Deployment and update process
Stages
- Analytics — evaluate the model, expected RPS, and latency SLA. Select GPU and quantization.
- Design — configure tensor parallel, batch size, quantization, and framework choice.
- Implementation — install CUDA, drivers, deploy vLLM/TGI as a systemd service with watchdog.
- Testing — load testing (locust/vegeta), measure p99 latency, check for long-tail issues.
- Deployment — nginx reverse proxy with rate limiting, SSL, Prometheus + Grafana monitoring.
- Documentation — describe endpoints, configurations, and update procedures.
Zero-downtime update
Classic blue-green: launch the new version on port 8001, test it, switch nginx upstream from port 8000 to 8001, then stop the old service. The entire process takes ~30 seconds; with proper keepalive settings, users don't notice the switch.
What's included in the work
- GPU selection and server configuration.
- Installation of CUDA and drivers (Docker if needed).
- Deployment of vLLM/TGI with systemd + watchdog.
- Nginx reverse proxy with rate limiting and SSL.
- Monitoring (Prometheus + Grafana — dashboards for GPU, latency, throughput).
- Zero-downtime model updates.
- Integration with MLOps pipelines.
- Documentation and team training.
- 30 days of post-deployment support.
Timeline: 2 to 7 days depending on complexity (GPU availability, number of models, HA). Cost is calculated individually — contact us for a project evaluation; it takes 1 business day.
Server configuration and infrastructure
Server setup
Before starting, check GPU availability and driver version:
nvidia-smi nvcc --version Install CUDA 12.1 and cuDNN. On Ubuntu 22.04:
apt-get install -y nvidia-driver-545 apt-get install -y cuda-toolkit-12-1 Check PyTorch: python3 -c "import torch; print(torch.cuda.get_device_name(0))"
Deploy vLLM as a systemd service
# /etc/systemd/system/vllm-llama.service [Unit] Description=vLLM LLaMA-3-8B Inference Server After=network.target [Service] Type=simple User=mlserving WorkingDirectory=/opt/vllm Environment="CUDA_VISIBLE_DEVICES=0,1" Environment="HF_TOKEN=hf_xxx" ExecStart=/opt/vllm/venv/bin/python -m vllm.entrypoints.openai.api_server \ --model /data/models/llama-3-8b-instruct \ --tensor-parallel-size 2 \ --max-model-len 8192 \ --max-num-seqs 128 \ --gpu-memory-utilization 0.92 \ --host 127.0.0.1 \ --port 8000 \ --log-level info Restart=always RestartSec=5 StandardOutput=journal StandardError=journal [Install] WantedBy=multi-user.target Nginx reverse proxy
# /etc/nginx/sites-available/vllm upstream vllm_backend { server 127.0.0.1:8000; keepalive 100; } limit_req_zone $binary_remote_addr zone=api_limit:10m rate=60r/m; server { listen 443 ssl http2; server_name llm.company.internal; ssl_certificate /etc/nginx/ssl/cert.pem; ssl_certificate_key /etc/nginx/ssl/key.pem; location /v1/ { limit_req zone=api_limit burst=20 nodelay; proxy_pass http://vllm_backend; proxy_http_version 1.1; proxy_set_header Connection ""; proxy_read_timeout 300s; proxy_buffering off; chunked_transfer_encoding on; } location /health { proxy_pass http://vllm_backend/health; } } Monitoring and auto-restart
We use Prometheus and Grafana with nvidia_gpu_exporter to track GPU temperature, VRAM utilization, and throughput. Alerts: temperature > 85°C, VRAM > 95%, service unavailable > 30 seconds.
systemd Restart=always plus a watchdog script that checks the health endpoint every 30 seconds. After three failures, it restarts the service.
#!/bin/bash while true; do if ! curl -sf http://127.0.0.1:8000/health > /dev/null; then systemctl restart vllm-llama echo "$(date) - vLLM restarted" >> /var/log/vllm-watchdog.log fi sleep 30 done Technical configuration details
We use AWQ quantization for 4-bit weight representation. This reduces VRAM by 75% with minimal accuracy loss (<1% on benchmarks). For MoE-architecture models (Mixtral), tensor parallelism is mandatory — without it, half the parameters won't fit in VRAM. We recommend --tensor-parallel-size equal to the number of GPUs.
Our engineers have over 5 years of ML/infra experience and have deployed more than 30 inference systems, with a 99.9% uptime guarantee. Every project comes with a configuration guarantee and post-deployment support. If you need integration with a RAG pipeline, fine-tuning, or MLOps CI/CD — we can discuss it during a consultation. Get a preliminary estimate for your project — contact us, we'll respond within a day.







