Real-Time Object Detection: Solving the Latency Problem
Let's note: when a site has 16 surveillance cameras and the system outputs 15 FPS — that's not real-time. The pipeline described below maintains 30+ FPS on each camera under a total load of up to 32 1080p streams. We guarantee latency below 10ms through hardware decoding (NVDEC) and GPU processing. In our practice — 20+ projects in object detection. Pipeline optimization reduces GPU infrastructure costs by 2–3 times. Contact us for a free preliminary assessment of your project — it takes no more than an hour.
Why Real-Time Detection Is a Complex Technical Challenge?
The task requires a balance between accuracy, speed, and hardware load. The naive approach — running every frame through a neural network — hits the GPU limit: modern architectures (YOLOv8, RT-DETR) require 10–30 ms per inference. Without optimization, latency easily exceeds 50 ms, which is critical for robotics or security systems. The solution lies in three areas: selecting a lightweight model (YOLOv8n/m), hardware acceleration (TensorRT, NVDEC), and removing redundant load through frame skipping.
System Architecture
Camera → Frame Capture → Preprocessing → Inference → Postprocessing → Output ↓ ↓ Frame Skipping TensorRT/ONNX Runtime Resize/Normalize GPU batching For RTSP/IP cameras, we use GStreamer or FFmpeg for stream capture with hardware decoding (NVDEC on NVIDIA):
import cv2 # Hardware-accelerated RTSP capture cap = cv2.VideoCapture( 'rtsp://camera_ip/stream?' 'pipeline=' 'rtspsrc location=rtsp://camera_ip/stream !' 'rtph264depay ! h264parse ! nvh264dec !' # NVDEC 'videoconvert ! appsink', cv2.CAP_GSTREAMER ) Dynamic batching groups frames from multiple cameras into a single GPU pass, increasing throughput. We support batch sizes up to 32 depending on GPU memory.
How We Achieve 280+ FPS on a Single Camera?
Optimization via TensorRT is the industry standard, confirmed by the NVIDIA TensorRT Developer Guide. Converting YOLOv8 to an FP16 engine provides a 2–5x speedup over native PyTorch. We also apply dynamic batching: frames from multiple cameras are grouped into a single batch, improving GPU utilization.
from ultralytics import YOLO model = YOLO('yolov8n.pt') # Export to TensorRT FP16 model.export( format='engine', half=True, # FP16 precision batch=1, # or batch=4 for batching device=0, workspace=4 # GB for optimization ) Frame skipping — we do not detect every frame. At 30 FPS video, detection on every 3rd frame (10 detections/sec) plus tracking for intermediate frames. Perceived quality is preserved.
Dynamic batching — we group frames from multiple cameras into a batch for a single GPU pass:
class MultiCameraInference: def __init__(self, model_path, num_cameras=8): self.model = load_trt_model(model_path) self.batch_size = num_cameras def process_batch(self, frames: list[np.ndarray]) -> list[list]: # Preprocessing batch batch = preprocess_batch(frames) # [N, 3, H, W] # Single GPU inference for all cameras results = self.model.infer(batch) return postprocess_batch(results) TensorRT is 3–4 times faster than native PyTorch. More details in the TensorRT documentation.
Comparison: TensorRT vs PyTorch
| Parameter | PyTorch FP32 | TensorRT FP16 | Speedup |
|---|---|---|---|
| YOLOv8n (640x640) | 12 ms | 4 ms | 3x |
| YOLOv8m (640x640) | 28 ms | 8 ms | 3.5x |
| YOLOv8l (640x640) | 55 ms | 14 ms | 4x |
What Does Using TensorRT Give for Multi-Camera Systems?
For monitoring with 8–32 cameras: one A100/H100 GPU processes up to 32 1080p@30fps streams with YOLOv8n. Architecture: shared inference server (Triton) + separate capture processes for each camera. Savings on GPU infrastructure — up to 3 times compared to a naive implementation.
Throughput:
- NVIDIA T4 (16GB): 8–12 cameras 1080p with YOLOv8m
- NVIDIA A100: 24–32 cameras 1080p with YOLOv8l
How We Optimize Latency?
Pipeline latency = capture + decode + preprocess + inference + postprocess + display
| Stage | Typical Time | Optimized Time |
|---|---|---|
| Frame capture | 5 ms | 2 ms (NVDEC) |
| Preprocessing | 8 ms | 1 ms (GPU preproc) |
| YOLOv8n inference | 12 ms | 4 ms (TRT FP16) |
| Postprocessing + NMS | 5 ms | 2 ms |
| Total | 30 ms | 9 ms |
Additionally, we use pipeline parallelism: capture, preprocessing, and inference execute concurrently on different GPU streams. This allows GPU utilization above 95%.
How We Implement the Solution: Step-by-Step Process
- Requirements analysis — determine number of cameras, object classes, acceptable latency.
- Data collection and labeling — if custom classes are needed, we prepare a dataset (1000+ frames).
- Training and quantization — select YOLOv8n/m, train on GPU, optimize to FP16/INT8.
- Integration with infrastructure — configure Triton Inference Server, RTSP capture, deploy in Docker.
- Monitoring and support — Grafana dashboards, alerting, model updates.
Scope of Work and Delivery
- Architecture: capture protocol, post-processing, tracking.
- Model: YOLO selection, dataset, training, quantization to FP16/INT8.
- Inference server: configure Triton or TorchServe with batching.
- Deployment: Docker image with CUDA 12.x, Helm chart for Kubernetes.
- Documentation: API, metrics, operator instructions.
- Training: 2–4 hour workshop for your personnel.
Deployment and Monitoring
Docker container with CUDA 12.x + TensorRT. Metrics: FPS per camera, inference latency, GPU utilization, detection count per class per minute. Alerting via Prometheus + Grafana.
| System Scale | Timeline |
|---|---|
| 1–4 cameras, basic detection | 2–3 weeks |
| 8–32 cameras, custom classes | 4–7 weeks |
| 50+ cameras, distributed architecture | 8–14 weeks |
Cost is calculated individually based on scale and complexity. Contact us for a consultation and preliminary estimate. Savings on GPU infrastructure can reach 2–3 times.







