Real-Time Object Detection: Up to 280 FPS on TensorRT

Real-Time Object Detection: Solving the Latency Problem

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1284
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1240
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_logo-aider_0.webp
    AIDER company logo development
    917
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1031

Real-Time Object Detection: Solving the Latency Problem

Let's note: when a site has 16 surveillance cameras and the system outputs 15 FPS — that's not real-time. The pipeline described below maintains 30+ FPS on each camera under a total load of up to 32 1080p streams. We guarantee latency below 10ms through hardware decoding (NVDEC) and GPU processing. In our practice — 20+ projects in object detection. Pipeline optimization reduces GPU infrastructure costs by 2–3 times. Contact us for a free preliminary assessment of your project — it takes no more than an hour.

Why Real-Time Detection Is a Complex Technical Challenge?

The task requires a balance between accuracy, speed, and hardware load. The naive approach — running every frame through a neural network — hits the GPU limit: modern architectures (YOLOv8, RT-DETR) require 10–30 ms per inference. Without optimization, latency easily exceeds 50 ms, which is critical for robotics or security systems. The solution lies in three areas: selecting a lightweight model (YOLOv8n/m), hardware acceleration (TensorRT, NVDEC), and removing redundant load through frame skipping.

System Architecture

Camera → Frame Capture → Preprocessing → Inference → Postprocessing → Output ↓ ↓ Frame Skipping TensorRT/ONNX Runtime Resize/Normalize GPU batching 

For RTSP/IP cameras, we use GStreamer or FFmpeg for stream capture with hardware decoding (NVDEC on NVIDIA):

import cv2 # Hardware-accelerated RTSP capture cap = cv2.VideoCapture( 'rtsp://camera_ip/stream?' 'pipeline=' 'rtspsrc location=rtsp://camera_ip/stream !' 'rtph264depay ! h264parse ! nvh264dec !' # NVDEC 'videoconvert ! appsink', cv2.CAP_GSTREAMER ) 

Dynamic batching groups frames from multiple cameras into a single GPU pass, increasing throughput. We support batch sizes up to 32 depending on GPU memory.

How We Achieve 280+ FPS on a Single Camera?

Optimization via TensorRT is the industry standard, confirmed by the NVIDIA TensorRT Developer Guide. Converting YOLOv8 to an FP16 engine provides a 2–5x speedup over native PyTorch. We also apply dynamic batching: frames from multiple cameras are grouped into a single batch, improving GPU utilization.

from ultralytics import YOLO model = YOLO('yolov8n.pt') # Export to TensorRT FP16 model.export( format='engine', half=True, # FP16 precision batch=1, # or batch=4 for batching device=0, workspace=4 # GB for optimization ) 

Frame skipping — we do not detect every frame. At 30 FPS video, detection on every 3rd frame (10 detections/sec) plus tracking for intermediate frames. Perceived quality is preserved.

Dynamic batching — we group frames from multiple cameras into a batch for a single GPU pass:

class MultiCameraInference: def __init__(self, model_path, num_cameras=8): self.model = load_trt_model(model_path) self.batch_size = num_cameras def process_batch(self, frames: list[np.ndarray]) -> list[list]: # Preprocessing batch batch = preprocess_batch(frames) # [N, 3, H, W] # Single GPU inference for all cameras results = self.model.infer(batch) return postprocess_batch(results) 

TensorRT is 3–4 times faster than native PyTorch. More details in the TensorRT documentation.

Comparison: TensorRT vs PyTorch

Parameter PyTorch FP32 TensorRT FP16 Speedup
YOLOv8n (640x640) 12 ms 4 ms 3x
YOLOv8m (640x640) 28 ms 8 ms 3.5x
YOLOv8l (640x640) 55 ms 14 ms 4x

What Does Using TensorRT Give for Multi-Camera Systems?

For monitoring with 8–32 cameras: one A100/H100 GPU processes up to 32 1080p@30fps streams with YOLOv8n. Architecture: shared inference server (Triton) + separate capture processes for each camera. Savings on GPU infrastructure — up to 3 times compared to a naive implementation.

Throughput:

  • NVIDIA T4 (16GB): 8–12 cameras 1080p with YOLOv8m
  • NVIDIA A100: 24–32 cameras 1080p with YOLOv8l

How We Optimize Latency?

Pipeline latency = capture + decode + preprocess + inference + postprocess + display

Stage Typical Time Optimized Time
Frame capture 5 ms 2 ms (NVDEC)
Preprocessing 8 ms 1 ms (GPU preproc)
YOLOv8n inference 12 ms 4 ms (TRT FP16)
Postprocessing + NMS 5 ms 2 ms
Total 30 ms 9 ms

Additionally, we use pipeline parallelism: capture, preprocessing, and inference execute concurrently on different GPU streams. This allows GPU utilization above 95%.

How We Implement the Solution: Step-by-Step Process

  1. Requirements analysis — determine number of cameras, object classes, acceptable latency.
  2. Data collection and labeling — if custom classes are needed, we prepare a dataset (1000+ frames).
  3. Training and quantization — select YOLOv8n/m, train on GPU, optimize to FP16/INT8.
  4. Integration with infrastructure — configure Triton Inference Server, RTSP capture, deploy in Docker.
  5. Monitoring and support — Grafana dashboards, alerting, model updates.

Scope of Work and Delivery

  • Architecture: capture protocol, post-processing, tracking.
  • Model: YOLO selection, dataset, training, quantization to FP16/INT8.
  • Inference server: configure Triton or TorchServe with batching.
  • Deployment: Docker image with CUDA 12.x, Helm chart for Kubernetes.
  • Documentation: API, metrics, operator instructions.
  • Training: 2–4 hour workshop for your personnel.

Deployment and Monitoring

Docker container with CUDA 12.x + TensorRT. Metrics: FPS per camera, inference latency, GPU utilization, detection count per class per minute. Alerting via Prometheus + Grafana.

System Scale Timeline
1–4 cameras, basic detection 2–3 weeks
8–32 cameras, custom classes 4–7 weeks
50+ cameras, distributed architecture 8–14 weeks

Cost is calculated individually based on scale and complexity. Contact us for a consultation and preliminary estimate. Savings on GPU infrastructure can reach 2–3 times.