Real-Time Object Detection: Up to 280 FPS on TensorRT

When video surveillance turns into a stream of delays rather than a control tool, both time and money are lost. We develop real-time object detection systems that process camera video without lags or stutters. Our team delivers the project turnkey—from architecture selection to deployment and ongoing support—ensuring stable performance even under high loads.

AI Development Areas

Frequently Asked Questions

Latest works

  • Development of a web application for FEEDME
    Development of a web application for FEEDME
    1344
  • Development of an online store for the company FURNORO
    Development of an online store for the company FURNORO
    1306
  • B2B Advance company logo design
    B2B Advance company logo design
    753
  • Development of a web application for Enviok
    Development of a web application for Enviok
    1049
  • AIDER company logo development
    AIDER company logo development
    992
  • CRM development for Chasseurs
    CRM development for Chasseurs
    1097

Real-Time Object Detection: Solving the Latency Problem

Let's note: when a site has 16 surveillance cameras and the system outputs 15 FPS — that's not real-time. The pipeline described below maintains 30+ FPS on each camera under a total load of up to 32 1080p streams. We guarantee latency below 10ms through hardware decoding (NVDEC) and GPU processing. In our practice — 20+ projects in object detection. Pipeline optimization reduces GPU infrastructure costs by 2–3 times. Contact us for a preliminary assessment of your project — it takes no more than an hour.

Why Real-Time Detection Is a Complex Technical Challenge?

The task requires a balance between accuracy, speed, and hardware load. The naive approach — running every frame through a neural network — hits the GPU limit: modern architectures (YOLOv8, RT-DETR) require 10–30 ms per inference. Without optimization, latency easily exceeds 50 ms, which is critical for robotics or security systems. The solution lies in three areas: selecting a lightweight model (YOLOv8n/m), hardware acceleration (TensorRT, NVDEC), and removing redundant load through frame skipping.

System Architecture

Camera → Frame Capture → Preprocessing → Inference → Postprocessing → Output
↓
↓
Frame Skipping
TensorRT/ONNX Runtime
Resize/Normalize
GPU batching

For RTSP/IP cameras, we use GStreamer or FFmpeg for stream capture with hardware decoding (NVDEC on NVIDIA):

import cv2 # Hardware-accelerated RTSP capture
cap = cv2.VideoCapture(
    'rtsp://camera_ip/stream?'
    'pipeline='
    'rtspsrc location=rtsp://camera_ip/stream !'
    'rtph264depay ! h264parse ! nvh264dec !' # NVDEC
    'videoconvert ! appsink',
    cv2.CAP_GSTREAMER
)

Dynamic batching groups frames from multiple cameras into a single GPU pass, increasing throughput. We support batch sizes up to 32 depending on GPU memory.

How We Achieve 280+ FPS on a Single Camera?

Optimization via TensorRT is the industry standard, confirmed by the NVIDIA TensorRT Developer Guide. Converting YOLOv8 to an FP16 engine provides a 2–5x speedup over native PyTorch. We also apply dynamic batching: frames from multiple cameras are grouped into a single batch, improving GPU utilization.

from ultralytics import YOLO
model = YOLO('yolov8n.pt')

# Export to TensorRT FP16
model.export(
    format='engine',
    half=True,  # FP16 precision
    batch=1,    # or batch=4 for batching
    device=0,
    workspace=4  # GB for optimization
)

Frame skipping — we do not detect every frame. At 30 FPS video, detection on every 3rd frame (10 detections/sec) plus tracking for intermediate frames. Perceived quality is preserved.

Dynamic batching — we group frames from multiple cameras into a batch for a single GPU pass:

class MultiCameraInference:
    def __init__(self, model_path, num_cameras=8):
        self.model = load_trt_model(model_path)
        self.batch_size = num_cameras

    def process_batch(self, frames: list[np.ndarray]) -> list[list]:
        # Preprocessing batch
        batch = preprocess_batch(frames)  # [N, 3, H, W]
        # Single GPU inference for all cameras
        results = self.model.infer(batch)
        return postprocess_batch(results)

TensorRT is 3–4 times faster than native PyTorch. More details in the TensorRT documentation.

Comparison: TensorRT vs PyTorch

Parameter PyTorch FP32 TensorRT FP16 Speedup
YOLOv8n (640x640) 12 ms 4 ms 3x
YOLOv8m (640x640) 28 ms 8 ms 3.5x
YOLOv8l (640x640) 55 ms 14 ms 4x

What Does Using TensorRT Give for Multi-Camera Systems?

For monitoring with 8–32 cameras: one A100/H100 GPU processes up to 32 1080p@30fps streams with YOLOv8n. Architecture: shared inference server (Triton) + separate capture processes for each camera. Savings on GPU infrastructure — up to 3 times compared to a naive implementation.

Throughput:

  • NVIDIA T4 (16GB): 8–12 cameras 1080p with YOLOv8m
  • NVIDIA A100: 24–32 cameras 1080p with YOLOv8l

How We Optimize Latency?

Pipeline latency = capture + decode + preprocess + inference + postprocess + display

Stage Typical Time Optimized Time
Frame capture 5 ms 2 ms (NVDEC)
Preprocessing 8 ms 1 ms (GPU preproc)
YOLOv8n inference 12 ms 4 ms (TRT FP16)
Postprocessing + NMS 5 ms 2 ms
Total 30 ms 9 ms

Additionally, we use pipeline parallelism: capture, preprocessing, and inference execute concurrently on different GPU streams. This allows GPU utilization above 95%.

How We Implement the Solution: Step-by-Step Process

  1. Requirements analysis — determine number of cameras, object classes, acceptable latency.
  2. Data collection and labeling — if custom classes are needed, we prepare a dataset (1000+ frames).
  3. Training and quantization — select YOLOv8n/m, train on GPU, optimize to FP16/INT8.
  4. Integration with infrastructure — configure Triton Inference Server, RTSP capture, deploy in Docker.
  5. Monitoring and support — Grafana dashboards, alerting, model updates.

Scope of Work and Delivery

  • Architecture: capture protocol, post-processing, tracking.
  • Model: YOLO selection, dataset, training, quantization to FP16/INT8.
  • Inference server: configure Triton or TorchServe with batching.
  • Deployment: Docker image with CUDA 12.x, Helm chart for Kubernetes.
  • Documentation: API, metrics, operator instructions.
  • Training: 2–4 hour workshop for your personnel.

Deployment and Monitoring

Docker container with CUDA 12.x + TensorRT. Metrics: FPS per camera, inference latency, GPU utilization, detection count per class per minute. Alerting via Prometheus + Grafana.

System Scale Timeline
1–4 cameras, basic detection 2–3 weeks
8–32 cameras, custom classes 4–7 weeks
50+ cameras, distributed architecture 8–14 weeks

Cost is calculated individually based on scale and complexity. Contact us for a consultation and preliminary estimate. Savings on GPU infrastructure can reach 2–3 times.