Custom Object Detection: Training YOLOv8, YOLO11, RT-DETR Models

You collected a dataset of production defects, ran `yolo train`, and [email protected] plateaued at 0.6. Sound familiar? We see this on nearly every other project. Our team helps companies train object detectors for their specific tasks: from custom dataset collection and image annotation to TensorRT optimiza

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1285
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1241
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    983
  • image_logo-aider_0.webp
    AIDER company logo development
    919
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1033

You collected a dataset of production defects, ran yolo train, and [email protected] plateaued at 0.6. Sound familiar? We see this on nearly every other project. Our team helps companies train object detectors for their specific tasks: from custom dataset collection and image annotation to TensorRT optimization. We guarantee mAP50 > 90% on a holdout set. Our services typically cost $2,500–$4,000, saving you up to 50% compared to in-house development. We will evaluate your project for free — contact us.

How to Tune Hyperparameters for Small Objects

YOLOv8 is the de facto standard for most production detection tasks. But the gap between hitting yolo train and achieving [email protected] > 0.85 on real-world data spans several iterations, each with specific decisions. For small objects (less than 5% of image area), careful hyperparameter selection is critical. Here's a typical config for medium-sized models:

Example configuration and training code
model: yolov8m.pt data: dataset.yaml imgsz: 640 batch: 16 epochs: 200 optimizer: AdamW lr0: 0.001 lrf: 0.01 momentum: 0.937 weight_decay: 0.0005 warmup_epochs: 3.0 mosaic: 1.0 mixup: 0.15 copy_paste: 0.1 degrees: 10.0 translate: 0.1 scale: 0.5 flipud: 0.0 fliplr: 0.5 hsv_h: 0.015 hsv_s: 0.7 hsv_v: 0.4 
from ultralytics import YOLO model = YOLO('yolov8m.pt') results = model.train( data='dataset/data.yaml', imgsz=640, batch=16, epochs=200, device='0', project='runs/detect', name='defect_v1', save_period=10, val=True, plots=True, patience=50 ) 

Increasing imgsz to 1280 boosts [email protected] for small objects (15–40 px) by 5–8%, but training time quadruples and VRAM usage jumps to 24 GB. For scenes with small objects, we also disable mosaic in the last 10 epochs and reduce its weight to 0.5.

Typical Training Problems: Three Root Causes

The most common scenario: loss drops, val mAP rises to ~0.6, then stagnates. Analysis of the confusion matrix shows systematic false positives for one class. Three main causes:

  1. Annotation errors. Even 5% incorrect bboxes ruin training for small classes. Our diagnostic tool is an audit script that checks for out-of-bounds, micro-bboxes, and duplicates.
import numpy as np from pathlib import Path def audit_annotation_quality(labels_dir: str) -> dict: issues = {'out_of_bounds': [], 'tiny_boxes': [], 'duplicates': []} for label_path in Path(labels_dir).glob('*.txt'): boxes = np.loadtxt(label_path, ndmin=2) if boxes.shape[0] == 0: continue cls_ids, cx, cy, bw, bh = (boxes[:, i] for i in range(5)) oob = (cx - bw/2 < 0) | (cx + bw/2 > 1) | \ (cy - bh/2 < 0) | (cy + bh/2 > 1) if oob.any(): issues['out_of_bounds'].append(str(label_path)) tiny = (bw * bh) < 0.0004 if tiny.any(): issues['tiny_boxes'].append(str(label_path)) return issues 
  1. Imbalanced dataset. Although YOLOv8 is anchor-free, spatial bias—objects of one class occupying the same image region—causes the model to learn correlations with background. Solution: stratified splitting and augmentations: RandomPerspective, Copy-Paste.

  2. Overly aggressive mosaic. For small objects, mosaic shrinks them 2–4 times, making them undetectable. We enable mosaic only for the first 90% of epochs; for datasets with objects <20 px, we reduce its weight to 0.5 and combine with MixUp.

How Training Works: Step-by-Step Plan

  1. Data analysis. Check objects per class, bbox sizes, distribution across images. If fewer than 500 objects per class, use pretrained YOLOv8m weights and freeze the backbone.
  2. Configuration. Select imgsz, batch size, optimizer, LR schedule. For small objects: imgsz=960 or 1280, reduce mosaic.
  3. Launch training. Use Ultralytics HUB or a local script with monitoring via TensorBoard/WandB. Stop at patience=50.
  4. Validation. Examine confusion matrix, Precision-Recall curves, [email protected]:0.95. If [email protected] < 0.8, return to step 1.
  5. Export. Convert to TensorRT (FP16) or ONNX. Check latency on target GPU.

When YOLO Falls Short: RT-DETR Advantages

RT-DETR (Real-Time DEtection TRansformer) is a transformer-based detector without NMS. It outperforms YOLOv8 on scenes with heavy occlusions and non-standard aspect ratios. On a small defect detection task (objects 15–40px), RT-DETR-L achieves 7% higher [email protected] than YOLOv8m with only 4ms extra latency. As noted in the Ultralytics documentation, RT-DETR provides a state-of-the-art accuracy-speed trade-off. Comparison:

Model [email protected] [email protected]:0.95 Latency (RTX3080) VRAM
YOLOv8n 0.724 0.421 2.3ms 2.1GB
YOLOv8m 0.811 0.513 5.1ms 5.8GB
YOLOv8l 0.837 0.541 8.2ms 8.1GB
RT-DETR-L 0.869 0.574 9.8ms 9.4GB
YOLO11l 0.845 0.553 7.9ms 7.8GB

Example RT-DETR training:

from ultralytics import RTDETR model = RTDETR('rtdetr-l.pt') model.train( data='dataset/data.yaml', imgsz=640, batch=8, epochs=100, device='0', optimizer='AdamW', lr0=0.0001, warmup_epochs=2 ) 

TensorRT for Production

For production, we export the trained model to TensorRT with FP16—achieving ~2x inference speedup over PyTorch with mAP drop of at most 0.5%. We also consider INT8 quantization if up to 1% accuracy loss is acceptable (memory savings up to 50%). Example export:

from ultralytics import YOLO trained_model = YOLO('runs/detect/defect_v1/weights/best.pt') trained_model.export( format='engine', device=0, half=True, dynamic=False, imgsz=640, batch=1, workspace=4 ) 

What's Included in Our Work

  • Problem analysis and architecture selection (YOLO / RT-DETR / Detectron2 / custom transformers)
  • Custom dataset collection and image annotation (conversion from COCO, Pascal VOC, Supervisely, CVAT)
  • Detector fine-tuning with hyperparameter and augmentation optimization (Grid Search, Bayesian Optimization)
  • Validation on a holdout set: mAP, confusion matrix, PR-curve, FPS
  • Model optimization for inference: export to TensorRT / ONNX with FP16/INT8 quantization
  • Model card documentation and inference script
  • Post-deployment support (2 weeks)

Timelines and Pricing

Task Timeline
Fine-tuning YOLOv8 (ready dataset) 1–2 weeks
Full cycle: data → training → optimization 4–7 weeks
Custom detector (new head architecture) 8–14 weeks

Pricing is assessed individually per project. Typical cost for fine-tuning YOLOv8 on a ready dataset ranges from $1,000 to $5,000, depending on data size and complexity. It includes a fixed SLA—we guarantee achieving the target metric ([email protected] > 90%) or we refine the model at no extra cost. Get a consultation on your project—we'll explain which approach delivers maximum accuracy given your budget and time constraints.

About Our Team

Over 10 years in computer vision. We have trained 50+ models for tasks including industrial defect detection, satellite object detection, people counting in retail, and animal recognition on farms. We work with YOLO, RT-DETR, Detectron2, DETR, Swin Transformer. Full stack: PyTorch, TensorRT, ONNX, NVIDIA Triton. Our specialists are Kaggle Grandmasters and authors of open-source CV libraries. Our team has 10+ years of experience and 50+ successful projects, providing reliable model training services.

Contact us to discuss your project.