01 Work

02 Writing

Nothing published yet

03 Deep dives

Nothing published yet

04 Teardowns

Nothing published yet

The story

↑↓ moveEnter openEsc close

All work

Custom RF-DETR Object Detection, Training to Production

A transformer detector fine-tuned, optimised and served at 28 ms per frame with a full MLOps loop.

The problem

The existing YOLO model had plateaued at 0.76 mAP@50, and the team needed better accuracy without giving up real-time latency at production load.

The approach

  1. Fine-tuned an RF-DETR transformer detector on a 40K-image labelled dataset (Roboflow, CVAT) with augmentation and class balancing.
  2. Exported to ONNX and TensorRT (FP16) and served through Triton Inference Server on GPU nodes in Kubernetes, with autoscaling.
  3. CI/CD for retraining, MLflow experiment tracking, and a model registry with canary rollouts.
  4. Drift monitoring in Prometheus/Grafana feeding a human-in-the-loop relabelling queue.

Architecture

How a request flows through the system, top to bottom.

  1. Data
    • 40K labelled images
    • Roboflow + CVAT
    • Relabelling queue
  2. Training
    • PyTorch RF-DETR
    • MLflow tracking
    • GitHub Actions retraining
  3. Optimisation
    • ONNX
    • TensorRT FP16
  4. Serving
    • Triton Inference Server
    • Kubernetes GPU autoscaling
    • Canary rollouts
  5. Monitoring
    • Prometheus
    • Grafana drift dashboards

Results

mAP@50, up from 0.76 for the YOLO baseline
0.89
per frame
28 ms
sustained with autoscaling
500+ rps

Stack

Python · PyTorch · RF-DETR · ONNX · TensorRT · Triton · Docker · Kubernetes · MLflow · GitHub Actions