All work
Custom RF-DETR Object Detection, Training to Production
A transformer detector fine-tuned, optimised and served at 28 ms per frame with a full MLOps loop.
The problem
The existing YOLO model had plateaued at 0.76 mAP@50, and the team needed better accuracy without giving up real-time latency at production load.
The approach
- Fine-tuned an RF-DETR transformer detector on a 40K-image labelled dataset (Roboflow, CVAT) with augmentation and class balancing.
- Exported to ONNX and TensorRT (FP16) and served through Triton Inference Server on GPU nodes in Kubernetes, with autoscaling.
- CI/CD for retraining, MLflow experiment tracking, and a model registry with canary rollouts.
- Drift monitoring in Prometheus/Grafana feeding a human-in-the-loop relabelling queue.
Architecture
How a request flows through the system, top to bottom.
- Data
- 40K labelled images
- Roboflow + CVAT
- Relabelling queue
- Training
- PyTorch RF-DETR
- MLflow tracking
- GitHub Actions retraining
- Optimisation
- ONNX
- TensorRT FP16
- Serving
- Triton Inference Server
- Kubernetes GPU autoscaling
- Canary rollouts
- Monitoring
- Prometheus
- Grafana drift dashboards
Results
- mAP@50, up from 0.76 for the YOLO baseline
- 0.89
- per frame
- 28 ms
- sustained with autoscaling
- 500+ rps
Stack
Python · PyTorch · RF-DETR · ONNX · TensorRT · Triton · Docker · Kubernetes · MLflow · GitHub Actions