Define latency before you measure it
Two numbers are often mixed up. Throughput is how many frames per second the system processes. Latency is how long it takes from something happening in front of a camera to your system reacting. A pipeline can have healthy throughput and still alert late.
For safety alerts such as a missing helmet or a person in an exclusion zone, latency is usually the number that matters. Agree a target with the people who will act on the alert.
The stages to budget
End-to-end latency is the sum of every stage. Give each one its own line in the budget:
- Camera and network: encoder settings, network jitter and the buffer your RTSP client keeps.
- Decode: Jetson devices have dedicated hardware video decoders, which frees the GPU for inference.
- Batching: in DeepStream, frames from several streams are gathered into a batch before inference, so a stream may wait for the batch to fill.
- Pre-processing and inference: the TensorRT engine, at FP16 or INT8 precision.
- Post-processing and tracking: decoding detections, non-maximum suppression, object tracking and rule checks such as zone entry.
- Output: the alert, message or database write.
Batching trades waiting time for throughput
Larger batches use the GPU more efficiently, which lets one device serve more cameras. The cost is waiting: the first frame in a batch waits for the others, up to a timeout you set. If your alert target is tight, prefer a smaller batch or a shorter timeout, and accept fewer cameras per device.
Precision and calibration
INT8 inference is usually faster and lighter than FP16, but it needs a calibration step. Calibrate with footage from the real cameras (lighting, angles, weather) rather than a generic dataset, and check accuracy again on your own data after quantisation. Quantisation can cost some accuracy, and how much depends on the model and the scene.
Measure on the real device, with real streams
- Use live RTSP feeds, not a video file that plays with perfect timing.
- Set the Jetson power mode and clocks you will use in production, because results change with them.
- Test inside the real enclosure and ambient temperature, because thermal throttling lowers speed over hours.
- Report percentiles such as p95 and p99, not only the average, and run long enough to see spikes.
Set the target from the use case
A dashboard that counts people can tolerate seconds. An intrusion alert cannot. Write the latency target, the number of cameras and the frame rate into the scope before building. In our projects this is settled in the first 48 hours of discovery, when we define the latency SLA.
