What the number is
A recogniser reads a plate one character at a time. For each position it produces a distribution over the candidate glyphs, and the reported confidence is derived from how decisively it picked each one. A high score means the model was not torn between an 8 and a B. It does not mean a human would agree with the result.
This distinction matters because the two most dangerous read errors are both high-confidence. A partially occluded plate can present a clean subset of characters that the model reads decisively and wrongly. A plate at a steep angle can compress into a different but perfectly legible string. In both cases the number beside the read is reassuring and incorrect.
There are three different confidences, not one
Corridor systems routinely collapse three separate measures into a single figure on screen. They should be reported separately because they fail independently:
- Character confidence — how sure the recogniser is of the plate string it produced.
- Match confidence — how close that string is to a watchlist entry. A single-character difference between a read and a BOLO plate is a very different situation from an exact match, and it should never be presented as the same event.
- Attribute confidence — colour, body type and make-model inference. This is the weakest of the three by a wide margin, and it should be treated as a filter over candidates rather than as an identification.
Where to set the threshold
Both directions of error have a cost, and they are not symmetrical for every use.
For traffic analytics — flow, journey time, origin-destination — a lower threshold is usually right. Individual misreads wash out in aggregate, and discarding marginal reads biases the sample toward the easy lanes and the slow vehicles.
For watchlist alerting the calculus reverses. Every false positive spends an operator's attention and, eventually, their trust in the channel. A BOLO channel that cries wolf gets muted, and a muted channel is worse than no channel because everyone believes it is working.
For anything that will be used as evidence, the threshold is not the control at all. The control is that a human looks at the frame. A read is a pointer to an image; the image is the thing that survives scrutiny.
Near misses need their own treatment
The reads just below threshold are not noise to be thrown away. On a serious matter they are exactly what an investigator wants: the vehicle that was probably there, on the corridor, in the window. Systems that silently discard sub-threshold reads destroy that, and the loss is invisible until someone asks for it.
Retain them, mark them clearly as unconfirmed, keep them out of automated alerting, and make them searchable by a role that understands what they are. The distinction that has to hold is between what the system asserts and what the system saw.
Measure your own rate, on your own road
Published accuracy figures describe a test set, not your corridor. The read rate you will actually get depends on plate standards in your jurisdiction, weather, speed, mounting geometry and how many plates are damaged, obscured or non-standard — and those vary enough between sites that a vendor figure is close to meaningless as a planning input.
Sample a few hundred reads per camera against the frames, by hand, at commissioning and then periodically. It is tedious and it is the only number worth quoting internally. It also tends to point straight back at placement, which is where most of the recoverable accuracy lives.