‹ Culture 3.14 · What we ask ourselves
Question · Decisions
When does a confidence metric create a certainty the system does not yet have?
A figure that is well calibrated in aggregate can be badly calibrated for exactly the case somebody has in front of them.
We used to publish a confidence figure alongside every prediction because it seemed the honest thing to do. We have changed our judgement.
The problem was not the aggregate calibration, which was reasonably good. It was the use. A number with decimal places conveys a precision the procedure does not have, and whoever receives it compares it with other numbers and decides on the basis of that comparison. Over time, a threshold becomes a business rule nobody approved.
Besides, a correct calibration in aggregate says nothing about the specific case. When the row in front of you bears little resemblance to what the model saw during training, the figure still comes out with the same typographic assurance.
We now show three things instead of one: a category with a defined operational meaning, an explicit warning when the case falls outside the usual distribution, and the model’s calibration in that segment. It takes up more space and it has changed whole conversations.
What we have not resolved is the uncertainty that comes from the data that is missing, which is the most frequent kind. No interval captures it, because it is not in the model: it is in what nobody recorded.