Skip to content
Questions

Confidence

Every answer comes with a sense of how sure it is. Here is how that number is computed and how to put it to work.

Why it matters

A bare pick cannot tell you when to trust it. A distribution can: a choice that put 0.93 on one option and a choice that split 0.40 / 0.38 / 0.22 both "pick" an option, but only the first should go straight into an automated workflow. Confidence condenses the distribution into one number between 0 and 1 so you can set a rule once and apply it everywhere.

Choice

How far the top probability stands above an even split, scaled to 0–1:

confidence = (p_max - 1/K) / (1 - 1/K)        K = number of options
Distribution (4 options)p_maxConfidence
0.25 / 0.25 / 0.25 / 0.250.250.00 (no idea)
0.55 / 0.25 / 0.15 / 0.050.550.40
0.88 / 0.08 / 0.03 / 0.010.880.84
1.00 / 0 / 0 / 01.001.00

A single-option choice always has confidence 1.

Score

How tightly the probability is packed around the most likely level, compared with a flat spread over the same scale:

mode       = the level with the highest probabilityspread     = Σ p_i · |i - mode|                 expected distance from the modeflat       = average distance of the levels from the middle of the scaleconfidence = max(0, 1 - spread / flat)          1.0 when there is one level

On a four-level scale, flat is 1.0. The urgency example puts 0.92 on level 3, 0.07 on level 2 and 0.01 on level 1: spread = 0.07 + 0.02 = 0.09, so confidence is 0.91. Probability split between two neighbouring levels lowers confidence less than probability split between the ends of the scale, which is what you want for an ordered scale.

Noul

A noul answer has no separate confidence field, because the probability already says it. If you want a number on the same 0–1 footing as choice, apply the choice formula with two options:

confidence = |2 · noul - 1|        0.5 → 0,  0.9 or 0.1 → 0.8

Choosing thresholds

  • Start from cost. Where a wrong automatic action is expensive (closing a ticket, rejecting an application), use a high bar such as 0.8 and send the rest to a person.
  • Measure on your own data. Take a few hundred cases where you know the right answer, run them, and plot accuracy against confidence. Pick the threshold where accuracy reaches what you need.
  • Re-check after you change the wording of a question or its options; confidence moves with the question, not only with the state.

For routing rules built on these numbers, see Patterns.

Loading the docs…