Drawing strong conclusions from small samples — treating a handful of observations as representative of a population or a pattern. The error is not in using small samples: in intelligence they are often all that exists. It is in failing to lower confidence accordingly.
In intelligence analysis
- Two or three incidents are described as a campaign; one or two intrusions become an actor's "characteristic" tradecraft.
- A pattern is asserted from a sample that cannot distinguish pattern from coincidence, and the reader is given no way to tell the difference.
- Small samples are systematically over-read where collection is genuinely thin — which is exactly where confidence should be lowest (see Absence of Evidence Fallacy).
- In cyber reporting the same weakness presents as "we have seen this TTP twice" generalised into an actor profile, and as prevalence claims from a single telemetry source.
Countermeasure
- State the n behind every pattern claim and let the reader see it.
- Where the sample is small, express the judgement as a possibility and name the evidence that would move it (Estimative Language).
Related
- Cognitive Bias — the taxonomy this bias sits within
- Base-Rate Neglect — the same failure applied to the prior rather than the evidence
- Overconfidence Bias — why the small sample produces confident prose
- Estimative Language — the mechanism for pricing small samples honestly