TrackingDesk

Glossary

Statistical significance

A statement about how likely a difference this large would be if there were no real effect. It is not a measure of how large, how important, or how certain the effect is.

Also called: significance, p-value

You ran a holdout test and the exposed group did better. Significance asks whether a gap that size would be surprising if the advertising had done nothing at all.

What it does not tell you. Not how big the effect is. Not whether it is worth acting on. Not the probability that your conclusion is correct — that last one is the most common misreading, and it is simply not what the calculation says.

A trivially small difference becomes significant with a large enough sample. A commercially important difference can miss significance in a small one. Significance and importance are separate questions and a result needs both answered.

Three ways teams fool themselves.

Peeking. Watching a running test and stopping when it crosses the threshold guarantees crossing it eventually, whether or not an effect exists. Decide the duration first and let it run.

Testing many things. Check enough variants, segments and metrics and some will look significant by chance alone. The more you look, the more you find that is not there.

Accepting “not significant” as “no effect”. It usually means the test could not resolve a difference at that sample size. Absence of evidence, not evidence of absence — the honest report is “we could not tell”, which is a real and useful answer.

The practical stance: significance is a filter against acting on noise, not a certificate. Decide the size of effect that would change your decision before you run the test, and check whether the result clears that bar as well as the statistical one.

Do not confuse with

Close enough to get mixed up, different enough that the mix-up costs something.