Bayesian Versus orthodox statistics: which side are you on? - summary of an article by Dienes, 2011
Critical thinking
Article: Dienes, Z, 2011
Bayesian Versus orthodox statistics: which side are you on?
doi: 10.1177/1745691611406920
The contrast: orthodox versus Bayesian statistics
The orthodox logic of statistics, starts from the assumption that probabilities are long-run relative frequencies.
A long-run relative frequency requires an indefinitely large series of events that constitutes the collective probability of some property (q) occurring is then the proportion of events in the collective with property q.
- The probability applies to the whole collective, not to any one person.
- One person may belong to two different collectives that have different probabilities
- Long run relative frequencies do not apply to the truth of individual theories because theories are not collectives. They are just true or false.
- Thus, when using this approach to probability, the null hypothesis of no population difference between two particular conditions cannot be assigned a probability.
- Given both a theory and a decision procedure, one can determine a long-run relative frequency with which certain data might be obtained. We can symbolize this as P(data| theory and decision procedure).
The logic of Neyman Pearson (orthodox) statistics is to adopt decision procedures with known long-term error rates and then control those errors at acceptable levels.
- Alpha: the error rate for false positives, the significance level
- Beta: the error rate for false negatives
Thus, setting significance and power controls long-run error rates.
- An error rate can be calculated from the tail area of test statistics.
- An error rate can be adjusted for factors that affect long-run error rates
- These error rates apply to decision procedures, not to individual experiments.
- An individual experiment is a one-time event, so does not constitute a long-run set of events
- A decision procedure can in principle be considered to apply over a indefinite long-run number of experiments.
The probabilities of data given theory and theory given data
The probability of a theory being true given data can be symbolized as P(theory|data).
This is what orthodox statistics tell us.
One cannot infer one conditional probability just by knowing its inverse. (So P(data|theory) is unknown).
Bayesian statistics starts from the premise that we can assign degrees of plausibility to theories, and what we want our data to do is to tell us how to adjust these plausibilities.
- When we start from this assumption, there is no longer a need for the notion of significance, p value, or power.
- Instead, we simply determine the factor by which we should change the probability of different theories given the data.
The likelihood
In the Bayesian approach, applies to the truth of theories.
We can answer the questions about:
- p(H), the probability of a hypothesis being true (our prior probability)
- p(H|D), the probability of a hypothesis given the data (our posterior probability).
Neither of these can be do using the orthodox approach.
Likelihood: the probability of obtaining the exact data given the hypothesis.
Posterior is given by likelihood times prior.
The likelihood principle: all information relevant to inference contained in data is provided by the likelihood.
When we are determining how given data changes the relative probability of our different theories, it is only the likelihood that connects the prior to the posterior.
The likelihood is the probability of obtaining the exact data obtained given a hypothesis (P(D|H).
This is different from a p value, which is the probability of obtaining the same or more extreme data given both a hypothesis and a decision procedure.
- A p-value for a t test is a tail area of the t distribution
- The corresponding likelihood is the height of the distribution at the point representing the data
In orthodox statistics, p values are changed according to the decision procedure; under what conditions one would stop collecting data, whether or not the test is post hoc, how many other test one conducted.
None of these factor influence the likelihood.
The Bayes factor
The Bayes factor pits one theory against another.
Prior probabilities and prior odds can be entirely personal and subjective.
There is no reason why people should agree about these before data are collected if they are not part of the publically presented inferential procedure.
If the priors form part of the inferential procedure, they must be fairly produced and subjected to the tribunal of peer judgement.
One data are collected we can calculate the likelihood for each theory.
These likelihoods are things we want researchers to agree on. Any probabilities that contribute to them should be plausibly or simply determined by determined by the specification of the theories.
The Bayes factor (B): the ratio of likelihoods.
Posterior odds = B x prior odds.
- If B is greater than 1, the data supported you experimental hypothesis over the null.
- If B is less than 1, the data supported the null hypothesis over the experimental one.
- If B is about 1, the experiment was not sensitive.
The evidence is continuous and there are not thresholds in Bayesian theory.
B automatically gives a notion of sensitivity, it directly distinguishes data supporting the null from data uninformative about whether the null or you theory was supported.
For both p values associated with a t test and for B, if the null is false, as a number of subjects increases, then test scores are driven in one direction.
- p values are expected to become smaller
- Both t and B values are expected to become larger
When the null hypothesis is true, p values are not driven in any direction, only B us. B is then driven to zero.
Problems with the Neyman Pearson approach
Stopping rule
In the Neyman Pearson approach, one must specify the stopping rule in advance.
Once those conditions are met, there is to be no more data collection.
Typically, this means one should use a power calculation to plan in advance how many subjects to run.
The Bayes factor behaves differently from p values as more data are run (regardless of stopping rule).
- For a p value, if the null is true, any value in the interval 0 to 1 is equally likely no matter how much data you collect
- For this reason, sooner or later, you are guaranteed to get a significant result if you run subjects long enough and stop when you get the p value you want
- When the null is true, as the number of subjects increases, the p value is not driven to any particular value.
- As the number of subjects increases and the null is true, the Bayes factor is driven toward zero.
Planned versus post hoc comparisons
When using Neyman Pearson, it matters whether you formulated your hypothesis before or after looking at the data (post hoc vs. planned comparisons).
Predictions made before rather than after looking at the data are treated differently.
- Post hoc fitting can involve preference for one auxiliary over may others of at least equal plausibility.
In Bayesian inference, the evidence for a theory is just as strong regardless of its timing relative to the data.
This is because the likelihood is unaffected by the time the data were collected.
The likelihood principle follows from the axioms of probability.
It is not the ability to predict in advance per se that is important, that ability is just an (imperfect) indicator of the prior probability of relevant hypotheses.
When performing Bayesian inference, there is no need to adjust for the timing of predictions per se.
Multiple testing
When using Neyman Pearson, one must correct for how many tests are conducted in a family of tests.
When using Bayes, it does not matter how many other statistical hypotheses are investigated. All that matters is the data relevant to each hypothesis under investigation.
Once one takes into account the full context, the axioms

















