One Head Tilts a Flat Prior to 2p, and the Average Bias to Two Thirds
A flat belief about a coin's bias is an input to the calculation, not a conclusion of it, and a single head does not leave it standing. The density tilts to 2p, the cumulative law becomes p squared, the average bias moves to 2/3, and the old answer of one half is demoted to the lower quartile. The general update is the Beta conjugate family, which sends 750 heads in 1000 to Beta(751, 251) with mean 0.749501.
A coin comes off a badly run production line. Its bias, meaning the probability that it lands heads, is equally likely to be any number between and . You toss it once. It lands heads. What do you now believe about ?
Most people answer that nothing has changed, or that the coin is presumably fair, which is the same answer said twice. One toss carries almost no information, the reasoning goes. The instinct about the quantity of information is right and the conclusion is wrong. After that single head the average value of is , and the shape of the belief has changed from a flat line into a straight ramp.
Where the answer one half legitimately lives
The reflex has a defensible version, and locating it is worth more than dismissing it. Had the problem stated that the coin was fair, one head would change nothing at all: a fair coin stays fair whatever it does. The problem says something else. It says the bias is unknown and flat across the unit interval, and a flat description of ignorance is an input to the calculation rather than a conclusion of it. Inputs do not survive observations.
There is exactly one place where belongs, and it is the denominator we are about to write down. It is also worth knowing what becomes of that number afterwards. Once the head has been observed, is the lower quartile of the belief about , so three quarters of the remaining credence sits above it. The trap answer does not simply move. It is demoted to the twenty-fifth percentile.
Bayes in continuous form, done by hand
Write for the density of before the toss. Uniformity means on , which integrates to one, so every value really is on an equal footing. The probability of the observation, given a particular bias , is itself. That is the entire likelihood. Bayes' rule in its density form then reads:
The denominator is the unconditional probability of seeing a head before you know anything, and with a flat prior it equals . Dividing by it forces the constant rather than leaving it to taste:
The mechanism is worth saying in words. Each candidate bias is reweighted in proportion to how readily it would have produced the head you actually saw. A bias of delivers a head nine times as often as a bias of , so it ends up nine times as credible. Renormalising that reweighted curve turns the horizontal line into a line of slope two.
The whole law, not only its average
Integrating the density gives the cumulative law, and from there every quantile follows:
Solving puts the median at , and the quartiles are exactly and . Note that the mean and the median disagree, which they must, because the density is skewed. When the answer to this question is quoted as it is the average bias being quoted, not the typical one. And since , the probability that the coin actually favours heads is after a single toss.
From one toss to a thousand
Nothing above depended on the toss count being one. With heads in independent tosses the likelihood is , and reweighting the flat prior by it produces a density proportional to the same expression. That is a Beta law:
Set and the mean returns . The single-toss answer is the general answer evaluated at a small argument, not a separate construction, and that is the cleanest reason to trust it.
A belief about a success probability, updated on successes and failures, returns a belief with mean . The family is closed under the update, which is what conjugate means. The uniform prior is , so one head sends it to , whose density is .
Now run the same machine on real data. Seven hundred and fifty heads in a thousand tosses: equation (4) gives , with mean and standard deviation . Its mode is exactly, the observed frequency. The cumulative law climbs from to , so essentially all of the belief lives inside a window one tenth of a unit wide.
Here a wording error is easy to make and worth naming, because it is the kind that survives editing. It is the cumulative law that is nearly a step at . The density is a narrow spike, and a spike is not a step. The two objects live on different vertical scales, the density peaking near while the cumulative law is capped at one by construction, and calling either one by the other name gives away that the distinction was never noticed.
What you should pay for the next toss
The distribution is the answer to the question asked, but the question a trader would ask next is what to bet on toss number two. Conditioning on the bias and averaging over it gives the predictive probability, and the tower property collapses the whole calculation to the posterior mean:
This is the rule of succession that Laplace wrote down in 1774. After one head it says , the same number as before. After 750 heads in a thousand it says rather than . The rule never quite reaches the observed frequency because the prior keeps pulling toward the middle, by an amount of order that dies out but never vanishes.
Where this stops working
Three hypotheses carry the whole derivation, and each fails in a way worth recognising.
The flat prior is a choice and not the absence of one. Jeffreys' prior for a Bernoulli parameter is , which is the one invariant under reparameterisation and therefore has a better claim to being uninformative. Update it on one head and you land on , whose mean is . Two defensible descriptions of ignorance, two different answers from the same single observation. The disagreement is fatal at and invisible by , where the Jeffreys posterior mean is against . Data washes out the prior, and how much data it takes depends on how much the two priors disagree.
Independence given is doing real work. It is what turns the likelihood into a product and hands you the Beta form. A coin whose surface wears with use, or a tosser who learns, breaks the product and everything downstream of it.
The concentration statement needs the true bias to be interior. Suppose the coin has heads on both faces, so . After heads the belief is with mean , which crawls toward one from below without arriving, and whose spread shrinks like rather than the usual . The comfortable picture of a symmetric spike with a two-sigma interval around it belongs to interior truths. At the boundary the spike is one-sided and the interval runs off the end of the parameter space, where it means nothing.
Sources and further reading
- Pierre-Simon Laplace, Mémoire sur la probabilité des causes par les événements (1774), where the rule of succession first appears.
- Harold Jeffreys, “An Invariant Form for the Prior Probability in Estimation Problems”, Proceedings of the Royal Society A186 (1946), 453–461.
- Wikipedia: Bayes' theorem, Beta distribution and Rule of succession.
Every quantity above was checked twice before publication: symbolically, by integrating the densities, and numerically, by a rejection sampler that reproduces the law without ever using the closed form.
Comentarios · 0
Sé el primero en comentar.