The Average of e to the X Is Never One
Moving the average inside the exponential returns 1, and 1 happens to be the exact median and the exact geometric mean of e^X, which is why the mistake survives every re-check of the arithmetic. Completing the square in the exponent gives the true value e^(sigma squared over two), or 1.6487 at unit spread, because multiplying a Gaussian density by e^x slides its centre and scales its mass. Convexity settles the direction before any integral is set up, and on a heavy-tailed variable the quantity stops being finite at all.
Let be normal with mean zero and variance . What is the average of ?
The answer is , which at comes to . The answer almost everybody gives first is , arrived at by moving the average inside the exponential: the average of is zero, and . That is a sixty-five per cent error, and the interesting part is that is the exact answer to a closely related question, which is why the mistake is so durable.
What moving the average inside actually assumes
Written out, the reflex is a chain of two equalities: . The second is arithmetic. The first is the whole problem. Expectation commutes with a function for every non-degenerate precisely when is affine, because expectation is a linear operator and nothing more. The exponential is strictly convex, so the identity fails, and convexity even tells you the direction of the failure before a single integral is set up.
The integral, in full
Take first, where the algebra is cleanest. By definition,
Everything now hangs on the exponent. Complete the square in it:
Expand the right-hand side if you want to check it, and note that the identity holds for every real rather than approximately near the origin. Substituting it back pulls the constant out of the integral and leaves a Gaussian density with its centre moved from to and its width untouched. A density integrates to one, so the remaining integral needs no work at all:
The content of the calculation is a single geometric fact. Multiplying a Gaussian density by does not change its shape. It slides the centre to the right by one unit and multiplies the total mass by . That surplus mass is the answer.
Any spread, and the function that carries all of them
The same completion works at general , with the shift and the constant both scaling:
The centre now moves to instead of one, and the mass multiplier is . Replacing by throughout gives the moment generating function of a centred Gaussian, which is the whole family of these calculations in a single expression:
Setting recovers the result. Setting at gives , which is also the average of for of spread two, so the two dials are interchangeable in the way the formula suggests. For a normal variable with a non-zero mean the statement becomes , and the ratio between two spreads is , which is a sharper thing to verify numerically than any single value.
The direction was settled before the integral
If is strictly convex and is not almost surely constant, then whenever both sides are finite. With this says the average of strictly exceeds , so the correction to the reflex answer is always an addition and can never be a subtraction.
You can see the mechanism without the Gaussian. Replace by a coin flip that pays or . Its average is still zero, and the average of is , already above one. Stepping up multiplies by ; stepping down multiplies by . Those two moves do not cancel, and there is a one-line reason why: the arithmetic mean of a positive number and its reciprocal is at least their geometric mean, which is exactly , with equality only at .
Why one is such a durable wrong answer
The reflex answer is not noise. It is the exact median of , and also its exact geometric mean. Half the mass of a centred normal sits below zero, and the exponential is increasing, so half the mass of sits below for every value of whatsoever. The geometric mean is by the same triviality. So the candidate who answers has computed a real summary of the right distribution and reported it under the wrong name, which is why re-checking the arithmetic never rescues them.
Stated properly, is log-normal with parameters . Its mean is , its median is , and its mode is , which at is . Three separated summaries in that order are the signature of a right-skewed law: the long upper tail drags the mean away while leaving the median where it was.
This has a consequence anyone modelling prices runs into. If the log-return over some period is normal with mean zero and spread , the simple return has average , which at is per cent. A strategy whose log return averages nothing still averages a gain. Turn that around and you get the reason the term appears in the solution of geometric Brownian motion: it is exactly the drift you must subtract from the log to leave the level with no average growth.
Where the average stops existing
Every step above used one property of the normal law that is easy to take for granted. Its density decays like , which beats at every scale, so the integral converges however large gets. Remove the Gaussian tail and the answer can stop existing altogether.
Let follow a Student law with any number of degrees of freedom. Its density decays polynomially, grows faster than any polynomial decays, and is infinite. The same thing happens for a positive that is itself log-normal, which is the standard cautionary example: a log-normal variable has finite moments of every order, and yet its moment generating function fails to exist for any positive . Having all moments finite is weaker than having a moment generating function, and the log-normal is the example that separates the two.
At the other extreme the reflex answer becomes correct. Setting makes the constant zero, Jensen's inequality loses its strictness, and both answers equal one. At the true average is , which is why the habit of pushing the expectation inside can survive years of low-volatility work before it costs anything. The hidden hypothesis is not normality, it is integrability, and integrability is the one thing careful arithmetic on the wrong formula will never reveal.
Sources and further reading
- Johan Ludwig Jensen, “Sur les fonctions convexes et les inégalités entre les valeurs moyennes”, Acta Mathematica30 (1906), 175–193, where the inequality is proved in the generality used here.
- C. C. Heyde, “On a Property of the Lognormal Distribution”, Journal of the Royal Statistical Society B25 (1963), 392–393, on the log-normal law not being determined by its sequence of moments.
- Wikipedia: Log-normal distribution, Jensen's inequality and Moment-generating function.
The value was confirmed two ways before publication: by completing the square symbolically, and by high-order Gauss–Hermite quadrature at five spreads, which agrees to the last digit a double carries.
Comentarios · 0
Sé el primero en comentar.