Lambdia

Counting Every McDonald's in America From the Two in Your Town

Two outlets in a town of fifty thousand is one per twenty-five thousand people, which scaled to the United States gives 13,600 against a published count near 13,500. That 0.74 per cent is luck, and the article shows why: the answer is exactly inversely proportional to the one density guess, and sweeping it across every defensible value spans 8,500 to 22,667. A second chain built from revenue, sharing no input at all, lands at 13,615.

How many McDonald's restaurants are there in the United States? You get no reference material and about thirty seconds. The point of the question is not the number. It is whether you reach for something you actually know or for something that sounds like a plausible size.

A hundred thousand sounds like a plausible size. It is roughly 7.4 times a published count of about 13,500 outlets, and it is worse than that, because it sits outside the entire range that any defensible assumption can produce. There is a better first move, and it is sitting outside your window.

One ratio, scaled

Think of a town you know well. Say it holds fifty thousand people and you can name two of these restaurants in it. That is the only genuinely new information in the whole problem, and it is a density:

ρ  =  50,000 people2 outlets  =  25,000 people per outlet\rho \;=\; \frac{50{,}000\ \text{people}}{2\ \text{outlets}} \;=\; 25{,}000\ \text{people per outlet}
(1)

Assume for a moment that the country resembles your town in this one respect. Then the outlet count is the population divided by that density:

N  =  Pρ  =  3.4×1082.5×104  =  13,600N \;=\; \frac{P}{\rho} \;=\; \frac{3.4 \times 10^{8}}{2.5 \times 10^{4}} \;=\; 13{,}600
(2)

Against a published figure near 13,500, that is 0.74 per cent off. Which is a piece of luck, and saying so out loud is the most useful thing you can do with it.

Why the agreement is luck

Equation (2) has exactly one soft input. The population of the United States is a number you either know or do not, and it is not in dispute. The density is a guess, and the answer is exactly inversely proportional to it:

Nρ  =  Pρ2dNN  =  dρρ\frac{\partial N}{\partial \rho} \;=\; -\frac{P}{\rho^{2}} \qquad\Longleftrightarrow\qquad \frac{\mathrm{d}N}{N} \;=\; -\frac{\mathrm{d}\rho}{\rho}
(3)

Double the density guess and the answer halves. There is nothing in the arithmetic that objects, no residual, no warning. So the error in the final figure is the error in that one guess, transmitted at full strength.

Sweeping the guess across every value a reasonable person might pick, one outlet per fifteen thousand people up to one per forty thousand, moves the answer from 22,667 down to 8,500. That is a spread of a factor of 2.67, and only one of those five densities lands within ten per cent of the published count. The one that does is the one in the story, which is precisely the point: the method delivers the decade, and the digits were a coincidence.

Fig. 1 — The whole sweep of the single soft input. Every bar is defensible, they span a factor of 2.67, and one of them happens to be nearly exact.

One tempting way to state the accuracy is that every point of the sweep gives the right power of ten. That claim is false, and it fails in an instructive way. The bottom of the sweep, 8,500, has already crossed below 10410^{4}, so it is a four-figure answer where the published count is five. The honest statement is a factor rather than a decade: every point of the sweep stays inside a factor of 1.7 of the published figure. An estimate that straddles a boundary between decades cannot claim decades.

A second chain that shares no input

A single chain of guesses is untestable, so build another one out of quantities that have nothing to do with towns or people. The chain publishes its systemwide United States sales, about 53.1 billion dollars, and average sales per outlet run about 3.9 million dollars. Dividing gives

N  =  5.31×10103.9×106    13,615N \;=\; \frac{5.31 \times 10^{10}}{3.9 \times 10^{6}} \;\approx\; 13{,}615
(4)

which agrees with equation (2) to 0.11 per cent. Neither the town nor the population appears anywhere in equation (4), and neither revenue figure appears anywhere in equation (2).

Fig. 2 — Two chains with no shared input. Their agreement is evidence; a single chain agreeing with itself would not have been.
What independence means here

Two estimation chains are independent when no quantity feeds both of them. That is a statement about inputs, not about publishers. Both figures above ultimately trace back to the same company, and that is fine: what would destroy the cross-check is a shared assumption, because then the two routes would be one route wearing two coats.

What the method actually buys

The defensible claim is a factor, not a figure. Two routes with disjoint inputs both land near fourteen thousand, the plausible range around each is a factor of two or so wide, and a blurted hundred thousand sits outside all of it. That is enough to reject the reflex answer and enough to sanity-check a number someone else hands you, which is what the question is really testing.

The temptation is to keep the 0.74 per cent, because it is a flattering number. Keeping it means presenting an estimate as a measurement, and an interviewer who understands the method will ask what happens if the town is atypical. The answer to that question is equation (3), and it is not comfortable.

Where the scaling breaks

Outlets are not sited by population. They are sited by traffic, and traffic correlates with population only loosely. A small town on a motorway junction carries far more restaurants per resident than a dense residential district of the same size, so the ratio you measured is one draw from a skewed distribution, and the error in equation (2) is dominated by that sampling rather than by anything in the arithmetic. Anchoring on a town you know is still better than guessing, because you know which way your town is unusual.

The revenue route has its own failure mode. It works here because a franchise chain runs a standardised format, so sales per outlet cluster tightly. Try the same division on hotels or on car dealerships, where revenue per site spans two decades, and the average carries almost no information about the count.

Finally, the target itself moves. The published United States count has run in the low thirteen thousands across recent years, so any claim tighter than a few per cent is a claim about a particular year rather than about the country. That is one more reason to state the answer as a bracket.

Sources and further reading

Nothing here asserts a true number of outlets. What was checked, in exact rational arithmetic so there is no rounding anywhere, is the arithmetic of both chains, that they share no input, the exact inverse proportionality of equation (3), and the full sweep: 8,500 to 22,667, a factor of 2.67, with one of the five densities inside ten per cent of the published count.

Comments · 0

Be the first to comment.