> For the complete documentation index, see [llms.txt](https://laurence-wilse-samson.gitbook.io/textbooks/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://laurence-wilse-samson.gitbook.io/textbooks/financial-economics-claims-prices-holders/part-ii-asset-pricing/chapter_05_equity_premium.md).

# Chapter 5: Consumption, Risk Premia, and the Equity Premium Puzzle

*Part II: Asset Pricing — Financial Economics: Claims, Prices, and Holders*

***

## Opening Episode: A Result Nobody Wanted

Rajnish Mehra and Edward Prescott had a paper by the end of the 1970s. It did not appear until 1985.

The exercise they set themselves was modest and, on its face, routine. Chapter 3 §3.5 ended with the statement that the stochastic discount factor is marginal utility growth, $$m = \delta u'(c\_1)/u'(c\_0)$$. That is a testable proposition, because consumption is measured. The national accounts report what Americans spent on nondurables and services every year going back to the nineteenth century, and the stock market reports what equities paid. Write down the simplest general-equilibrium economy consistent with the theory — one representative household, an endowment of consumption arriving each period, a claim to that endowment traded at a price — calibrate the household's consumption process to the observed data, and see what risk premium the model generates. Mehra and Prescott used ninety years, 1889 to 1978, and a two-state Markov process for consumption growth matched to its mean, its variance, and its first-order autocorrelation. The one free preference parameter was the coefficient of relative risk aversion, $$\gamma$$, which they allowed to run anywhere from zero to ten. Ten was already generous: a household with $$\gamma = 10$$ will pay a substantial share of its wealth to avoid a coin flip over modest stakes.

The economy would not produce the premium. Across the whole admissible parameter space, the largest equity premium the model could generate was about a third of a percentage point. The data said just over six.

An order of magnitude is not a calibration disagreement. It is the kind of gap that ordinarily means someone has made a mistake, and that is what almost everyone told them. Mehra's retrospective accounts of the episode describe seminars in which the discussion consisted of suggestions about where the error might be — the consumption series, the deflator, the return series, the Markov chain, the code. Referees said the same thing in writing. An earlier version of the paper was framed as a test of the intertemporal asset-pricing model and rejected as such; the reframing that eventually got it published was the decision to stop apologizing for the result and name it. The published title is four words long and one of them does the work: *The equity premium: A puzzle*.

That decision is the reason the paper matters, and it is why this chapter opens with it rather than with an equation. A rejected model is a dead end; a puzzle is a research program. By calling the gap a puzzle, Mehra and Prescott asserted that the theory was too useful to abandon and too wrong to keep as it stood, and they handed the profession a number to beat. Four decades of asset pricing have been organized around that number. Habit formation, long-run risk, rare disasters, recursive preferences, incomplete markets, limited participation, and intermediary asset pricing are all, in the first instance, attempts to write down an $$m$$ volatile enough to explain a six-percent premium without implying things about interest rates and consumption that are obviously false.

Two lessons are worth stating before the algebra begins. The first is that the puzzle is a *quantitative* failure of a *qualitatively correct* framework. Nothing in Chapter 3 breaks. The pricing equation still holds; assets that pay off in bad states are still expensive. What fails is the claim that aggregate consumption is a good enough measure of "bad states" to price equity. The second is that a model's most informative output is often the size of its residual. Mehra and Prescott's contribution was not a theory. It was a well-measured discrepancy, stated precisely enough that everyone afterwards had to answer to it.

***

## 5.1 The Consumption-Based Model

Chapter 3 §3.5 derived $$m = \delta u'(c\_1)/u'(c\_0)$$ from a household's first-order condition and then set it aside, treating $$m$$ as an object recovered from prices. This chapter takes the derivation literally and asks what happens when the $$c$$ in that expression is the consumption series in the national accounts.

Write $$c\_t$$ for real consumption per capita at date $$t$$ and

$$
g\_{t+1} \equiv \ln\left(\frac{c\_{t+1}}{c\_t}\right)
$$

for log consumption growth. The household maximizes $$u(c\_t) + \delta E\_t\[u(c\_{t+1})]$$ subject to a budget constraint, and the Euler equation for any traded asset with gross return $$R\_{t+1}$$ is

$$
1 = E\_t\left\[\delta\frac{u'(c\_{t+1})}{u'(c\_t)}R\_{t+1}\right]
$$

The workhorse specification of $$u$$ is **power utility** (constant relative risk aversion),

$$
u(c) = \frac{c^{1-\gamma} - 1}{1 - \gamma}, \qquad u'(c) = c^{-\gamma}
$$

with the case $$\gamma = 1$$ read as $$\ln c$$. Power utility is used for three reasons and only one of them is empirical. It makes risk aversion independent of the level of wealth, which is what a century of growth without a trend in risk premia seems to require. It makes portfolio shares independent of wealth, so a scale-free economy has a balanced growth path. And it has one parameter. The stochastic discount factor becomes

$$
m\_{t+1} = \delta\left(\frac{c\_{t+1}}{c\_t}\right)^{-\gamma} = \delta e^{-\gamma g\_{t+1}}
$$

Everything in this chapter follows from that one line, so read it slowly. The discount factor is high exactly when consumption growth is low, and $$\gamma$$ governs *how much* higher. A recession in which consumption falls two percent raises $$m$$ by roughly $$2\gamma$$ percent. With $$\gamma = 2$$ that is a four-percent movement in the price of a state-contingent dollar; with $$\gamma = 50$$ it is a hundred-percent movement. The single parameter $$\gamma$$ is doing double duty — it is the curvature of the utility function, and it is the amplifier that converts small consumption movements into large movements in state prices. The puzzle, in one sentence, is that the amplifier has to be turned up impossibly far.

Substituting into $$1 = E\[mR]$$ gives the two equations the rest of the chapter works with. For the riskless asset, whose return is known at $$t$$,

$$
R\_{f,t+1} = \frac{1}{E\_t\left\[\delta e^{-\gamma g\_{t+1}}\right]}
$$

and for equity, with gross return $$R\_{t+1}$$,

$$
1 = E\_t\left\[\delta e^{-\gamma g\_{t+1}} R\_{t+1}\right]
$$

Chapter 3 §3.5's beta representation, $$E\[R] - R\_f = -R\_f\mathrm{Cov}(m, R)$$, now has content, because $$m$$ is a function of measured data. An asset is expensive — earns a low expected return — if it pays off when consumption growth is low. An asset that pays off in booms is a bet on more of what you already have, and the market will not pay much for it. Equity is the second kind of claim, so it should earn a premium. The question is how large a premium a given amount of covariance with consumption can justify.

Two remarks on what this model is and is not. It is not an alternative to the CAPM; it is the more primitive statement of which the CAPM is a special case. Chapter 4 §4.6 showed the CAPM to be the restriction $$m = a - b R\_M$$, which asserts that the only thing making a state bad is a low market return. The consumption model asserts that the only thing making a state bad is low consumption, which is closer to what "bad" means and has the advantage that the conditioning variable is not itself an asset price. And it is not a partial-equilibrium device: in a representative-agent economy consumption equals output, so the model ties asset prices to the macroeconomy directly. That is the ambition, and it is what makes the failure interesting rather than merely technical.

***

## 5.2 The Lognormal Two-Equation System

To get numbers out of §5.1 we need a distributional assumption, and the standard one is that log consumption growth and log returns are jointly normal and identically distributed over time. Write $$\mu\_g = E\[g]$$, $$\sigma\_g = \sigma(g)$$, and let $$r = \ln R$$.

The derivation is three lines. For a lognormal variable, $$\ln E\[e^y] = E\[y] + \tfrac{1}{2}\mathrm{Var}(y)$$. Apply that to $$1 = E\[mR]$$, taking logs of both sides:

$$
0 = E\[\ln m] + E\[r] + \tfrac{1}{2}\mathrm{Var}(\ln m) + \tfrac{1}{2}\mathrm{Var}(r) + \mathrm{Cov}(\ln m, r)
$$

Apply it first to the riskless asset, for which $$r = r\_f$$ is a constant, so the last two terms vanish. Since $$\ln m = \ln\delta - \gamma g$$, we have $$E\[\ln m] = \ln\delta - \gamma\mu\_g$$ and $$\mathrm{Var}(\ln m) = \gamma^2\sigma\_g^2$$, giving the **risk-free rate equation**

$$
r\_f = -\ln\delta + \gamma\mu\_g - \tfrac{1}{2}\gamma^2\sigma\_g^2
$$

Now subtract the riskless case from the general one and use $$\mathrm{Cov}(\ln m, r) = -\gamma\mathrm{Cov}(g, r)$$:

$$
E\[r] - r\_f + \tfrac{1}{2}\mathrm{Var}(r) = \gamma\mathrm{Cov}(g, r)
$$

The left-hand side is the log premium plus a Jensen correction, which together are approximately the arithmetic premium. So the **premium equation** is

$$
E\[R] - R\_f \approx \gamma\mathrm{Cov}(g, r) = \gamma\rho\_{g,r}\sigma\_g\sigma\_r
$$

These two equations are the whole of the empirical content, and each has a clean reading.

The risk-free rate equation says the safe rate is determined by three forces. Impatience ($$-\ln\delta$$) pushes it up: impatient people must be paid to save. Expected growth ($$\gamma\mu\_g$$) pushes it up, and *scaled by* $$\gamma$$, because a household that expects to be richer tomorrow wants to borrow against that, and the more curved its utility the more sharply it wants consumption smoothed across time. Uncertainty ($$-\tfrac{1}{2}\gamma^2\sigma\_g^2$$) pushes it down, through precautionary saving. Note that $$\gamma$$ here is playing a third role, on top of the two in §5.1: with power utility, the willingness to substitute consumption *across time* is $$1/\gamma$$, the reciprocal of risk aversion. That forced identification is the seam that Epstein-Zin preferences cut open in §5.6, and it is the reason the risk-free rate equation turns into a second puzzle in §5.3.

The premium equation says something sharper. **Only the covariance of returns with consumption growth is priced, and the price of that covariance is** $$\gamma$$**.** Not the variance of returns; not the total risk of the asset. This is Chapter 3 §3.5's beta representation with a name attached to the factor. Set the covariance to zero and the premium is zero regardless of how volatile the asset is — a claim to a fair coin flip, however large the stakes, is worth its expected value. The premium equation is therefore a constraint linking three measurable quantities and one free parameter, and §5.3 measures the three.

***

## 5.3 The Puzzle, Quantitatively

Start with Mehra and Prescott's own arithmetic, which is where the puzzle got its numbers.

**Table 5.1: US moments, 1889-1978, as reported by Mehra and Prescott**

| Quantity                                        | Value  |
| ----------------------------------------------- | ------ |
| Mean real return on the S\&P index              | 6.98%  |
| Mean real return on the short riskless security | 0.80%  |
| Equity premium                                  | 6.18%  |
| Standard deviation of the real equity return    | 16.54% |
| Mean growth of real per-capita consumption      | 1.83%  |
| Standard deviation of consumption growth        | 3.57%  |

*Source: Mehra and Prescott (1985), summary statistics as reported by the authors; annual real figures for 1889-1978.*

Two features of the table matter more than the levels. Consumption growth is **smooth**: a standard deviation of three and a half percent over a sample containing the Great Depression, and something closer to one to one and a half percent in postwar data, against sixteen to twenty percent for equity returns. And consumption growth is **weakly correlated** with equity returns: annual correlations in US data are conventionally quoted in the range of 0.1 to 0.3, but the estimate is unusually sensitive to how the year is cut. On 1948 to 2026 data, matching calendar-year averages gives about 0.07, fourth quarter to fourth quarter about 0.42, and first quarter to first quarter about −0.20 — so the sign itself is a timing convention before it is a fact, and the choice between nondurables and services or total consumption moves it again. Aggregate consumption barely moves, and what movement there is only loosely tracks the stock market.

![Figure 5.3: The equity premium by decade](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-4bc7cba3f3953fda1f5a0737ba51af62e344c0c1%2Ffig_05_03_the_equity_premium_by_decade.png?alt=media)

**Figure 5.3: The equity premium by decade.** Panel (a) is the realized real equity premium over long bonds, decade by decade since 1872, from Shiller's long-run total-return series; panel (b) is the same premium divided by its own standard deviation within the decade. The premium is over *bonds* rather than bills, because Shiller's file carries a long-bond total return and no short riskless rate before 1934; measured against bills the premium would be larger, which strengthens rather than weakens what follows. Two readings. The dashed line is the full-sample mean, 5.9 percent a year, and the shaded band is two standard errors around it — 3.0 to 8.8 percent. That band is the honest answer to the question a reader should ask of Table 5.1's 6.18: a century and a half of annual data pins the premium to within about three points either way. And the decade means run from −5 in the 2000s to +19 in the 1950s, a spread of twenty-four points across periods in which the underlying economy did not change by anything like that much. Neither reading rescues the model. A premium of three percent still requires a risk aversion in the dozens once it is divided by the covariance Figure 5.4 measures; the puzzle is not a puzzle about the second decimal place. *Source: Robert Shiller's ie\_data.xls, real total-return series for the S\&P composite and for long US government bonds. Author's calculations.*

![Figure 5.4: Consumption growth and equity returns](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-b4b86b9a70cfa933c54a41fadb1985c816d8aa2a%2Ffig_05_04_consumption_growth_and_equity_returns.png?alt=media)

**Figure 5.4: Consumption growth and equity returns.** Annual real per-capita consumption of nondurables and services against the annual real total return on equities, 1948 to the present, with the fitted line. This is the denominator of the premium equation, and the figure's point is that the cloud is round: consumption growth has a standard deviation of about 1.7 percent against roughly 18 for the equity return, and the two barely move together. The inset reports the covariance under three timing conventions, because the correlation is unusually sensitive to which twelve months of consumption are matched to which twelve months of return — the annual-average convention a reader replicating from FRED would use gives +0.07, a fourth-quarter-to-fourth-quarter convention gives +0.42, and a first-quarter convention gives −0.20. What survives all three is the magnitude. The implied risk aversion is 278 on the first convention and 45 on the second, and undefined on the third because the covariance has the wrong sign. There is no timing convention on which aggregate consumption covaries with the stock market enough to price a six-percent premium at a plausible risk aversion, and that — not any particular correlation — is the puzzle. *Source: Bureau of Economic Analysis and Bureau of Labor Statistics via FRED (PCND, PCESV, CPIAUCSL, and mid-period population), with Shiller's ie\_data.xls for the equity return. Author's calculations.*

Now run the premium equation backwards. Take a six-percent premium and a twenty-percent standard deviation of equity returns — round figures consistent with Shiller's long-run US series and with Table 5.1 — and ask what $$\gamma$$ is needed at various assumptions about the smoothness and comovement of consumption. Then feed that $$\gamma$$ into the risk-free rate equation, with a two-percent mean growth rate and, generously, no impatience at all ($$\delta = 1$$), and see what safe rate the same household implies.

**Table 5.2: Implied risk aversion, and the riskless rate it produces**

| $$\sigma\_g$$ | $$\rho\_{g,r}$$ | $$\mathrm{Cov}(g,r)$$ | Implied $$\gamma$$ | Implied $$r\_f$$ at that $$\gamma$$ |
| ------------- | --------------- | --------------------- | ------------------ | ----------------------------------- |
| 1.5%          | 0.10            | 0.00030               | 200                | −50.0%                              |
| 1.5%          | 0.20            | 0.00060               | 100                | 87.5%                               |
| 1.5%          | 0.50            | 0.00150               | 40                 | 62.0%                               |
| 1.5%          | 1.00            | 0.00300               | 20                 | 35.5%                               |
| 3.57%         | 0.20            | 0.00143               | 42                 | −28.5%                              |
| 3.57%         | 1.00            | 0.00714               | 8.4                | 12.3%                               |

*Source: Author's calculation from the equations of §5.2, with an equity premium of 6 percent, a return standard deviation of 20 percent, mean consumption growth of 2 percent, and a time discount factor of 1.*

Read the fourth column first. Under postwar-like smoothness and a correlation of 0.2, the premium equation demands a coefficient of relative risk aversion of one hundred. Even granting the model a *perfect* correlation between consumption growth and equity returns — an assumption the data reject flatly, and the most generous one available — it still needs twenty. Mehra and Prescott's cap of ten was not a rhetorical device; it corresponds to a household that would pay a large fraction of its wealth to avoid a gamble it could self-insure. A household with $$\gamma = 100$$ is one that would refuse almost any actuarially favorable bet on any stake, and whose implied behavior toward small risks is not recognizable as human.

Now read the fifth column, which is Weil's (1989) contribution and the reason the puzzle has two halves. Fixing the premium by raising $$\gamma$$ does not leave everything else alone. Raising $$\gamma$$ raises the household's desire to smooth consumption *over time*, and since consumption has grown at roughly two percent a year, a household with a strong smoothing motive wants desperately to borrow from its richer future self. In equilibrium nobody can, so the interest rate must rise until it stops wanting to. At $$\gamma = 20$$ the model's riskless rate is 35.5 percent. At $$\gamma = 40$$ it is 62 percent. The realized US real short rate over the last century is on the order of one percent. This is the **risk-free rate puzzle**: the parameter that fixes the premium destroys the interest rate.

There is a mechanical curiosity in the negative entries worth understanding rather than glossing. The two $$\gamma$$ terms in the risk-free rate equation pull in opposite directions, and the precautionary term is quadratic, so the implied $$r\_f$$ rises with $$\gamma$$, peaks at $$\gamma = \mu\_g/\sigma\_g^2$$, and then falls, going negative for large enough $$\gamma$$. With $$\sigma\_g = 1.5$$ percent the peak sits at $$\gamma \approx 89$$ and reaches about 89 percent, so the $$\gamma = 200$$ row is on its far side; with $$\sigma\_g = 3.57$$ percent the peak sits at $$\gamma \approx 16$$, which is why the $$\gamma = 42$$ row is negative too. This does not rescue the model. It says that the only way power utility can deliver a low safe rate alongside a huge premium is by making precautionary saving enormous, which requires consumption risk the data do not contain. Figure 5.1 draws Table 5.2's two columns on a shared $$\gamma$$ axis, hump and all.

![Figure 5.1: The premium against gamma, and the rate it implies](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-d3b31b444a61eca3e0946d06ec80f9fc5ddde09f%2Ffig_05_01_premium_against_gamma.png?alt=media)

**Figure 5.1: The premium against gamma, and the rate it implies.** Table 5.2 drawn, on a shared $$\gamma$$ axis, at that table's parameters: a six percent premium, twenty percent equity volatility, two percent mean consumption growth, and no impatience ($$\delta = 1$$). Panel (a): the premium power utility delivers, $$\gamma\rho\_{g,r}\sigma\_g\sigma\_r$$, against the six percent to be explained — even a perfect correlation between consumption growth and equity returns needs $$\gamma \approx 20$$, and a postwar-like correlation of 0.2 needs 100. Panel (b): the riskless rate that same $$\gamma$$ implies, $$\gamma\mu\_g - \tfrac{1}{2}\gamma^2\sigma\_g^2$$ — 35.5 percent at $$\gamma = 20$$ and 62 percent at $$\gamma = 40$$, against a realized US real short rate on the order of one percent — Weil's (1989) half of the puzzle. The parameter that fixes the premium destroys the interest rate. In panel (a) a chevron on the top frame marks a ray that leaves the panel rather than ending there. *Source: Author's calculation from the two equations of Section 5.2, at Table 5.2's parameters.*

One more way to state the failure, because it removes the last escape route. Suppose we abandon $$\delta \le 1$$ and simply solve for the discount factor that reconciles $$\gamma = 20$$ with a one-percent riskless rate. The answer is $$\delta = 1.41$$. The household must value consumption next year forty-one percent *more* than consumption this year — negative time preference at a rate of about twenty-nine percent a year. At $$\gamma = 100$$ the required $$\delta$$ is 2.38. The model can always be forced to fit two moments with two free parameters. What it cannot do is fit them with parameters anyone is willing to defend.

> **Box 5.1 — How the national accounts measure consumption**
>
> Every calibration in this chapter divides the equity premium by a covariance with consumption growth, and the consumption series is a construct with choices in it. Four of those choices reach the answer.
>
> **What is counted.** The theory wants a flow of services consumed. The accounts publish personal consumption expenditure, which includes durable goods — a car bought in one quarter is spent in that quarter and consumed over a decade. That is why the literature uses nondurables and services rather than total consumption, and it means Table 5.1's smoothness is partly a definitional choice: the excluded component is the volatile one.
>
> **What is imputed rather than observed.** A substantial share of measured consumption is not a transaction anybody made. Owner-occupied housing enters as imputed rent on a house the household already owns; some financial services enter through the margin between interest rates rather than a billed fee; employer-paid insurance enters through the employer's payment. None of these is a household deciding at a price, and all of them are inside the series whose covariance with returns the model is about.
>
> **The timing convention.** Consumption is a flow over a period; a return is a point-to-point change. A quarterly consumption figure is an average across three months, so its growth rate is a difference of averages, which is not the object an Euler equation refers to. Figure 5.4 shows how much this matters: the correlation between annual consumption growth and the annual real equity return moves from +0.07 to +0.42 to −0.20 depending on nothing but which twelve months of consumption are matched to which twelve months of return.
>
> **Revisions.** The series is revised, sometimes substantially, in annual and comprehensive updates. A calibration run today on 1889-1978 is not run on the numbers Mehra and Prescott ran it on.
>
> The puzzle survives all four, which is why they are worth stating rather than hiding. No timing convention, no exclusion and no vintage produces a covariance large enough to price a six-percent premium at a plausible risk aversion. But a reader who takes the second decimal place of any published estimate seriously has not read the source note.

***

## 5.4 ★ Hansen-Jagannathan Bounds

*Starred. This section states the general form of the constraint that §5.3 illustrated with one utility function. A reader who skips it loses the diagnostic, not the argument.*

The arithmetic of §5.3 is specific to power utility and lognormality. Hansen and Jagannathan (1991) showed that the essential difficulty survives without either assumption, and their statement of it is the standard diagnostic in this literature.

Take the pricing equation for an excess return $$R^e$$ — the difference between two returns, so it costs nothing today:

$$
0 = E\[m R^e] = E\[m]E\[R^e] + \mathrm{Cov}(m, R^e) = E\[m]E\[R^e] + \rho\_{m,R^e}\sigma(m)\sigma(R^e)
$$

Rearrange and use $$|\rho| \le 1$$:

$$
\frac{\big|E\[R^e]\big|}{\sigma(R^e)} = \big|\rho\_{m,R^e}\big|\frac{\sigma(m)}{E\[m]} \le \frac{\sigma(m)}{E\[m]}
$$

The left-hand side is the Sharpe ratio of the excess return. The right-hand side involves no asset except the riskless one. So:

$$
\boxed{\frac{\sigma(m)}{E\[m]} \ge \mathrm{SR}}
$$

**The volatility of the discount factor, scaled by its mean, must be at least as large as the highest Sharpe ratio available in the market.** No preferences, no distributional assumption, no completeness assumption — only $$p = E\[mx]$$ and the definition of a correlation. This is the same inequality Chapter 3 §3.5 derived in one line at the end of "The beta representation," where it was noted that in that chapter's complete two-state economy the bound held with *equality*: $$\sigma(m)/E\[m] = 0.459$$ and the stock's Sharpe ratio was also 0.459. The equality there was not luck. With two states and two independent assets the market is complete, $$m$$ lies entirely in the payoff space, and the correlation between $$m$$ and the stock's excess return is exactly one. Chapter 3's example was sitting on the Hansen-Jagannathan frontier because it had nowhere else to sit.

Real markets are not complete, so the bound is an inequality, and it becomes a test. Measure the Sharpe ratio; that is a floor on the volatility of any admissible discount factor. Then ask whether a candidate model can clear it.

The consumption model cannot. With $$m = \delta e^{-\gamma g}$$ and lognormal $$g$$, exactly

$$
\frac{\sigma(m)}{E\[m]} = \sqrt{e^{\gamma^2 \sigma\_g^2} - 1} \approx \gamma\sigma\_g
$$

so the bound reduces to the requirement $$\gamma\sigma\_g \gtrsim \mathrm{SR}$$. Table 5.3 puts numbers to it.

**Table 5.3: What power utility delivers against what the market requires**

| $$\gamma$$ | $$\sigma(m)/E\[m]$$ | Largest premium consistent with $$\sigma\_r = 20$$ percent |
| ---------- | ------------------- | ---------------------------------------------------------- |
| 2          | 0.030               | 0.60%                                                      |
| 5          | 0.075               | 1.50%                                                      |
| 10         | 0.151               | 3.02%                                                      |
| 25         | 0.389               | 7.77%                                                      |
| 50         | 0.869               | 17.38%                                                     |

*Source: Author's calculation, consumption growth volatility of 1.5 percent; the third column is the second column times the 20 percent return standard deviation, the premium attainable if the asset's return were perfectly correlated with the discount factor.*

A long-run US equity Sharpe ratio of about 0.30 (a six-percent premium on twenty-percent volatility) requires $$\gamma \approx 19.6$$ exactly, or about 20 by the approximation. A postwar Sharpe ratio nearer 0.50 requires about 31.5. And these are floors that assume the equity return is *perfectly* correlated with the discount factor; the realized correlation between consumption growth and returns pushes the requirement up by a factor of three to five, which is how Table 5.2 got to one hundred.

Figure 5.2 is the table as a curve, with both floors drawn and each requirement read off it.

![Figure 5.2: The Hansen-Jagannathan bound](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-7bf09186c2b2ac592f17e87f56222ff34f22ea73%2Ffig_05_02_hansen_jagannathan_bound.png?alt=media)

**Figure 5.2: The Hansen-Jagannathan bound.** What power utility delivers for the volatility of the discount factor, against what the market requires. The rising curve is the exact expression, and the dotted line beneath it is the approximation that reads the ratio as risk aversion times consumption growth volatility — accurate to two decimals until risk aversion passes about twenty, and increasingly optimistic after that. The markers are Table 5.3's five rows. The two horizontal rules are the Sharpe ratios the section names, and where each crosses the curve is the smallest risk aversion consistent with it: 19.6 for the long-run US ratio of 0.30, and 31.5 for a postwar 0.50. The right-hand axis restates the same curve as Table 5.3's third column — the largest equity premium the model can deliver at 20 percent return volatility — so a reader can enter the figure from either the discount factor or the premium. Two things are worth carrying away from the geometry rather than the arithmetic. The curve is close to a straight line through the origin over the entire admissible range, which is why the puzzle has the shape it does: there is no risk aversion at which the model suddenly starts working, only a proportional cost in units of implausibility for every extra point of Sharpe ratio. And both crossings are floors, not estimates. They assume the equity return is perfectly correlated with the discount factor, which it is not; at the realized correlation the requirement is three to five times higher, and the figure's horizontal axis would have to run past a hundred.

The bound also tells us what a successful $$m$$ has to look like, which is more useful than knowing that a particular one fails. It must be **volatile** — swinging by tens of percent, not by the two or three percent that $$\gamma\sigma\_g$$ delivers at defensible $$\gamma$$ — and it must be **strongly correlated with equity returns**. Aggregate consumption growth is neither. Every resolution in §5.5 is an attempt to manufacture one or both properties: habit formation makes $$m$$ volatile by making the *effective* risk aversion move; long-run risk makes it volatile by adding a persistent component that revalues long-horizon claims; disasters make it volatile by putting probability mass where marginal utility explodes; limited participation makes it volatile by changing whose consumption is in the exponent.

One further property sets the terms of every test that follows: the floor computed from a single excess return is the weakest version of the bound available, because the floor is the highest Sharpe ratio attainable from *any* portfolio of the assets in the test set, so adding bonds, managed portfolios or options can only raise it — which is why stating the equity premium alongside the term premium, as Campbell's lectures do, is the more demanding way to pose the puzzle.

Hansen and Jagannathan's second contribution, the *distance* measure, extends this from a scalar bound to a full region in mean-standard deviation space and gives a metric for how badly a candidate model misses. Cochrane's Chapter 21 works it through; the scalar version above is what carries the argument.

***

## 5.5 The Resolutions and Their Costs

Four decades of work have produced three main lines of attack inside the representative-agent framework, plus the participation line that §5.7 treats as this book's own. Each buys the premium. Each buys it with something, and the something is usually a new prediction that can be tested.

### Habit formation

Campbell and Cochrane (1999) keep power utility's functional form but change its argument. Instead of ranking consumption levels, the household ranks consumption relative to a slow-moving **habit** $$X\_t$$ — a benchmark formed from its own and its neighbors' past consumption. Utility depends on the **surplus consumption ratio**

$$
S\_t \equiv \frac{c\_t - X\_t}{c\_t}
$$

and the local coefficient of relative risk aversion becomes $$\gamma/S\_t$$ rather than $$\gamma$$. Because $$X\_t$$ moves slowly and $$c\_t$$ moves a little, $$S\_t$$ moves a lot: a two-percent fall in consumption when the household is already consuming only five percent above habit is a forty-percent fall in surplus, and effective risk aversion doubles. Small consumption movements are levered into large movements in the price of risk.

What it buys is more than the premium. Effective risk aversion is now **countercyclical** — high in recessions when $$S\_t$$ is low, low in booms — so the model generates a time-varying risk premium, high volatility of stock returns relative to dividends, and return predictability from valuation ratios. Those are the phenomena of Chapter 7, and habit was the first consumption model to produce them jointly with a large unconditional premium. The model is also engineered so that the riskless rate is roughly constant, sidestepping §5.3's second puzzle.

What it costs. The underlying $$\gamma$$ is modest but the *effective* risk aversion in bad times reaches values as extreme as anything in Table 5.2 — the puzzle is relocated into the state variable rather than removed. The habit process is specified to deliver the required behavior rather than derived, and the smoothness of the riskless rate is a design choice, not a result. And the model's sharpest testable implication is that risk premia should be high precisely when the ratio of consumption to its recent past is low; the evidence on that conditional prediction is real but weaker than the unconditional fit.

### Long-run risk

Bansal and Yaron (2004) attack the other input. If consumption growth is smooth in the short run but contains a small, highly persistent component in its conditional mean — call it $$\eta\_t$$, nearly invisible in a variance decomposition but very long-lived — then the *level* of consumption in the distant future is far more uncertain than one-year growth volatility suggests. An asset that is a claim to a stream of dividends extending decades forward is exposed to that long-horizon uncertainty, and the exposure is priced.

For this to generate a premium, the household must care about the timing of the resolution of uncertainty, and that requires the Epstein-Zin preferences of §5.6, with an elasticity of intertemporal substitution $$\psi$$ above one. With $$\psi > 1$$ the price-dividend ratio falls when expected growth falls, so the asset does badly exactly when the news about long-run consumption is bad, and the covariance the premium equation needs appears — not from the realized consumption growth of §5.3, but from *news about future* consumption growth. Adding a persistent stochastic volatility process generates time-varying premia as well.

What it buys: the premium, a low and stable riskless rate, return predictability, and — attractively — an account of why the covariance in Table 5.2 looks so small at annual frequency. The risk is there; it is simply not visible in one-year growth rates.

What it costs. The persistent component is close to unidentifiable in the length of macro data available: a series with a nearly-unit-root conditional mean and small innovations is statistically almost indistinguishable from i.i.d. growth. The model therefore rests on a feature of the data that the data cannot confirm. Beeler and Campbell (2012) pressed the auxiliary predictions and found several that fail: the model implies more predictability of consumption growth from the price-dividend ratio than appears in the data, and it requires an EIS above one, against a large microeconometric literature estimating it below one. And it is sensitive to $$\psi$$: the sign of the price-dividend response to a growth shock flips at $$\psi = 1$$.

### Rare disasters

Rietz (1988) proposed, and Barro (2006) made empirically serious, the possibility that the sample is the problem. Suppose that in each year there is a small probability $$\pi\_D$$ of a **disaster** — a war, a depression, a revolution — in which consumption falls by a fraction $$\kappa$$. The state price of a dollar delivered in such a state is enormous, because marginal utility is proportional to $$(1-\kappa)^{-\gamma}$$, so equity's exposure to it commands a large premium even at modest $$\gamma$$. Barro's contribution was to show that the required parameters are not fanciful: assembling twentieth-century data on many countries, he found contractions of fifteen percent or more occurring at an annual rate on the order of 1.5 to 2 percent, with an average size around a third.

The approximate premium in this economy is

$$
E\[R] - R\_f \approx \gamma\sigma\_g^2 + \pi\_D\left\[(1-\kappa)^{-\gamma} - 1\right]\kappa
$$

and the exact version is easy to compute. Table 5.4 does so for a household with $$\gamma = 4$$ — squarely inside Mehra and Prescott's admissible range — holding a claim to consumption.

**Table 5.4: Disaster risk at** $$\gamma = 4$$

| Disaster probability $$\pi\_D$$ | Disaster size $$\kappa$$ | $$r\_f$$ | $$E\[R]$$ | Premium |
| ------------------------------- | ------------------------ | -------- | --------- | ------- |
| none                            | —                        | 11.3%    | 11.5%     | 0.18%   |
| 1.7%                            | 20%                      | 8.7%     | 9.4%      | 0.69%   |
| 1.7%                            | 29%                      | 6.0%     | 7.7%      | 1.64%   |
| 1.7%                            | 40%                      | −0.1%    | 4.3%      | 4.39%   |
| 1.0%                            | 40%                      | 4.3%     | 7.2%      | 2.85%   |

*Source: Author's calculation. Power utility with risk aversion 4 and a time discount factor of 0.97; log consumption growth normal with mean 2 percent and standard deviation 2 percent in normal times, scaled by one minus the disaster size in a disaster; equity is a one-period claim to consumption. Exact numerical integration, not the approximation above.*

Read the first and fourth rows together. The same household, with the same modest risk aversion, moves from a 0.18-percent premium and an 11.3-percent riskless rate — the two puzzles, in miniature — to a 4.4-percent premium and a riskless rate of essentially zero, purely by admitting a 1.7-percent annual chance of a forty-percent contraction. Disasters fix both halves at once, and that is the mechanism's real appeal: the same fear that makes equity expensive makes safe claims expensive too.

What it costs. The premium is extraordinarily sensitive to the tail parameters, as the table's third and fourth rows show — moving $$\kappa$$ from 0.29 to 0.40 nearly triples it — and those parameters are estimated from the thinnest part of the historical record. The model is close to unfalsifiable on its own terms, since the explanation for why we do not see the disasters is that they are rare. It has a sharper problem too: if investors genuinely feared a 1.7-percent annual chance of a forty-percent collapse, deep out-of-the-money index put options would be far more expensive than they are, and the term structure of option-implied tail probabilities does not match the calibration cleanly (Backus, Chernov and Martin, 2011). Wachter (2013) responded by making disaster *probability* time-varying, which restores volatility and predictability at the cost of another latent state variable.

**Table 5.5: The three resolutions compared**

|                           | Habit (Campbell-Cochrane)                                             | Long-run risk (Bansal-Yaron)                                                   | Rare disasters (Rietz-Barro)                                  |
| ------------------------- | --------------------------------------------------------------------- | ------------------------------------------------------------------------------ | ------------------------------------------------------------- |
| What makes $$m$$ volatile | Effective risk aversion $$\gamma/S\_t$$ moves with the business cycle | News about a persistent growth component, priced through recursive preferences | A small probability of states where marginal utility explodes |
| Preferences               | Power utility over surplus consumption                                | Epstein-Zin, $$\psi > 1$$                                                      | Power utility, modest $$\gamma$$                              |
| Required $$\gamma$$       | Low on average, very high in bad times                                | Around 10                                                                      | 3 to 5                                                        |
| Also delivers             | Countercyclical premia, excess volatility, predictability             | Predictability, stable $$r\_f$$, small annual covariance                       | Low $$r\_f$$, fat-tailed returns                              |
| Main vulnerability        | Habit process is engineered; extreme conditional risk aversion        | Key state variable is not identifiable in available data                       | Tail parameters unmeasurable; option prices disagree          |
| Testable elsewhere        | Conditional premia vs. consumption-to-habit                           | Consumption predictability; the value of $$\psi$$                              | Index option prices; international disaster records           |

*Source: Author's summary of the cited literature.*

The pattern across the three is worth naming. Each buys a volatile $$m$$ by adding a state variable that is hard to measure — a habit level, a persistent growth component, a disaster probability. That is not cheating; it is what the Hansen-Jagannathan bound demands. But it means the models are difficult to separate on the moments they were built to match, and the honest comparison happens on their *other* predictions: what they imply for option prices, for the predictability of consumption, for the term structure of equity claims, and for the conditional behavior of premia. Chapter 7 takes up the last of these.

> **Box 5.2 — Mehra and Prescott's referees**
>
> The equity premium puzzle was not found by someone looking for a puzzle. Rajnish Mehra and Edward Prescott set out to *use* the consumption-based model, in the form Lucas and Breeden had just given it, to account for the observed premium on US equities. The 1985 paper is the report of that attempt failing.
>
> The failure has an unusual structure, and it is why the result has lasted. Most negative findings in economics say that a parameter is imprecisely estimated or that a specification is fragile. This one says the opposite. The model is entirely well behaved, the arithmetic is elementary, and the answer is wrong by orders of magnitude: within a broad class of preferences, an economy calibrated to the observed smoothness of consumption growth cannot produce a premium of more than a small fraction of a percentage point, against the six-plus percent in the authors' own table. There is no better estimator to reach for, because nothing is being estimated.
>
> The paper had a long gestation and a difficult reception, and the scepticism it met was of a specific and instructive kind: if the number is so obviously wrong, something obvious must be wrong with the calculation. Successive readers proposed the obvious things — the sample period, the choice of riskless rate, taxes, survivorship in the US series — and the puzzle absorbed every one of them. That is how a rejected calibration became a research program with subfields of its own.
>
> Two features of that history carry into §5.5. The first is that every resolution in this chapter was proposed as an answer to this specific paper, which is why they all take the same three shapes: change the preferences, change the consumption process, or change the tail of the distribution. The second is that the authors' own reading stayed conservative. Their later surveys treat the puzzle as unresolved rather than as solved by any of the candidates, and this chapter follows them.

***

## 5.6 ★ Epstein-Zin Preferences

*Starred. One equation and its consequences; the full recursive-utility apparatus is in the readings.*

Section 5.2 flagged the defect that long-run risk needs repaired. Under power utility, $$\gamma$$ measures aversion to risk *across states* and $$1/\gamma$$ measures willingness to substitute consumption *across time*. Those are conceptually unrelated attitudes — one is about gambling, the other about growth — and tying them together with a single parameter is an artifact of the additively separable functional form, not an economic proposition. It is also exactly what breaks the model in §5.3: raising $$\gamma$$ to fix the premium mechanically lowers the intertemporal elasticity, which is what sends the riskless rate to thirty-five percent.

Epstein and Zin (1989), building on Kreps-Porteus, break the link with a recursive aggregator. Lifetime utility is defined not over consumption levels but over today's consumption and a **certainty equivalent** of tomorrow's continuation utility:

$$
U\_t = \left\[(1-\delta)c\_t^{1-1/\psi} + \delta\left(E\_t\left\[U\_{t+1}^{1-\gamma}\right]\right)^{\frac{1-1/\psi}{1-\gamma}}\right]^{\frac{1}{1-1/\psi}}
$$

Two parameters, two jobs. The inner exponent $$1-\gamma$$ governs how the household aggregates across *states*: $$\gamma$$ is relative risk aversion. The outer exponent $$1-1/\psi$$ governs how it aggregates across *time*: $$\psi$$ is the elasticity of intertemporal substitution. Power utility is the special case $$\psi = 1/\gamma$$, where the two collapse.

The separation has a consequence beyond parameter freedom, and it is the one long-run risk exploits. When $$\gamma \neq 1/\psi$$ the household is no longer indifferent to *when* uncertainty is resolved. With $$\gamma > 1/\psi$$ it prefers early resolution, and news about the distant future is therefore priced today. That is why a persistent component in expected growth can command a premium under Epstein-Zin and cannot under power utility, where only realized consumption growth enters the discount factor.

The arithmetic payoff for §5.3's second puzzle is immediate. In the i.i.d. lognormal case the Epstein-Zin risk-free rate is

$$
r\_f = -\ln\delta + \frac{\mu\_g}{\psi} - \frac{1}{2}\left\[\gamma + \frac{\gamma-1}{\psi}\right]\sigma\_g^2
$$

while the premium remains $$\gamma\mathrm{Cov}(g,r)$$, exactly as in §5.2. The growth term is now scaled by $$1/\psi$$ rather than by $$\gamma$$. Set $$\gamma = 20$$, $$\mu\_g = 2$$ percent, $$\sigma\_g = 1.5$$ percent, $$\delta = 0.98$$. With power utility ($$\psi = 1/20$$) the model's riskless rate is 37.5 percent. With $$\psi = 1.5$$ it is 3.0 percent. The premium is unchanged; only the interest rate moved, and it moved because $$\psi$$ was allowed to stop being the reciprocal of $$\gamma$$.

That is not a resolution of the equity premium puzzle — $$\gamma = 20$$ is still $$\gamma = 20$$, and the Hansen-Jagannathan bound is still binding on the same object. Epstein-Zin dissolves the *risk-free rate* puzzle and, more consequentially, makes the timing of uncertainty a priced attribute. It is a piece of machinery rather than an explanation, which is why it appears inside Bansal-Yaron and inside most modern consumption models rather than as a resolution in its own right.

***

## 5.7 Who Holds Equity Risk, and What Their Constraints Do to Its Price

Every model in §5.5 keeps the representative agent and works on the shape of its preferences or the process for its consumption. There is another move available, and it is the one this book's framing points to: keep the preferences and change whose consumption is in the exponent.

The premium equation prices equity off the covariance between returns and the marginal utility of **the investor who is marginal in equity**. Chapter 3 §3.7 raised this as the book's organizing question; here it has a number attached. The representative-agent step is legitimate only if everyone holds equity at the margin, because only then does everyone's marginal utility move together with aggregate consumption. Chapter 14 §14.2 documents that a large minority of US households hold no equity at all, and Chapter 14 §14.6, reading the Federal Reserve's Distributional Financial Accounts, shows that the top decile of the wealth distribution holds something on the order of eighty-five percent of directly held corporate equity, and the bottom half about one percent. Aggregate per-capita consumption averages over a population most of whom are not in the market. Their marginal utility is not in any equity price.

Mankiw and Zeldes (1991) drew the obvious inference and did the measurement. Using the Panel Study of Income Dynamics, they split households into stockholders and non-stockholders and built a consumption series for each. Stockholders' consumption growth is substantially more volatile — roughly double, in their food-consumption data — and substantially more strongly correlated with the excess return on equity. Both changes move the premium equation in the right direction, because $$\gamma = (E\[R]-R\_f)/(\rho\_{g,r}\sigma\_g\sigma\_r)$$ falls when either $$\sigma\_g$$ or $$\rho\_{g,r}$$ rises. Hold the premium at six percent and equity volatility at twenty percent, and vary only whose consumption series is used.

**Table 5.6: Whose consumption prices equity**

| Consumption series                   | $$\sigma\_g$$ | $$\rho\_{g,r}$$ | Implied $$\gamma$$ |
| ------------------------------------ | ------------- | --------------- | ------------------ |
| Aggregate per-capita                 | 1.5%          | 0.20            | 100                |
| Stockholders                         | 3.2%          | 0.20            | 47                 |
| Stockholders, with higher comovement | 3.2%          | 0.40            | 23                 |
| Wealthy stockholders                 | 5.0%          | 0.40            | 15                 |

*Source: Author's calculation from §5.2's premium equation. The volatility figures are illustrative magnitudes in the range reported by Mankiw and Zeldes (1991) and the subsequent participation literature; the premium and return volatility are held at 6% and 20% throughout.*

Note what the table does and does not do. Moving from the aggregate household to a wealthy stockholder cuts the required risk aversion by nearly a factor of seven, from a number nobody defends to a number that is merely uncomfortable. **Limited participation shrinks the puzzle; it does not close it.** Chapter 14 §14.6 states the same conclusion from the holdings side, and the subsequent literature — Vissing-Jørgensen's work on participation and the elasticity of substitution, Brav, Constantinides and Geczy's household-level Euler equation tests, Malloy, Moskowitz and Vissing-Jørgensen's finding that long-run stockholder consumption risk does better still — has narrowed the gap without eliminating it.

The framing payoff is larger than the arithmetic. Once the marginal holder is identified as a wealthy household rather than an average one, facts about that holder start to matter for prices, and none of them is a preference parameter. The marginal holder is **levered to equity**: for households in the top percentiles, public and private business equity is the dominant asset, so their consumption is exposed to equity returns structurally rather than incidentally, and entrepreneurial and capital income is far more cyclical than wage income. Chapter 14 §14.6's second reading — that the household sector supplies two demand curves at once — makes the point concrete. On one side is the default-driven contributor whose purchases are governed by payroll dates and a glide path, whose demand curve for equities is close to vertical and who is nobody's marginal investor because she is not choosing. On the other is the wealthy direct holder who can lever, short, and reallocate, and whose realization decisions are price- and tax-sensitive. It is the second curve that has slope, and a demand curve with slope is what a marginal condition describes.

And the marginal holder is **not always a household at all**. The proportion of equity risk borne through intermediaries — mutual funds, pension funds, insurers, dealers — has risen for fifty years, and an intermediary's shadow price of capital is not any household's marginal utility. Chapter 19 §19.5 replaces $$u'(c)$$ with the marginal value of intermediary net worth and finds that it prices assets across classes better than consumption does; Chapter 16 §16.5 supplies the constrained-capital machinery it needs.

So the closing statement of this chapter is a negative one, and it is the reason Chapter 3 §3.7 exists. **The representative consumer prices nothing, because no one holds the average portfolio.** The consumption-based model is not wrong about the mechanism — claims that pay off in bad states are expensive, and that is as true in this chapter as it was in Chapter 3's two-state economy. It is wrong about the identity of the person for whom the states are bad. Fixing that requires knowing who holds the claim, what their balance sheet looks like, and what constrains them. Chapter 20 turns that requirement into an estimation strategy: measure the demand curves directly and derive prices from market clearing, rather than positing a consumer and inverting.

In Chapter 1 §1.2's terms, the equity premium is a standing constraint binding rather than news about consumption: most of the households the aggregate series averages over are kept out of the market by participation costs and have no marginal utility in any equity price, and the premium is what the ones who are in it require.

Two forward pointers close the theory. The time-varying premia that habit and long-run risk generate are the subject of **Chapter 7**, where predictability from valuation ratios is treated as evidence about discount rates rather than about beliefs. And the observation that a single factor — consumption, or the market — cannot price the cross-section leads directly to **Chapter 6**, which asks what the priced factors actually are and whether the resulting zoo is measuring risk or measuring the researcher.

***

## Elsewhere in the Series

* **The macroeconomics of aggregate consumption** — *Institutionalist Macroeconomics*, Chapters 5 and 25. This chapter uses the consumption series as an input and takes its measurement on trust; the permanent-income and buffer-stock literatures, the excess-sensitivity and excess-smoothness findings, and the distribution of consumption across households are developed there. The traffic runs one way: nothing in this chapter is a prerequisite for that material, and nothing there is needed to follow this chapter.
* Everything else here is this book's own. The pricing equation it tests is Chapter 3's; the marginal-holder question it answers partially is Chapter 3 §3.7's; the participation evidence it uses is Chapter 14's.

***

## Summary

1. **Mehra and Prescott's finding is a quantitative failure of a qualitatively sound framework.** Calibrating a representative-agent economy to ninety years of US data and allowing risk aversion up to ten, the largest equity premium the model could produce was about a third of a percentage point against an observed 6.18. The paper's contribution was a well-measured residual, and the decision to call it a puzzle rather than a rejection is what made it a research program.
2. **The consumption-based model is** $$p = E\[mx]$$ **with** $$m$$ **made observable.** With power utility, $$m\_{t+1} = \delta (c\_{t+1}/c\_t)^{-\gamma}$$. The discount factor is high when consumption growth is low, and $$\gamma$$ is the amplifier converting consumption movements into movements in state prices.
3. **Two equations carry the empirics.** Under lognormality, $$r\_f = -\ln\delta + \gamma\mu\_g - \tfrac{1}{2}\gamma^2\sigma\_g^2$$ and $$E\[R]-R\_f \approx \gamma\mathrm{Cov}(g,r)$$. Only covariance with consumption growth is priced, and its price is $$\gamma$$.
4. **Consumption is too smooth and too weakly correlated with returns.** With $$\sigma\_g$$ around 1.5 percent and a correlation of 0.2, a six-percent premium requires $$\gamma = 100$$. Granting a perfect correlation — the most generous assumption available — still requires 20.
5. **The risk-free rate puzzle is the second half.** Because power utility ties intertemporal substitution to $$1/\gamma$$, the $$\gamma$$ that fixes the premium implies a riskless rate of 35 percent at $$\gamma = 20$$ and 62 percent at $$\gamma = 40$$, against a realized figure near one. Forcing a one-percent rate at $$\gamma = 20$$ requires $$\delta = 1.41$$ — negative time preference.
6. **The Hansen-Jagannathan bound generalizes the difficulty.** From $$0 = E\[mR^e]$$ and $$|\rho| \le 1$$, $$\sigma(m)/E\[m] \ge \mathrm{SR}$$, with no assumptions about preferences or distributions. Chapter 3 §3.5's complete two-state economy hit the bound exactly, at 0.459 on both sides, because there $$m$$ lay entirely in the payoff space. A Sharpe ratio of 0.30 with $$\sigma\_g = 1.5$$ percent requires $$\gamma \approx 20$$ as a floor.
7. **A successful discount factor must be volatile and strongly correlated with returns.** That is the specification every resolution is written to satisfy, and it is why each of them adds a hard-to-measure state variable.
8. **The three representative-agent resolutions each buy the premium with something.** Habit makes effective risk aversion $$\gamma/S\_t$$ countercyclical, delivering time-varying premia at the cost of an engineered habit process and extreme conditional risk aversion. Long-run risk prices news about a persistent growth component, at the cost of resting on a state variable macro data cannot identify and requiring $$\psi > 1$$. Disasters deliver a 4.4-percent premium and a near-zero riskless rate at $$\gamma = 4$$ (Table 5.4), at the cost of tail parameters that cannot be measured and option prices that do not fully agree.
9. **Epstein-Zin preferences separate** $$\gamma$$ **from** $$\psi$$**.** The recursive aggregator makes risk aversion and the elasticity of intertemporal substitution independent parameters and makes the timing of the resolution of uncertainty a priced attribute. At $$\gamma = 20$$ and $$\psi = 1.5$$ the implied riskless rate falls from 37.5 to 3.0 percent with the premium unchanged. This dissolves the risk-free rate puzzle and is machinery rather than explanation.
10. **The representative consumer prices nothing, because no one holds the average portfolio.** Stockholders' consumption is roughly twice as volatile as the aggregate and more strongly correlated with returns; using it cuts implied risk aversion from 100 to something in the twenties, and using wealthy stockholders' consumption cuts it further. Limited participation shrinks the puzzle without closing it, and it relocates the question from preferences to holders: who is marginal in equity, how levered are they, and what constrains them. Chapters 19 and 20 answer it.

***

## Key Terms

* **Consumption-based model**: The asset-pricing model that takes $$m = \delta u'(c\_{t+1})/u'(c\_t)$$ literally and measures $$c$$ from national accounts data
* **Power utility (CRRA)**: $$u(c) = (c^{1-\gamma}-1)/(1-\gamma)$$; risk aversion independent of wealth, with a single parameter $$\gamma$$
* **Coefficient of relative risk aversion** $$\gamma$$: The curvature of the utility function; under power utility it is also the price of consumption covariance and the reciprocal of the elasticity of intertemporal substitution
* **Elasticity of intertemporal substitution** $$\psi$$: Willingness to move consumption across time in response to the interest rate; forced to equal $$1/\gamma$$ under power utility, free under Epstein-Zin
* **Equity premium puzzle**: The finding that the consumption-based model requires an implausibly large $$\gamma$$ to match the historical equity premium
* **Risk-free rate puzzle**: Weil's complement — the large $$\gamma$$ that fixes the premium implies a counterfactually high riskless rate, or a discount factor above one
* **Hansen-Jagannathan bound**: $$\sigma(m)/E\[m] \ge \mathrm{SR}$$; the volatility of any admissible discount factor is at least the maximum Sharpe ratio in the market
* **Habit formation**: Preferences defined over consumption relative to a slow-moving benchmark $$X\_t$$; makes effective risk aversion $$\gamma/S\_t$$ countercyclical
* **Surplus consumption ratio** $$S\_t$$: $$(c\_t - X\_t)/c\_t$$; the state variable governing effective risk aversion in the habit model
* **Long-run risk**: A small, highly persistent component $$\eta\_t$$ in expected consumption growth, priced under recursive preferences because it moves the value of long-horizon claims
* **Rare disasters**: A small annual probability $$\pi\_D$$ of a large consumption contraction $$\kappa$$; generates a premium at modest $$\gamma$$ because marginal utility in the disaster state is enormous
* **Epstein-Zin (recursive) preferences**: A utility recursion over current consumption and the certainty equivalent of continuation utility, separating $$\gamma$$ from $$\psi$$
* **Limited participation**: The fact that many households hold no equity, so aggregate consumption is not the consumption of the marginal equity holder
* **Marginal holder**: The investor whose Euler equation actually binds in a given claim, and whose marginal utility therefore appears in its price

***

## Readings

### Required

* Mehra, R. and E. Prescott (1985). "The Equity Premium: A Puzzle." *Journal of Monetary Economics* 15(2): 145-161. *The paper. Read it for the calibration discipline more than the result: the argument is that the model is being given every advantage and still misses by an order of magnitude.*
* Cochrane, J. (2005). *Asset Pricing*, revised edition. Princeton University Press. Chapter 21. *The standard graduate treatment of the puzzle and its resolutions, and the source of this chapter's organizing claim that a successful discount factor must be volatile; read it alongside §§5.3-5.5.*

### Recommended

* Hansen, L. P. and R. Jagannathan (1991). "Implications of Security Market Data for Models of Dynamic Economies." *Journal of Political Economy* 99(2): 225-262. *The bound of §5.4, plus the mean-standard-deviation region and the distance measure that turn it into a full model-diagnostic apparatus.*
* Campbell, J. Y. and J. Cochrane (1999). "By Force of Habit: A Consumption-Based Explanation of Aggregate Stock Market Behavior." *Journal of Political Economy* 107(2): 205-251. *The habit model, engineered so that effective risk aversion is countercyclical and the riskless rate is nearly constant; the joint delivery of the premium and of return predictability is the achievement.*
* Bansal, R. and A. Yaron (2004). "Risks for the Long Run: A Potential Resolution of Asset Pricing Puzzles." *Journal of Finance* 59(4): 1481-1509. *Long-run risk, and the clearest demonstration of why Epstein-Zin preferences are needed for news about the distant future to be priced.*
* Beeler, J. and J. Y. Campbell (2012). "The Long-Run Risks Model and Aggregate Asset Prices: An Empirical Assessment." *Critical Finance Review* 1(1): 141-182. *The audit §5.5 owes long-run risk: the model is confronted with the predictability and consumption-growth implications it makes along the way, not only with the moments it was built to match. Read it against Table 5.5's "testable elsewhere" row.*
* Barro, R. (2006). "Rare Disasters and Asset Markets in the Twentieth Century." *Quarterly Journal of Economics* 121(3): 823-866. *Takes Rietz's idea and measures it, assembling international twentieth-century contractions into a disaster distribution; read it for the empirical construction, which is what made the mechanism respectable.*
* Dimson, E., P. Marsh and M. Staunton (2014). *Credit Suisse Global Investment Returns Yearbook 2014*. Credit Suisse Research Institute. *Long-run equity, bond and bill returns for roughly twenty countries since 1900 — the cross-country twentieth-century record §5.5's disaster subsection describes and the chapter's own data exercise otherwise takes only from the United States.*
* Mankiw, N. G. and S. Zeldes (1991). "The Consumption of Stockholders and Nonstockholders." *Journal of Financial Economics* 29(1): 97-112. *The participation paper behind §5.7: separate the two groups in panel data and stockholders' consumption is more volatile and more correlated with returns, which shrinks the required risk aversion substantially without eliminating the gap.*
* Bodie, Z. (1995). "On the Risk of Stocks in the Long Run." *Financial Analysts Journal* 51(3): 18-22. *Prices the cost of insuring an equity position over a long horizon and finds it rising, not falling, with the horizon; the sharpest available correction to the time-diversification intuition behind §5.7's account of who can bear equity risk.*
* Campbell, J. Y. (2008). "Risk and Return in Stocks and Bonds." Lecture series. *States the premium for two asset classes at once, which is the form §5.4's bound actually wants; the term premium is the second observation any candidate discount factor has to price.*
* Carroll, C. D. Graduate lecture notes on the consumption-based capital asset pricing model and on the equity premium puzzle, Johns Hopkins University. *Two short notes that work the derivation §5.1 compresses and the calibration §5.3 reports; work them as a "do the arithmetic yourself" companion.*

***

## Discussion Questions

1. **Puzzle or measurement artifact?** Three objections say the premium was never really six percent: the US sample is the winner of a survivorship contest among twentieth-century equity markets; realized returns include an unrepeatable revaluation as dividend yields fell and participation broadened; and measured consumption is a poor proxy for the flow of services households actually enjoy. Take each in turn. What would each imply about the *forward-looking* premium, and what evidence would distinguish it from a genuine risk premium? Does any of the three also explain the risk-free rate puzzle, and does it matter that a resolution should explain both?
2. **Whose consumption?** Section 5.7 shows that using stockholders' consumption cuts the implied $$\gamma$$ from about 100 to the twenties. Suppose you could measure the consumption of the top one percent of the wealth distribution perfectly. Would you expect the puzzle to disappear? Name two reasons the measured consumption of very wealthy households might *fail* to reveal their marginal utility even if the data were perfect, and say what Chapter 14 §14.6's "two demand curves" observation implies about which households the Euler equation should be applied to at all.
3. **What would falsify a disaster model?** Table 5.4 shows the premium is highly sensitive to $$\kappa$$. Design a test of the disaster explanation that does not use equity returns. What would you look at, what would the model predict, and what result would you count as a rejection? Why is "we have not observed a disaster recently" not evidence against it?
4. **Relocated or resolved?** The habit model uses a low structural $$\gamma$$ but generates effective risk aversion in bad times as large as the numbers in Table 5.2. Is that a resolution of the puzzle or a relabelling of it? Formulate the criterion you are using to answer, and apply the same criterion to long-run risk.
5. **The bound as a discipline.** Section 5.4's bound uses only the Sharpe ratio of a traded excess return. If a researcher proposes a new discount factor and reports that it prices the market portfolio, what should you check first? Why does the bound get tighter when you add more assets — including managed portfolios and options — to the test set, and what does that do to the models in §5.5?

***

## Problems

**Problem 1 — Implied risk aversion from moments.** An economy has an equity premium of 5.5 percent, an equity return standard deviation of 18 percent, consumption growth with mean 1.9 percent and standard deviation 1.2 percent, and a correlation between consumption growth and the equity return of 0.15.

(a) Compute $$\mathrm{Cov}(g, r)$$ and the coefficient of relative risk aversion implied by the premium equation. (b) Recompute assuming the correlation is one. Explain why this is a lower bound on the implied $$\gamma$$ and why it is not an attainable one. (c) Using the $$\gamma$$ from (b), $$\delta = 1$$, and the risk-free rate equation, compute the implied riskless rate. Compare it to a realized rate of 1 percent. (d) What value of $$\delta$$ would reconcile the $$\gamma$$ from (b) with a 1 percent riskless rate? Interpret the sign of the implied rate of time preference.

**Problem 2 — Hansen-Jagannathan arithmetic.** An investor can form a portfolio with a Sharpe ratio of 0.45. Consumption growth has a standard deviation of 1.1 percent.

(a) State the lower bound the data place on $$\sigma(m)/E\[m]$$. (b) Using the exact power-utility expression $$\sigma(m)/E\[m] = \sqrt{e^{\gamma^2\sigma\_g^2}-1}$$, find the smallest $$\gamma$$ that clears the bound. Compare it to the approximation $$\gamma \ge \mathrm{SR}/\sigma\_g$$. (c) A researcher insists $$\gamma = 10$$ is the largest defensible value. Compute $$\sigma(m)/E\[m]$$ at $$\gamma = 10$$ and the largest equity premium the model can support if $$\sigma\_r = 18$$ percent. By what factor does it fall short of a 6 percent premium? (d) A colleague proposes adding an option-based strategy with a Sharpe ratio of 0.9 to the test set. Does this change the bound? Does it change the model's ability to clear it? Explain the asymmetry.

**Problem 3 — A disaster model.** Consumption growth is lognormal with mean 2 percent and standard deviation 2 percent in normal times. With probability $$\pi\_D = 0.02$$ each year, consumption is additionally multiplied by $$(1-\kappa)$$ with $$\kappa = 0.35$$. The household has power utility with $$\gamma = 4$$ and $$\delta = 0.97$$, and equity is a one-period claim to consumption.

(a) Using the approximation $$E\[R]-R\_f \approx \gamma\sigma\_g^2 + \pi\_D\[(1-\kappa)^{-\gamma}-1]\kappa$$, compute the premium. How much of it comes from the normal-times term? (b) Compute the premium exactly by numerical integration and compare. Why does the approximation overstate it? (c) Compute the riskless rate with and without the disaster branch. Explain in one sentence why disasters address both puzzles at once. (d) Hold the expected loss $$\pi\_D\kappa$$ fixed and halve $$\kappa$$ while doubling $$\pi\_D$$. What happens to the premium, and what does the answer tell you about which moment of the disaster distribution the model is sensitive to?

**Problem 4 ★ — Epstein-Zin comparative static.** Consumption growth is i.i.d. lognormal with $$\mu\_g = 2$$ percent and $$\sigma\_g = 1.5$$ percent. Take $$\delta = 0.98$$ and $$\gamma = 15$$.

(a) Compute the riskless rate under power utility. (b) Compute it under Epstein-Zin with $$\psi = 1.5$$ and with $$\psi = 0.5$$, using

$$
r\_f = -\ln\delta + \frac{\mu\_g}{\psi} - \frac{1}{2}\left\[\gamma + \frac{\gamma-1}{\psi}\right]\sigma\_g^2
$$

(c) Verify that setting $$\psi = 1/\gamma$$ recovers your answer to (a). (d) The premium in this i.i.d. economy is $$\gamma\mathrm{Cov}(g,r)$$ regardless of $$\psi$$. Explain what Epstein-Zin does and does not fix, and state the extra ingredient long-run risk requires in order for $$\psi$$ to affect the premium as well.

**Problem 5 — Reading the bound backwards.** You observe an equity premium of 6 percent and $$\sigma\_r = 20$$ percent. You are told the true marginal investor's consumption growth has a correlation of 0.4 with equity returns.

(a) What is the smallest $$\sigma\_g$$ consistent with $$\gamma \le 10$$? (b) Aggregate per-capita consumption growth has $$\sigma\_g \approx 1.5$$ percent. By what factor must the marginal investor's consumption be more volatile than the aggregate? (c) Chapter 14 §14.6 reports that the top decile holds roughly 85 percent of directly held equity. Sketch, in words, how you would use household-level data to test whether the required volatility is present, and state one reason your test would be biased even with perfect data.

***

## Selected Solutions

*Solutions to Problems 1 and 4 follow. Solutions to the remainder are in the instructor materials.*

**Problem 1.**

(a) $$\mathrm{Cov}(g,r) = \rho\sigma\_g\sigma\_r = 0.15 \times 0.012 \times 0.18 = 0.000324$$. Then $$\gamma = 0.055/0.000324 = 169.8$$.

(b) With $$\rho = 1$$, $$\mathrm{Cov}(g,r) = 0.012 \times 0.18 = 0.00216$$ and $$\gamma = 0.055/0.00216 = 25.5$$. It is a lower bound because $$\gamma$$ is inversely proportional to $$\rho$$ and a correlation cannot exceed one; it is not attainable because a correlation of one would make equity a deterministic function of consumption growth, which the data reject decisively.

(c) $$r\_f = 0 + 25.5(0.019) - \tfrac{1}{2}(25.5)^2(0.012)^2 = 0.4838 - 0.0467 = 0.4371$$, or **43.7 percent**. Against a realized 1 percent, the model misses the interest rate by more than forty percentage points using the *most favorable* risk-aversion parameter available. This is the risk-free rate puzzle in its cleanest form: the parameter chosen to fit one moment destroys another.

(d) Set $$-\ln\delta = 0.01 - 0.4371 = -0.4271$$, so $$\ln\delta = 0.4271$$ and $$\delta = 1.533$$. The household values consumption a year hence more than half again as much as consumption today — a rate of time preference of about $$-35$$ percent. Nothing in the theory forbids $$\delta > 1$$, but nothing in observed behavior supports it either, and a model that needs it has stopped being a description of preferences.

**Problem 4.**

(a) Power utility: $$r\_f = -\ln(0.98) + 15(0.02) - \tfrac{1}{2}(15)^2(0.015)^2 = 0.0202 + 0.30 - 0.0253 = 0.2949$$, or **29.5 percent**.

(b) Epstein-Zin with $$\psi = 1.5$$:

$$
r\_f = 0.0202 + \frac{0.02}{1.5} - \tfrac{1}{2}\left\[15 + \frac{14}{1.5}\right] (0.015)^2 = 0.0202 + 0.01333 - 0.00274 = 0.03080
$$

or **3.1 percent**. With $$\psi = 0.5$$: $$r\_f = 0.0202 + 0.04 - \tfrac{1}{2}\[15+28] (0.000225) = 0.0202 + 0.04 - 0.00484 = 0.0554$$, or **5.5 percent**. Lowering $$\psi$$ raises the rate, because a household less willing to substitute across time needs a larger inducement to postpone consumption in a growing economy.

(c) With $$\psi = 1/15$$: $$\mu\_g/\psi = 15(0.02) = 0.30$$, and $$\tfrac{1}{2}\[15 + 15 \times 14] (0.000225) = \tfrac{1}{2}(225)(0.000225) = 0.0253$$. So $$r\_f = 0.0202 + 0.30 - 0.0253 = 0.2949$$, matching (a) exactly, as it must: $$\gamma + (\gamma-1)/\psi = \gamma + \gamma(\gamma-1) = \gamma^2$$.

(d) Epstein-Zin fixes the *risk-free rate* puzzle by unlinking the intertemporal channel from the risk channel; the premium in this i.i.d. economy still requires $$\gamma = 15$$, so the equity premium puzzle is untouched and the Hansen-Jagannathan bound is still binding on the same $$m$$. For $$\psi$$ to affect the premium, consumption growth must be predictable — a persistent component in its conditional mean — so that continuation utility becomes a genuine source of risk. That is precisely the ingredient long-run risk supplies, and it is why the two pieces of machinery always appear together.

***

## Data Exercise: How Big Is the Premium, and Can Consumption Price It?

**Part A — The premium by decade (free data: Shiller).** Download Robert Shiller's long-run US series from his Yale website: monthly S\&P Composite price, dividends, earnings, the consumer price index, and the long interest rate, beginning 1871.

1. Construct an annual real total return on the index (price change plus dividends, deflated by the CPI) and an annual real return on the short-term rate series. Compute the arithmetic and geometric mean premium over the full sample.
2. Report the mean premium and the realized Sharpe ratio *by decade*. Comment on the dispersion: how many decades would, on their own, have produced a puzzle, and how many would not? What does this say about the standard error attached to the six-percent figure?
3. Split the sample at 1945. Compare the pre- and postwar premium, volatility, and Sharpe ratio, and state which of Discussion Question 1's three objections your split speaks to.

**Part B — Consumption moments (free data: FRED).** From FRED, download real personal consumption expenditures on nondurable goods and services, and total population, at quarterly frequency (series for real PCE by major type of product; construct per-capita figures yourself).

1. Build annual real per-capita consumption growth. Report $$\mu\_g$$ and $$\sigma\_g$$ over the full postwar sample and over the period since 1990. Note how much of the volatility comes from 2008-09 and 2020, and report the moments with those years excluded.
2. Compute the correlation and covariance between consumption growth and the annual real equity return from Part A. Be explicit about the timing convention — consumption is a flow measured over the year, returns are measured point to point — and report the moments under at least two conventions.
3. Using the premium equation, report the implied $$\gamma$$ for each combination of sample period and timing convention. Present it as a table. The spread across specifications is itself the answer to a question worth asking.

**Part C — The bound.**

1. Compute the realized Sharpe ratio of the equity excess return over your postwar sample. That is the Hansen-Jagannathan floor on $$\sigma(m)/E\[m]$$.
2. Plot $$\sqrt{e^{\gamma^2\sigma\_g^2}-1}$$ against $$\gamma$$ for your estimated $$\sigma\_g$$, and mark the floor. Read off the smallest $$\gamma$$ that clears it.
3. Now add a second asset: construct a portfolio that is long equities and short bonds at a leverage that maximizes the in-sample Sharpe ratio, or use a value-minus-growth return from Ken French's data library. Recompute the floor. Report how much the bound tightens, and explain why adding assets can only move the floor up.

**Part D ★ (if you have WRDS).** Repeat Part B using the Consumer Expenditure Survey to construct separate consumption growth series for households reporting equity holdings and those reporting none, following Mankiw and Zeldes (1991). Report $$\sigma\_g$$ and $$\rho\_{g,r}$$ for each group and the implied $$\gamma$$, and reproduce Table 5.6 with estimated rather than illustrative inputs. State clearly what the CEX's measurement error does to your estimates and in which direction.

**Part E — The international record.** Take the long-run country series in the Dimson-Marsh-Staunton yearbook (equity, bond and bill returns since 1900) and tabulate the realized equity premium over bills for each country, ranked. Where does the United States sit in that ranking, and by how much does the world average fall short of the US figure? Then ask what the ranking would look like if the markets that closed or expropriated their investors during the century were carried at their true realized return rather than dropped. Report which of §5.5's resolutions the exercise supports and which it embarrasses.
