> For the complete documentation index, see [llms.txt](https://laurence-wilse-samson.gitbook.io/textbooks/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://laurence-wilse-samson.gitbook.io/textbooks/financial-economics-claims-prices-holders/part-ii-asset-pricing/chapter_07_information_efficiency.md).

# Chapter 7: Information, Efficiency, and Price Discovery

*Part II: Asset Pricing — Financial Economics: Claims, Prices, and Holders*

***

## Opening Episode: One Prize, Two Verdicts

On the morning of October 14, 2013, the Royal Swedish Academy of Sciences announced that the year's economics prize would be shared three ways, "for their empirical analysis of asset prices." The recipients were Eugene Fama and Lars Peter Hansen, both of the University of Chicago, and Robert Shiller of Yale.

The pairing was noticed immediately, and not kindly. Fama had spent five decades arguing that competitive markets impound information into prices so quickly that the price series is close to unforecastable, and that apparent anomalies are usually failures of the model of expected returns rather than failures of the market. Shiller had spent four decades arguing that stock prices move far more than any plausible account of fundamentals can justify, that the movements are partly predictable, and that the predictability reflects waves of enthusiasm and gloom. He had said so at book length in *Irrational Exuberance*, which reached bookshops in March 2000, within weeks of the NASDAQ peak. In the American financial press the joint award was reported as a prize for a proposition and its negation.

The reading is understandable and it is wrong, and getting clear about why is the work of this chapter.

Start with what the two men agree about, which is nearly all of the evidence. Both accept that stock returns over the next day, week, or month are close to unforecastable from past returns. Both accept that returns over the next five to ten years *are* forecastable, and forecastable from the same variable: the level of prices relative to dividends or earnings. Neither disputes the regressions — Fama himself, with Kenneth French, published two of the founding papers on long-horizon predictability in 1988 and 1989. The disagreement is entirely about what the coefficient means.

Fama's reading: the expected return investors require for holding equity is not a constant. It rises in bad times, when wealth is low and risk-bearing capacity is scarce, and falls in good times. A high required return is, mechanically, a low price relative to dividends. So low prices forecast high returns because they *are* high required returns, and the regression measures a risk premium rather than a mispricing. Chapter 5's habit model is one formalization of exactly this, built to deliver it.

Shiller's reading: the required return is roughly stable, and prices swing around fundamental value because investors extrapolate. High prices forecast low returns because investors bought at prices their own beliefs could not support and were subsequently disappointed. Chapter 15 §15.3 supplies the sharpest testimony for this side — survey measures of investor expectations that are high when prices are high, that fail to forecast returns, and that move opposite to the model-implied expected return the regression recovers.

Hansen sits between them, and his placement in the citation is the committee's most interesting judgment. His generalized method of moments made it possible to test asset-pricing restrictions without committing to a full distributional specification, which is what turns "prices reflect information" from a slogan into a hypothesis with a rejection region. Hansen's tools are why the disagreement between the other two is scientific rather than temperamental: both positions are stated as restrictions on the same data, and both survive.

The contribution being honored, then, was not a conclusion. It was a well-posed and unresolved question — what is the object that a return-forecasting regression measures? — together with the apparatus that makes the question answerable in principle. Sections 7.1 and 7.2 establish what efficiency claims and why it can never be complete. Sections 7.3 and 7.4 are the two great testing traditions and where they landed. Section 7.5 asks why the money that would settle the matter does not arrive, and §7.6 asks what "prices reflect information" means when the information includes the models themselves.

***

## 7.1 What Efficiency Actually Claims

The efficient markets hypothesis is often stated as "prices reflect all available information," which is not a hypothesis because it does not say which information, reflected how, or by what standard.

Fama's 1970 survey supplied the missing pieces, and the discipline it imposed is the reason the paper is still read. Efficiency is defined **relative to an information set**, written $$\Omega\_t$$, and the content of the hypothesis is that no trading rule based on $$\Omega\_t$$ earns an abnormal return. Fama's three nested cases have become standard vocabulary:

* **Weak form**: $$\Omega\_t$$ is the history of prices and returns. Technical trading rules do not work.
* **Semi-strong form**: $$\Omega\_t$$ adds all public information — earnings, filings, news, macroeconomic releases. Trading on a public announcement after it is public does not work.
* **Strong form**: $$\Omega\_t$$ adds private information. Even insiders cannot beat the market.

The nesting matters. Strong-form efficiency implies semi-strong implies weak, so evidence against the weak form is evidence against all three, and evidence against the strong form says almost nothing about the other two. Nobody has ever seriously defended the strong form; corporate insiders' trades earn abnormal returns, which is why they are regulated and disclosed. The live hypothesis is the semi-strong form, and the interesting evidence is about it.

**The joint-hypothesis problem.** The phrase "abnormal return" is where the trouble enters, and it is not a technicality. Abnormal relative to what? A return is abnormal only against a benchmark of what the return *should have been*, and that benchmark is a model of equilibrium expected returns. So every test of market efficiency is a test of two things at once: that prices reflect $$\Omega\_t$$, and that the model of expected returns is correct. When the test rejects, the rejection cannot be assigned. It may be that the market is inefficient, or that the model of expected returns is wrong, and no amount of additional data separates them, because the data are what both hypotheses are being fitted to.

Chapter 6's opening episode is the canonical illustration. Fama and French found in 1992 that beta had no reliable relation to average returns once size and book-to-market were controlled. Read one way this is a large anomaly and evidence of inefficiency; read the other way it is a rejection of the Sharpe-Lintner CAPM as the model of expected returns. Fama took the second reading and published a three-factor replacement within a year, reclassifying size and value from anomalies into priced state variables. No experiment adjudicates. The joint-hypothesis problem does not make efficiency untestable; it makes it untestable *in isolation*, which is why the two dominant testing traditions both work by narrowing the window in which the model of expected returns has to be right. Event studies (§7.3) shrink the window to days, so that any plausible expected return is small. Long-horizon regressions (§7.4) do the opposite and accept the joint hypothesis openly, then argue about which half moved.

**What efficiency does not claim.** Three misreadings are worth clearing. Efficiency does not claim that prices equal fundamental value; it claims that prices are the market's best conditional forecast of fundamental value given $$\Omega\_t$$, which permits large errors ex post. Efficiency does not claim that returns are unpredictable: predictable variation in expected returns is what a time-varying risk premium *is*, and unpredictability follows only when one adds the auxiliary assumption that expected returns are constant — exactly the assumption the modern literature abandons. And efficiency does not claim that no one profits from information. If it did it would contradict itself, which is the subject of §7.2.

**Efficiently inefficient, as a preview.** The modern resolution, and the one this book adopts, is stated most compactly by Pedersen: markets are *efficiently inefficient*. Prices are inefficient enough that skilled investors who bear the costs of information and the risks of arbitrage are compensated for doing so, and efficient enough that no one else is. Mispricings exist and are the wage of the people who remove them. This is not a compromise between Fama and Shiller. It is the equilibrium implication of taking the cost of information seriously, and Grossman and Stiglitz wrote it down in 1980.

***

## 7.2 Grossman-Stiglitz: Why Prices Cannot Be Fully Revealing

The argument is a paradox before it is a model, and the paradox is worth stating in three sentences.

If prices reflected all information, an investor could learn everything by reading the price and would never pay for research. If nobody pays for research, nobody trades on information, and the price reflects nothing. So a fully revealing price cannot be an equilibrium, and whatever the equilibrium is, it must leave the informed strictly compensated.

Grossman and Stiglitz (1980) made this precise. Their model is the canonical statement of what a price is: not a sufficient statistic for information, but a noisy aggregator whose accuracy is an endogenous, priced quantity. It is the backbone of Chapter 11's microstructure models, of Chapter 17's treatment of what the passive shift does to price discovery, and of the active-management industry's claim to exist.

**The economy.** One risky claim, one riskless claim. The risky claim pays

$$
x = y + \varepsilon
$$

where $$y \sim N(\bar y, \sigma\_y^2)$$ is the *learnable* component and $$\varepsilon \sim N(0, \sigma\_\varepsilon^2)$$, independent of $$y$$, is the part nobody can learn. The riskless claim returns $$R\_f$$. Traders have constant absolute risk aversion $$\tau$$ and, before trading, may pay a cost $$\xi$$ in units of date-0 wealth to observe $$y$$ exactly. Write $$n \in \[0,1]$$ for the fraction who do.

The one piece of machinery that makes the model solvable is the **CARA-normal setup**, and it is worth stating in words rather than deriving. Under constant absolute risk aversion with normally distributed payoffs, an investor's optimal holding of the risky claim is the expected excess payoff divided by risk aversion times the conditional variance:

$$
D = \frac{E\[x \mid \Omega] - R\_fp}{\tau\mathrm{Var}\[x \mid \Omega]}
$$

Two features drive everything that follows. Demand does not depend on wealth, so traders aggregate without tracking the wealth distribution. And demand is *linear* in the conditional mean, which makes the equilibrium price a linear function of the underlying shocks and therefore itself a normal signal.

The last ingredient is the one the paradox requires. The per-capita supply of the risky claim is random: $$z \sim N(\bar z, \sigma\_z^2)$$, independent of everything else. This is **noise**, and it stands for whatever moves the market for reasons unrelated to $$y$$ — liquidity needs, rebalancing, index flows, taxes, retirements. Without it the model has no equilibrium at all, and that is the point rather than a repair.

**The price as a signal.** Informed traders know $$y$$, so their conditional variance is $$\sigma\_\varepsilon^2$$ and their demand is $$(y - R\_f p)/(\tau\sigma\_\varepsilon^2)$$. Uninformed traders condition on the price. Setting aggregate demand equal to supply,

$$
n\frac{y - R\_f p}{\tau \sigma\_\varepsilon^2} + (1-n)\frac{E\[x\mid p] - R\_f p}{\tau\mathrm{Var}\[x\mid p]} = z
$$

and rearranging shows that the price carries exactly the information in the single statistic

$$
w = y - \frac{\tau \sigma\_\varepsilon^2}{n}z
$$

Read that line slowly, because the entire economics of the model is in it. The price tells you $$y$$ contaminated by supply noise, and the contamination is scaled by $$1/n$$. When many traders are informed, their aggregate demand responds sharply to $$y$$, so a given supply shock moves the price little relative to the information it carries, and the price is a precise signal. When few are informed, the same supply shock swamps the signal. Price informativeness is increasing in the number of people paying to be informed — and that is exactly why it cannot be complete, because each of them has to be paid.

**Free entry.** Under CARA-normal, the ratio of the two traders' ex-ante expected utilities depends only on the ratio of their conditional variances and on the cost (the algebra is in the starred subsection). Setting it to one — nobody wants to switch type — gives the **free-entry condition**

$$
\frac{\mathrm{Var}\[x \mid p]}{\mathrm{Var}\[x \mid y]} = e^{2\tau\xi}
$$

with $$\mathrm{Var}\[x \mid y] = \sigma\_\varepsilon^2$$. This one equation is the model's whole content, and it says something striking before any numbers are attached. **In equilibrium the informational disadvantage of the uninformed is pinned down by the cost of information and nothing else.** Not by how much there is to learn, not by how much noise trading there is, not by how many analysts there are. Those things determine how many informed traders it takes to reach the equilibrium gap; they do not change the gap.

To turn it into a number, define the signal-to-noise ratio of the price,

$$
\Psi(n) = \frac{\sigma\_y^2}{\big(\tau\sigma\_\varepsilon^2/n\big)^2 \sigma\_z^2} = \frac{n^2\sigma\_y^2}{\tau^2\sigma\_z^2\sigma\_\varepsilon^4}
$$

so that $$\mathrm{Var}\[y \mid p] = \sigma\_y^2/(1+\Psi)$$ by the standard normal updating formula, and $$\mathrm{Var}\[x\mid p] = \sigma\_\varepsilon^2 + \sigma\_y^2/(1+\Psi)$$. Substituting into the free-entry condition and solving,

$$
\Psi^{\ast} = \frac{\sigma\_y^2}{\left(e^{2\tau\xi} - 1\right)\sigma\_\varepsilon^2} - 1, \qquad n^{\ast} = \frac{\tau\sigma\_z\sigma\_\varepsilon^2}{\sigma\_y}\sqrt{\Psi^{\ast}}
$$

truncated to $$\[0,1]$$. Price informativeness — the fraction of the variance of the learnable component that the price reveals — is $$\Psi^{\ast}/(1+\Psi^{\ast})$$.

**Table 7.1: Equilibrium information production as the cost of information varies**

| Cost $$\xi$$ | $$e^{2\tau\xi}$$ | Equilibrium $$\Psi$$ | Informed fraction $$n^{\ast}$$ | Price informativeness |
| ------------ | ---------------- | -------------------- | ------------------------------ | --------------------- |
| 0.10         | 1.492            | 1.000                | 1.000 (corner)                 | 0.500                 |
| 0.20         | 2.226            | 1.000                | 1.000 (corner)                 | 0.500                 |
| 0.275        | 3.000            | 1.000                | 1.000                          | 0.500                 |
| 0.30         | 3.320            | 0.724                | 0.851                          | 0.420                 |
| 0.35         | 4.055            | 0.309                | 0.556                          | 0.236                 |
| 0.40         | 4.953            | 0.012                | 0.109                          | 0.012                 |
| 0.45         | 6.050            | 0                    | 0                              | 0                     |

*Source: Author's calculation from the equilibrium conditions above, with risk aversion 2, a learnable-component variance of 4, and unit variances for the unlearnable part of the payoff and for noise supply. The interior region runs from a cost of 0.275 to a cost of 0.402 — one-quarter the log of 3 and one-quarter the log of 5; below it every trader informs, above it none does. In the two corner rows the reported signal-to-noise ratio is the value realized when every trader informs, not the interior solution, which there exceeds it.*

Three readings of the table.

Figure 7.2 draws both of them, and adds the comparative static the table cannot show.

![Figure 7.2: Grossman-Stiglitz equilibrium](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-3b0223af464b8cb32b407c9413c88c213e3b01f9%2Ffig_07_02_grossman_stiglitz.png?alt=media)

**Figure 7.2: Grossman-Stiglitz equilibrium.** Panel (a) is Table 7.1 as a continuum, with the informed fraction and price informativeness plotted against the cost of information and the table's seven rows marked on them. The two shaded regions are the corners — every trader informs below a cost of 0.275, nobody does above 0.402 — and between them the equilibrium is interior everywhere, with the informed fraction falling continuously and price informativeness with it. Nothing forces a corner, and nothing makes the price fully revealing: at the lower boundary informativeness is one half, not one, which is the model's central claim drawn rather than asserted. Panel (b) holds the cost at 0.30 and varies the amount of noise trading instead. The informed fraction rises in exact proportion to it and price informativeness does not move at all, staying at 0.420 across a fivefold change in the amount of noise — analysts are hired to offset noise, not to overcome it. Past a noise level of about 1.18 even that stops: everybody is informed already, the informed fraction cannot rise further, and additional noise dilutes the price rather than supporting more research. That last stretch is the one part of the picture the chapter's table does not reach, and it is where the model's logic runs backwards on itself.

The equilibrium is **interior** over the whole intermediate range, and the informed fraction falls continuously with the cost of information. Nothing forces it to a corner, and nothing makes prices fully revealing. The last row is the other corner: when information is expensive enough relative to what there is to learn, nobody buys it and the price reveals nothing — a market that has stopped doing price discovery, which is a fair description of some corners of the municipal and private-credit markets.

The **comparative static in noise** is the one Chapter 17 needs, and it is invisible in the table by construction. Holding $$\xi$$ at 0.30 and varying the amount of noise trading, $$\Psi^{\ast}$$ stays at 0.724 and price informativeness stays at 0.420, while the informed fraction moves in exact proportion to $$\sigma\_z$$: $$n^{\ast} = 0.213$$ at $$\sigma\_z = 0.25$$, $$0.425$$ at $$\sigma\_z = 0.5$$, $$0.851$$ at $$\sigma\_z = 1$$. More noise trading supports a larger research industry and buys no additional price accuracy. Analysts are hired to offset the noise, not to overcome it.

And the **impossibility result** falls out of the same algebra. Let the noise vanish, $$\sigma\_z \to 0$$. Then for any $$n > 0$$ the statistic $$w$$ becomes a deterministic function of $$y$$, the price is fully revealing, $$\mathrm{Var}\[x \mid p] = \sigma\_\varepsilon^2$$, and the left side of the free-entry condition equals one while the right side exceeds one for any $$\xi > 0$$. No informed trader is compensated, so $$n = 0$$. But at $$n = 0$$ the price reveals nothing and information is strictly worth buying. Neither candidate survives: **with costly information and no noise, competitive equilibrium does not exist.** That is Grossman and Stiglitz's title — "On the impossibility of informationally efficient markets" — and it is not a claim about frictions or behavior. It is a claim that a price reflecting all information is internally inconsistent as soon as information costs anything.

### ★ Solving the linear equilibrium

*Starred. The steps between the market-clearing condition and the free-entry condition. A reader who skips it loses the construction, not the result.*

Two gaps remain in the account above. The first is why $$w$$ exhausts the price's information. Multiply the market-clearing condition through by $$\tau\sigma\_\varepsilon^2/n$$ and collect:

$$
y - \frac{\tau\sigma\_\varepsilon^2}{n}z = R\_f p - \frac{(1-n)\sigma\_\varepsilon^2}{n\mathrm{Var}\[x\mid p]}\Big(E\[x\mid p] - R\_f p\Big)
$$

The right-hand side is a function of $$p$$ alone. So observing $$p$$ is informationally equivalent to observing $$w$$, whatever the coefficients of the price function turn out to be — which is what licenses conditioning on $$w$$ rather than solving for the price function first. Since $$w$$ is normal with mean $$\bar y - (\tau\sigma\_\varepsilon^2/n)\bar z$$ and noise variance $$(\tau\sigma\_\varepsilon^2/n)^2\sigma\_z^2$$, normal updating gives $$\mathrm{Var}\[y\mid w] = \sigma\_y^2/(1+\Psi)$$ directly.

The second gap is the free-entry condition. With CARA utility $$U(W) = -e^{-\tau W}$$ and normal payoffs, a trader with information set $$\Omega$$ who chooses $$D$$ optimally attains a certainty equivalent $$R\_f W\_0 + (E\[x\mid\Omega] - R\_f p)^2 / \big(2\tau\mathrm{Var}\[x\mid\Omega]\big)$$, so realized expected utility is an exponential of a quadratic in a normal variable. Taking its unconditional expectation and forming the ratio for the two types, the terms involving the *level* of the expected excess payoff cancel, and what survives is

$$
\frac{EU\_{\text{informed}}}{EU\_{\text{uninformed}}} = e^{\tau\xi}\left(\frac{\mathrm{Var}\[x\mid y]}{\mathrm{Var}\[x\mid p]}\right)^{1/2}
$$

Both expected utilities are negative, so the informed are better off when this ratio is *below* one, and indifference is the ratio equal to one — which on squaring is the free-entry condition in the text. The cost enters as $$e^{\tau\xi}$$ because under CARA a fixed wealth charge multiplies expected utility by exactly that factor.

Two structural features are worth naming because Chapter 11 reuses them. Information enters through the conditional *variance*, not through any expected-return advantage: the informed earn more because they take larger positions when they are right, not because the asset is mispriced from their point of view. And the model is a **rational expectations equilibrium** in the demanding sense — uninformed traders know the mapping from $$(y, z)$$ to $$p$$ and invert it correctly. They are not naive. They are simply unable to separate a high price caused by good news from a high price caused by low supply, and it is that inseparability, not any error, that leaves room for the informed.

***

## 7.3 Event Studies: Efficiency Tested in a Narrow Window

Event studies exist because of the joint-hypothesis problem, not in spite of it.

Over a one-day window, expected return is on the order of four basis points. Whether the correct model of expected returns is the CAPM, a five-factor model, or a constant, the implied benchmark differs by a fraction of a basis point, while the price reaction to a large corporate announcement is measured in whole percentage points. The joint hypothesis has not disappeared, but its second half has been shrunk until it cannot plausibly account for the finding. This is the sharpest available test of semi-strong efficiency, and it is why the method spread out of finance into law, accounting, antitrust, and regulatory economics.

**The logic.** Fama, Fisher, Jensen and Roll's 1969 study of stock splits established the template. Fix an event date $$t = 0$$ and estimate a benchmark return model on a window well before the event — commonly a market model, $$R\_{i,t} = \alpha\_i + \beta\_i R\_{M,t} + \varepsilon\_{i,t}$$, fitted over the prior year of daily data. Then in the event window compute the **abnormal return**

$$
AR\_{i,t} = R\_{i,t} - \left(\hat\alpha\_i + \hat\beta\_i R\_{M,t}\right)
$$

and cumulate it over the window to obtain $$CAR\_i(t\_1, t\_2) = \sum\_{t=t\_1}^{t\_2} AR\_{i,t}$$. Average across many events to kill the idiosyncratic noise, and test whether the average is zero. The full mechanics — estimation-window choice, clustering when events share a calendar date, the correction to standard errors for estimation error in $$\hat\alpha$$ and $$\hat\beta$$, and the treatment of long-horizon windows where benchmark misspecification does return with force — are in **Appendix A**, which is self-contained.

**A worked window.** Suppose a firm's market model over the prior year gives $$\hat\alpha = 0.02$$ percent per day and $$\hat\beta = 1.2$$, with residual standard deviation $$1.0$$ percent per day. It announces earnings after the close on day 0.

**Table 7.2: An earnings-announcement window**

| Day | Firm return | Market return | Benchmark $$\hat\alpha + \hat\beta R\_M$$ | Abnormal return |
| --- | ----------- | ------------- | ----------------------------------------- | --------------- |
| −1  | 0.40%       | 0.50%         | 0.62%                                     | −0.22%          |
| 0   | 5.80%       | 0.30%         | 0.38%                                     | 5.42%           |
| +1  | 1.10%       | −0.20%        | −0.22%                                    | 1.32%           |
| +2  | 0.60%       | 0.40%         | 0.50%                                     | 0.10%           |
| +3  | 0.50%       | 0.10%         | 0.14%                                     | 0.36%           |

*Source: Author's constructed example.*

The announcement-day abnormal return is 5.42%, and $$CAR(-1,+1) = 6.52$$ percent, with standard error $$1.0 \times \sqrt{3} = 1.73$$ percent and $$t = 3.8$$. That part is efficiency working exactly as advertised: public information arrives and is in the price the same day.

The interesting number is the other one. $$CAR(+1,+3) = 1.78$$ percent — an abnormal return earned entirely *after* the information was public. In a single firm this is noise ($$t = 1.0$$). Averaged across tens of thousands of announcements sorted by the size of the earnings surprise, it is not.

Figure 7.3 lays the whole apparatus out: where the benchmark comes from, and what the window shows once it has been subtracted.

![Figure 7.3: An event study window](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-1f13f63e5b7d5d115b5e3b876277a690c9bba762%2Ffig_07_03_event_study_window.png?alt=media)

**Figure 7.3: An event study window.** Panel (a) is the design. The market model is fitted over a year of daily data that ends well before the event — the gap is there so that anticipation of the announcement cannot contaminate the benchmark — and only then is the five-day event window scored against it. Almost every methodological choice in event studies is a choice about this rule: how long the estimation window is, how wide the gap is, and what is done when many firms announce on the same date. Panel (b) is Table 7.2 with the abnormal returns as bars and their running sum as a line, inside a band of plus and minus 1.96 standard errors that widens with the square root of the days cumulated, because independent daily residuals accumulate variance and not standard deviation. Two readings, and they point in opposite directions. The three-day window around the announcement is 6.52 percent against a standard error of 1.73, a t of 3.8, and it is efficiency working exactly as advertised: public information arrives and is in the price the same day. The three days after it sum to 1.78 percent — an abnormal return earned entirely after the information was public, and one that a mechanical rule using nothing private could have collected. In one firm that is noise. The whole of post-earnings-announcement drift is the claim that across tens of thousands of announcements it is not.

**Post-earnings-announcement drift.** Ball and Brown documented in 1968 that prices continue to drift in the direction of an earnings surprise for weeks after the announcement; Bernard and Thomas established in the late 1980s that the drift survives risk adjustment, transaction-cost estimates, and every specification test thrown at it, and that its magnitude across the extreme deciles of the surprise distribution runs to several percentage points over the following quarter. It is the anomaly that survived. Chapter 15 §15.3 gives the leading behavioral account — limited attention, identified off the fact that the immediate response is weaker and the drift stronger for Friday announcements and for days crowded with other firms' announcements. Nothing about the information differs on those days; only the number of eyes on it.

Note what drift does and does not show. It is a violation of semi-strong efficiency in the most literal sense: a mechanically applied trading rule using only public information earned an abnormal return. It is *not* an interesting instance of the joint hypothesis, because no model of expected returns generates quarter-long swings of that size in response to an accounting number. And it has decayed, as McLean and Pontiff's post-publication evidence (Chapter 6 §6.4) would predict. Efficiency is a margin that moves as capital arrives, not a property a market has or lacks.

***

## 7.4 Return Predictability and the Fama-Shiller Disagreement

**Short horizons.** The oldest finding in the field is that returns over short horizons are nearly unforecastable from past returns. Autocorrelations of daily and weekly index returns are small; the profitable exceptions — short-horizon reversal, momentum at intermediate horizons — are cross-sectional rather than aggregate, and the aggregate index is close to a martingale after adjustment for a small drift. Nothing in the last fifty years has overturned this, and it is the part of Fama's position that has never been in dispute.

**Long horizons.** The pattern changes with the horizon. Regress the return on the aggregate market over the next $$H$$ years on the dividend-price ratio, or on Shiller's cyclically adjusted price-earnings ratio (**CAPE**, price divided by a ten-year moving average of real earnings), and the coefficient is large, correctly signed, and grows with the horizon; the $$R^2$$ rises from under ten percent at one year to roughly thirty to forty percent at five to ten years in long US samples. Fama and French established the dividend-yield result in 1988 and 1989, Campbell and Shiller the earnings-ratio version in 1988. This is the central empirical fact of the last forty years of asset pricing, and Cochrane's 2011 presidential address is its statement of record. Figure 7.1 is the picture: the subsequent ten-year real return on the market against log CAPE, one point per month of the long US sample.

![Figure 7.1: CAPE and the next ten years](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-5ed4b20f25e6bc2853e7baf3bae58794d44839a3%2Ffig_07_01_cape_next_decade.png?alt=media)

**Figure 7.1: CAPE and the next ten years.** Subsequent annualized ten-year real total return against log CAPE, every month for which ten years of subsequent data exist, with 1929, 1966, 1982, 2000 and 2009 labelled and the fitted line and prediction band drawn — the classic scatter, specified in the data exercise Part A. Overlapping ten-year windows mean the fit's conventional standard error is wrong; Appendix A §A.1.3 supplies the correction and the effective sample size. *Source: Robert J. Shiller, ie\_data.xls (shillerdata.com); author's calculations.*

**Excess volatility.** Shiller's 1981 paper and LeRoy and Porter's, published independently the same year, made the same point from the other direction. If the price is the conditional expectation of the discounted stream of future dividends, then the *actual* price is the conditional expectation of the *perfect-foresight* price $$P^{\ast}$$ — the value computed ex post from realized dividends. A conditional expectation is less variable than the thing it forecasts, so $$\sigma(P) \le \sigma(P^{\ast})$$ must hold. It fails in the data, and by a wide margin: realized dividends are smooth, so the perfect-foresight price is smooth, while actual prices swing enormously. Shiller reported violations by a factor of four or five in variance.

![Figure 7.4: Excess volatility](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-57d17c7325e0f31be4c109b39f72d80e2bd1f4c0%2Ffig_07_04_excess_volatility.png?alt=media)

**Figure 7.4: Excess volatility.** Shiller's 1981 figure, rebuilt on his current file by the recipe this chapter's own data exercise gives: the perfect-foresight price is the discounted stream of realized real dividends at a constant real rate equal to the sample's average real total return, terminated with the realized real price in the final year, and both series are divided by the same exponential trend fitted to the actual price. The orange line is smooth because realized dividends are smooth. The blue line is not. Over 1871 to 2001 the standard deviation of the detrended actual price is 2.1 times that of the perfect-foresight price — a variance ratio of 4.5, which is the order of magnitude Shiller reported. The shaded years are excluded from that calculation, and the reason is worth stating because it is the construction's main weakness rather than a presentational choice: near the end of the sample the sum is dominated by its own terminal value rather than by realized dividends, so the perfect-foresight price inherits the actual price's volatility and the bound stops binding for an arithmetic reason that has nothing to do with markets. The other weaknesses — the constant discount rate, the stationarity assumption, the small-sample bias in the variance estimates — are the subject of the next paragraph and are not repaired by anything in this figure. *Source: Robert Shiller's ie\_data.xls; author's calculations.*

The bound as originally stated has real weaknesses, and Marsh and Merton and Kleidon pressed them: it assumes a constant discount rate, and it assumes the dividend process is stationary around a trend, which makes $$P^{\ast}$$ artificially smooth. Both objections are fair. What survived them, due to Campbell and Shiller and to Cochrane, matters more than the original statement. Take the accounting identity linking the log dividend-price ratio, the log return, and log dividend growth,

$$
dp\_t \approx r\_{t+1} - \Delta d\_{t+1} + \rho dp\_{t+1}
$$

where $$\rho \approx 0.96$$ is a linearization constant. Regress each term on $$dp\_t$$ and the identity forces the coefficients to satisfy $$\beta\_r - \beta\_d + \rho\phi = 1$$, where $$\phi$$ is the persistence of $$dp\_t$$. **Something must be forecastable.** If price-dividend ratios move at all, then either returns are predictable, or dividend growth is predictable, or the ratio is not persistent. In US data $$\phi \approx 0.94$$, so $$\rho\phi \approx 0.90$$: ten percent of the variation must be accounted for by the other two terms, and empirically essentially all of it sits in $$\beta\_r$$ and none in $$\beta\_d$$. Dividend growth is not forecastable from the dividend-price ratio.

That is the modern form of the result, and it collapses the two literatures into one. **Excess volatility and return predictability are the same fact seen from two angles.** Prices move too much relative to dividends *because* discount rates move, and discount rates moving is exactly what makes returns forecastable. The debate is no longer about whether either is true.

**The two readings.** What the debate is about is the nature of the moving discount rate.

The *rational* reading is Fama's, and Cochrane's. The expected return required to hold equity is countercyclical: high when consumption is low relative to habit (Chapter 5 §5.5), when intermediary capital is scarce (Chapter 19 §19.5), when the marginal holder's constraints bind (Chapter 16 §16.5). The regression measures a risk premium. It has real support: the same countercyclical pattern appears in corporate bond spreads, in Treasury term premia, in the variance risk premium, and in international equity markets, which is hard to attribute to one population's sentiment. And it is disciplined — the habit and long-run-risk models of Chapter 5 were built to deliver predictability jointly with the equity premium, and they do.

The *behavioral* reading is Shiller's, and Chapter 15 §15.3 makes its strongest case. Greenwood and Shleifer assembled six survey series of investor return expectations spanning several decades. The series agree with one another; they are strongly extrapolative, high after the market has risen and when prices are high relative to fundamentals; they do not forecast returns; and fund flows follow them, so respondents act on what they say. The model-implied expected return the predictive regression recovers moves *in the opposite direction* from the measured expectations of actual investors. Under the rational reading, investors at market peaks are demanding the least compensation they will ever demand; under the surveys, they are expecting the most they will ever expect. Both accounts fit the regression; only one fits the surveys.

**What settled and what did not.** Settled: the facts. Short-horizon aggregate returns are close to unforecastable; long-horizon returns are forecastable from valuation ratios; the price-dividend ratio's variation is almost entirely expected-return news; prices are more volatile than any constant-discount-rate model permits. Fama and Shiller agree on all of it, which is why the committee could honor both. Not settled: the interpretation, and it resists settlement because a time-varying rational risk premium and a time-varying expectational error have the same reduced form. Distinguishing them requires evidence outside the return series — survey expectations, cross-sectional patterns in who trades on what, the behavior of quantities as well as prices — which is why Part IV of this book exists.

Two statistical caveats belong on the record, because the regressions are weaker than their $$R^2$$ suggests. The predictor is highly persistent and its innovations are strongly negatively correlated with returns, which biases the estimated slope upward and the conventional standard error downward — the Stambaugh bias. And Goyal and Welch showed that most popular predictors, the dividend yield included, fail to beat a rolling historical mean out of sample over much of the postwar period. Predictability is a real feature of the data and a poor investment strategy, which is itself informative: if it were easy, §7.2 says it would not be there.

**UIP failure, in one paragraph.** The same shape of evidence appears in currencies. Uncovered interest parity says the interest-rate differential between two currencies should forecast depreciation of the high-rate currency by exactly enough to equalize expected returns. Regressions of exchange-rate changes on interest differentials produce coefficients not merely below the theoretical value of one but reliably of the *wrong sign*, so that high-interest currencies have tended to appreciate — the forward premium puzzle. This is predictability of an excess return from a public, contemporaneously observable variable, and it admits the same two readings: a time-varying currency risk premium compensating for crash risk and skewness, or a persistent expectational error. The mechanics, the return and Sharpe-ratio evidence, and the crash-risk interpretation are developed in *International Finance* Chapter 11 §11.5; the factor-model treatment of currency carry is Chapter 6 §6.6. The point here is only that the Fama-Shiller disagreement is not a fact about equities.

### ★ Why long-horizon $$R^2$$ rises

*Starred. A caution about reading long-horizon regressions as stronger evidence than short-horizon ones.*

The rise in $$R^2$$ with the horizon is often presented as though long-horizon predictability were a separate and stronger finding than one-year predictability. It is mostly the same finding restated.

Suppose one-period returns are generated by $$r\_{t+1} = \beta x\_t + \varepsilon\_{t+1}$$ with a persistent predictor, $$x\_{t+1} = \phi x\_t + u\_{t+1}$$. Then the $$H$$-period cumulative return has slope $$\beta(1-\phi^H)/(1-\phi)$$ on $$x\_t$$, which grows toward $$\beta/(1-\phi)$$, while the variance of the cumulative return grows roughly linearly in $$H$$. The ratio rises. Calibrate to $$\phi = 0.94$$, an annual return standard deviation of 18 percent, and a one-year $$R^2$$ of 0.08, with shocks taken independent for clarity:

**Table 7.3: Horizon and fit, from a one-year** $$R^2$$ **of 0.08**

| Horizon $$H$$ (years) | 1     | 2     | 3     | 5     | 7     | 10    | 15    | 20    |
| --------------------- | ----- | ----- | ----- | ----- | ----- | ----- | ----- | ----- |
| $$R^2\_H$$            | 0.080 | 0.140 | 0.185 | 0.245 | 0.278 | 0.300 | 0.297 | 0.277 |

*Source: Author's calculation from the stated data-generating process.*

Nothing has been added between the columns. One number — a one-year $$R^2$$ of eight percent — plus a persistence parameter generates the entire profile, including the eventual decline at very long horizons. Long-horizon regressions are not independent evidence; they are the short-horizon regression viewed through a magnifier, and they carry no additional degrees of freedom because the overlapping windows share almost all of their data. Their value is expositional and economic, not statistical: they make visible that a small, persistent movement in expected returns has large cumulative consequences for prices.

The chapter teaches a disagreement rather than a conclusion, so the line between the facts and their readings is worth marking explicitly, as Chapter 20 §20.5 does for the demand system.

**Table 7.4: The state of the evidence on efficiency**

| Claim                                                                                                                        | Status                                                                                                                                            |
| ---------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| Public information is in the price the same day                                                                              | **Established.** The event-study window is the sharpest test in the field, and it passes (§7.3)                                                   |
| Post-earnings-announcement drift is a genuine violation of semi-strong efficiency                                            | **Established**, and decaying — no model of expected returns produces quarter-long swings of that size                                            |
| Short-horizon aggregate returns are close to unforecastable                                                                  | **Established.** The one part of Fama's position never in dispute                                                                                 |
| Long-horizon returns are forecastable from valuation ratios, and price-dividend variation is almost all expected-return news | **Established.** Fama and Shiller agree on it; the Campbell-Shiller identity forces it                                                            |
| The moving discount rate is a rational risk premium rather than an expectational error                                       | **Contested**, and structurally so: the two have the same reduced form, and only evidence outside the return series separates them                |
| Predictability is exploitable out of sample                                                                                  | **Contested**, and the weight of it is negative — Stambaugh bias inflates the slope, and Goyal and Welch beat most predictors with a rolling mean |

*Source: Author's assessment of the literature discussed in §§7.3-7.4.*

***

## 7.5 Why Smart Money Does Not Correct Mispricing

The standard objection to every finding in §§7.3 and 7.4 is that if the pattern were real, someone would trade it away. This section begins the chapter's holder work that §7.7 closes: it asks who the someone is, and what constrains them. The full mechanics are developed later in the book, and the pointers here are exact.

**Noise-trader risk.** An arbitrageur who shorts an overpriced asset bears the fundamental risk of the asset *and* the risk that sentiment gets worse before it gets better, pushing the price further away and forcing a loss on a position that was correct. This risk is created by the noise traders, cannot be hedged, and limits the size of the rational position. Its consequences — mispricing survives in equilibrium, noise traders can earn higher expected returns than the arbitrageurs trading against them, and prices are excessively volatile relative to fundamentals — are worked through in **Chapter 15 §15.4**.

**Capital constraints and delegation.** Arbitrage is performed by specialists using other people's money, and the investors supplying that money can judge the manager only by realized returns. When a mispricing widens, the opportunity improves and the fund's capital contracts. **Chapter 15 §15.5** develops this as performance-based arbitrage: the supply of arbitrage capital shrinks precisely when the demand for it is highest, so the standard defense of efficiency — large mispricings attract capital — is backwards about the states in which it is invoked. **Chapter 16 §16.5** is this book's canonical statement of the general case, covering leverage ratios, haircuts, margin spirals, and fire sales, and it is where the balance-sheet mechanics live.

**Career risk.** One level down, inside the institution, a contrarian position that is wrong for two quarters ends a career in a way that a consensus position wrong for two quarters does not. **Chapter 16 §16.3** traces this through benchmark-hugging and the asymmetric flow-performance relationship.

**Mechanical limits.** Short-sale constraints are the hard case, and Chapter 3's opening episode is the worked exhibit: a twenty-billion-dollar mispricing, published in the financial press, surviving because the lendable float was tiny and the convergence date was contingent.

The relation to §7.2 is worth making explicit. Grossman and Stiglitz explain why the informed must be *compensated*; the limits-to-arbitrage literature explains why the compensation can be large, and largest exactly when mispricing is worst. Together they give the efficiently-inefficient equilibrium a mechanism as well as a name. And they answer part of Chapter 3 §3.7's question: where information is costly and arbitrage capital constrained, the price is set by whoever is unconstrained at the margin, which in stressed states is a much smaller and more specialized group than in calm ones.

> **Box 7.1 — Efficiency violations that persist: the bubble case**
>
> This section explains why a mispricing can survive being noticed. A bubble is the limiting case: an asset trades above any defensible present value of its cash flows, and does so for long enough that the deviation is the market's normal state rather than an interval between corrections.
>
> The category is harder to define than it sounds. In an infinite-horizon model a price can exceed fundamental value indefinitely provided the excess grows at the discount rate, since each holder expects to sell to the next at a price that compensates him for waiting — a **rational bubble**. Transversality conditions rule these out for most claims of interest, which is why the instructive cases are the ones that need neither an infinite horizon nor an irrational buyer. Allen and Gorton (1993) build one out of an agency relation alone: portfolio managers are paid a share of gains and cannot be charged for losses beyond their own capital, so buying an asset known to be overpriced is a rational use of somebody else's money, and the price sustains itself for as long as the delegation chain does.
>
> That construction is why the topic belongs here and not only in the behavioral chapter. It requires no mistaken beliefs — only a holder whose objective is not the one an efficiency test assumes, which is Chapter 3 §3.7's question asked of a market rather than of a security. **Chapter 15 §15.4** supplies the mechanisms that do run through beliefs, extrapolation and disagreement under short-sale constraints among them, and reports why the econometric detection problem is close to hopeless; the crisis-specific version belongs to the 2008 volume. What this chapter retains is the negative result: how long a mispricing persists is evidence about the constraints on its correctors, not about the conviction of its buyers.

***

## 7.6 Models as Market Infrastructure: Performativity

Everything above treats information as something that exists prior to the price and is then, more or less accurately, incorporated. A body of work argues that this is the wrong picture in an important class of cases, and it deserves a hearing in a chapter about what prices reflect.

The claim is Donald MacKenzie's, in *An Engine, Not a Camera* (2006). The title comes from Milton Friedman's description of theory as an engine for analyzing the world rather than a photographic reproduction of it, and MacKenzie's argument is that in finance the metaphor is literal. Financial models are not only descriptions of markets. They are used *by* markets — as pricing conventions, as risk systems, as the basis of products, as the language in which traders talk to one another — and their use changes the thing described.

**The canonical exhibit is Black-Scholes.** The Chicago Board Options Exchange opened on April 26, 1973, weeks before the Black-Scholes paper appeared in the *Journal of Political Economy*. At the outset, traded option prices departed substantially from the formula's values. Within a few years they did not. MacKenzie and Millo's account traces the mechanism concretely: traders carried sheets of model prices onto the floor, Texas Instruments sold a calculator with the formula built in, market makers quoted in implied volatility rather than in dollars, and the model's assumptions became the shared framework in which quotes were formed and disputes settled. The fit between model and market improved not because traders discovered the model was right but because they adopted it. MacKenzie calls this **Barnesian performativity**: use of a model makes the model more accurate.

This is not a debunking. The model was adopted because it worked — the replication argument it identifies is genuinely there. The observation is about the *direction of fit*. A theory used as a coordination device supplies a focal point, and a focal point can be self-validating in a way a description of planetary orbits cannot.

**Counterperformativity is the other half, and it is the more consequential half.** Use of a model can also undermine the conditions under which it holds. Leland, O'Brien and Rubinstein Associates, founded in 1981 by two Berkeley finance professors and a marketing partner, sold **portfolio insurance**: a synthetic protective put, manufactured by exactly the dynamic hedging argument that underlies Black-Scholes, sold to pension funds that wanted a floor without buying options that did not exist in size. The strategy sells the underlying as it falls and buys it as it rises. By 1987 something in the range of sixty to ninety billion dollars of US equity was covered by such programs, on the Brady Commission's later estimate. On October 19, 1987 the mechanical selling that the replication argument required arrived all at once, into a market with no one on the other side, and the Commission's report identified it as a principal amplifier of the decline.

The replication argument had assumed the hedger was small relative to the market and could trade without moving prices. That assumption is innocuous for one user and false for all of them together. The model's own success destroyed its premise. And the aftermath is visible in prices to this day: before October 1987, implied volatilities across strikes on index options were roughly flat, consistent with the lognormal assumption; afterward the skew appeared and never left. Chapter 8 §8.5 develops the smile as the empirical payoff of the Black-Scholes framework, and its Box 8.1 carries this episode.

Two further instances complete the pattern. **Index models and the passive complex**: the CAPM's "market portfolio" was a theoretical construct, and index funds were built to hold it. Trillions of dollars now track indices whose composition is decided by committees at a handful of firms, so the object the theory took as exogenous is now an administered artifact whose reconstitution dates move prices. Chapter 17 §§17.4 and 17.7 develop this, and Chapter 17's opening episode — Tesla's addition to the S\&P 500 in December 2020 — is the exhibit. **VaR as convention**: value-at-risk began as one bank's internal reporting system, became a public methodology in 1994, and was written into the Basel market-risk framework in 1996. Once a common risk measure binds for many institutions at once, a volatility spike tightens everyone's constraint simultaneously and produces the correlated selling that raises measured volatility further. Chapter 26 §26.3 treats VaR's construction and its failure modes, Box 26.1 traces the path from one bank's internal report to public infrastructure, and §26.6 turns the procyclicality into a pricing mechanism; Chapter 16 §16.5 has the balance-sheet version.

**A version of the argument that needs no sociology.** A neighboring literature inside corporate finance reaches the same structure from the other direction. Managers do not know their own investment opportunities perfectly, and the market price aggregates information they do not have, so they learn from it and invest accordingly. The price is then an input to the cash flows it is supposed to be forecasting. The feedback runs in both directions and is not benign: a speculator trading on his estimate of what a firm will do changes what the firm does, sometimes making the estimate correct; a short seller who drives the price down can make the financing the firm needed unobtainable, and thereby produce the distress he was betting on. This is performativity in a purely economic register, with a measurable object attached — corporate investment. Chapter 22 §22.4 owns the link from market valuations to corporate investment, and Chapter 12 §12.5 owns the issuance decision it runs through.

**What this does to the efficiency claim.** The serious epistemic point is not that models are wrong or that markets are irrational. It is that "prices reflect information" presupposes an information set existing independently of the price system. In the cases above, part of the information set *is* the price system's own conventions: what an option is worth depends on what everyone's model says it is worth, and what a portfolio's risk is depends on the risk measure everyone has agreed to use. The same holds one level down, where §7.5's arbitrageurs work: to call two assets substitutes, and a gap between their prices a mispricing, requires a theory of their similarity rather than an observation of it, and the ethnographic work on arbitrage desks finds traders treating that theory as the modeling choice it is. Efficiency then becomes a fixed point rather than a correspondence with an external fact — a condition of mutual consistency among the models in use. That is a weaker and more interesting property than the one Fama defined, and it has a corollary both Fama and Shiller can accept: when a model becomes infrastructure, its assumption failures stop being academic and become systemic, because everyone is wrong in the same way at the same moment.

One extension generalizes the argument past the case where the model is a good one. A model's worth as a shared language is separable from its worth as a representation, so a model everybody knows to describe the world badly may be retained anyway, because the alternative to one inaccurate common language is two accurate rival ones — which is worse for anyone who has to agree on a price with a counterparty. The strongest form of the performativity claim is therefore not that models make themselves true but that they make coordinated action possible; and when such a model fails, what unravels is not primarily its equations but the whole ordered pattern of behavior they had been coordinating.

Boundaries. The crisis-specific version of this argument — what the models did in 2007-2009, and the epistemics of ratings and correlation assumptions — belongs to the 2008 volume. The philosophy-of-science treatment, including when social-scientific knowledge is self-fulfilling and what that does to the notion of a scientific law, belongs to the Philosophy volume. Retained here is only the part that changes how a reader should interpret an efficiency test.

***

## 7.7 Who Holds the Information, and What Their Constraints Do to Its Price

Information is not a substance that prices absorb; it is a trading process, executed by holders, and two things this chapter has treated as primitives are what the rest of the book makes of that.

The first is *how* information gets into a price. Section 7.2 has informed traders submitting demand curves to a Walrasian auctioneer, which is a device, not a description. Real information enters through orders, hitting a limit book or a dealer's quote, and the dealer widens the spread precisely because some orders are informed. **Chapter 11 §§11.2-11.3** replaces the auctioneer with Glosten and Milgrom's sequential market maker and Kyle's strategic informed trader, and derives price impact, spreads, and the speed of price discovery as equilibrium objects. This chapter's $$\Psi$$ becomes the informativeness of order flow there; the noise supply $$z$$ becomes the liquidity trader whose orders let the informed hide.

The second is *who* the informed and uninformed traders are. This chapter's $$n$$ is a fraction of an anonymous population. Part IV replaces it with an ecology: households who under-diversify and extrapolate (**Chapters 14** and **15**), institutions whose mandates and capital charges determine what they may hold and when they must sell (**Chapter 16**), index funds whose demand is by construction insensitive to information (**Chapter 17**), hedge funds who are the arbitrage capital of §7.5 (**Chapter 18**), and dealers whose balance sheets set the price of immediacy (**Chapter 19**). Chapter 17 §17.6 applies this chapter's model directly: every dollar moving from an active to an index mandate is a dollar withdrawn from information production, the informed fraction cannot go to zero because the returns to information rise as it falls, and the open question is the level at which it stabilizes.

Efficiency is not a property a market has or lacks. It is the outcome of a contest between the people paying to learn things and the people trading for reasons that have nothing to do with learning, and the terms of that contest are set by who holds the claim and what constrains them.

In Chapter 1 §1.2's terms, this chapter has been an argument about how much of a price move the first reading can carry: news is what informed holders are paid to put into the price, and what they are not paid enough to put in shows up instead as a constraint binding on the correctors and as a model that has become convention.

***

## Elsewhere in the Series

* **Uncovered interest parity, the forward premium puzzle, and the carry trade** — *International Finance*, **Chapter 11 §11.5**, with return and Sharpe-ratio tables and the crash-risk and skewness interpretation. Section 7.4 uses UIP failure as one more instance of the predictability evidence and does not re-derive the parity conditions; Chapter 6 §6.6 owns the factor-model treatment of currencies.
* **What models did in the crisis** — the 2008 volume. The epistemics of ratings, correlation assumptions in structured credit, and model failure under stress are developed there; §7.6 retains only the general performativity argument.
* **Performativity as philosophy of science** — the Philosophy volume, which takes up self-fulfilling social-scientific knowledge in general. Section 7.6 is deliberately confined to what changes a reader's interpretation of an efficiency test.
* Within this book: **Chapter 11 §§11.2-11.4** owns microstructure and the mechanics of price impact, and reuses this section's $$z$$ and $$n$$; **Chapter 15 §§15.3-15.5** owns extrapolative expectations, noise-trader risk, and performance-based arbitrage; **Chapter 16 §16.5** owns the canonical constrained-capital statement; **Chapter 17 §17.6** applies §7.2 to the passive shift; **Chapter 8** owns the Black-Scholes framework and the smile; **Chapter 26** owns VaR; **Appendix A** owns event-study mechanics.

***

## Summary

1. **The 2013 prize honored a disagreement, not a conclusion.** Fama and Shiller agree on the evidence — short-horizon returns are close to unforecastable, long-horizon returns are forecastable from valuation ratios — and disagree entirely about what the forecasting coefficient measures. Hansen's place in the citation is the reason the disagreement is scientific: his methods make both readings testable restrictions on the same data.
2. **Efficiency is defined relative to an information set.** Weak form (past prices), semi-strong form (public information), strong form (all information, including private). The nesting means evidence against the weak form damages all three; nobody defends the strong form. The live hypothesis is the semi-strong form.
3. **The joint-hypothesis problem makes efficiency untestable in isolation.** "Abnormal return" requires a model of expected returns, so every rejection can be assigned either to inefficiency or to the model. Chapter 6's 1992 episode is the standing illustration: Fama assigned the death of beta to the model and published a replacement within a year.
4. **Grossman and Stiglitz proved that a fully revealing price cannot be an equilibrium.** If prices revealed everything, nobody would pay for information; if nobody paid, prices would reveal nothing. With costly information and no supply noise, competitive equilibrium does not exist.
5. **The equilibrium is interior and its informativeness is pinned by the cost of information.** With CARA-normal traders, price informativeness is a function of $$\Psi^{\ast} = \sigma\_y^2/\[(e^{2\tau\xi}-1)\sigma\_\varepsilon^2] - 1$$, which depends on the cost $$\xi$$, risk aversion, and the payoff variances — but not on the amount of noise trading. More noise trading raises the *number* of informed traders in exact proportion to $$\sigma\_z$$ and leaves price accuracy unchanged. Analysts are hired to offset noise, not to overcome it.
6. **Event studies dodge the joint hypothesis by shrinking the window.** Over a day, expected return is four basis points and the announcement effect is percentage points, so any plausible benchmark gives the same answer. The template is Fama, Fisher, Jensen and Roll (1969): market model, abnormal returns, cumulated and averaged. Mechanics in Appendix A.
7. **Post-earnings-announcement drift is the anomaly that survived.** Prices continue to move in the direction of an earnings surprise for weeks after the announcement — a literal violation of semi-strong efficiency that no model of expected returns can absorb. It has decayed as capital arrived, which is what efficiency looks like as a margin rather than a state.
8. **Excess volatility and return predictability are the same fact.** The Campbell-Shiller identity $$\beta\_r - \beta\_d + \rho\phi = 1$$ forces something to be forecastable if valuation ratios move at all. With $$\rho \approx 0.96$$ and $$\phi \approx 0.94$$, roughly ten percent of the ratio's variation must be explained, and empirically essentially all of it is expected-return news and none of it dividend-growth news. Shiller's 1981 bound, once purged of its stationarity and constant-discount-rate assumptions, is this result.
9. **The rational and behavioral readings have the same reduced form.** Countercyclical required returns and extrapolative expectations both make high prices forecast low returns. Chapter 5's habit model formalizes the first; Chapter 15 §15.3's survey evidence — expectations that are high at peaks, do not forecast returns, and are acted on through flows — is the sharpest evidence for the second, because it moves opposite to the model-implied expected return.
10. **The regressions are weaker than their** $$R^2$$ **suggests.** Long-horizon $$R^2$$ rises mechanically with predictor persistence: a one-year $$R^2$$ of 0.08 and $$\phi = 0.94$$ generate 0.30 at ten years with no additional information (Table 7.3). Stambaugh bias inflates the slope and deflates the standard error, and Goyal and Welch showed most predictors fail out of sample.
11. **Smart money does not correct mispricing because arbitrage is risky, delegated, and constrained.** Noise-trader risk (Chapter 15 §15.4), performance-based arbitrage (Chapter 15 §15.5), balance-sheet constraints and fire sales (Chapter 16 §16.5), career risk (Chapter 16 §16.3), and short-sale constraints each limit the correcting mechanism, and several of them bind hardest exactly when mispricing is largest.
12. **Models are market infrastructure, and that complicates the efficiency claim.** Black-Scholes fitted better after traders adopted it; portfolio insurance built on the same replication argument became a mechanical amplifier in October 1987 and left a volatility skew that never went away; index construction turned the theoretical market portfolio into an administered object; VaR turned a risk measure into a common constraint. When part of the information set is the price system's own conventions, efficiency is a fixed point rather than a correspondence with an external fact.

***

## Key Terms

* **Information set** $$\Omega\_t$$: The conditioning set relative to which efficiency is defined; efficiency claims are meaningless without one
* **Weak, semi-strong, strong form efficiency**: The nested cases in which $$\Omega\_t$$ is past prices, all public information, and all information including private
* **Joint-hypothesis problem**: The impossibility of testing efficiency separately from a model of equilibrium expected returns, since "abnormal return" presupposes a benchmark
* **Efficiently inefficient markets**: The modern resolution — prices are inefficient enough to compensate those who bear the costs of information and arbitrage, and efficient enough that no one else profits
* **Grossman-Stiglitz paradox**: A fully revealing price destroys the incentive to acquire the information it reveals, so informationally efficient prices cannot be an equilibrium when information is costly
* **Noise (supply) shock** $$z$$: Random asset supply unrelated to fundamentals; the ingredient that lets an equilibrium with costly information exist at all
* **Price informativeness**: The fraction of the variance of the learnable payoff component revealed by the price, $$\Psi/(1+\Psi)$$ in the model of §7.2
* **CARA-normal**: The tractability assumption pairing constant absolute risk aversion with normal payoffs, giving wealth-independent demands linear in the conditional mean
* **Rational expectations equilibrium**: An equilibrium in which uninformed traders correctly invert the mapping from fundamentals and noise to prices, and are limited by inseparability rather than by error
* **Event study**: A test of semi-strong efficiency using a narrow window around a dated event, so that misspecification of expected returns cannot account for the finding
* **Abnormal return** $$AR\_{i,t}$$: Realized return minus a benchmark, commonly the fitted market model $$\hat\alpha\_i + \hat\beta\_i R\_{M,t}$$
* **Cumulative abnormal return** $$CAR\_i(t\_1,t\_2)$$: The sum of abnormal returns across an event window
* **Post-earnings-announcement drift**: Continued price movement in the direction of an earnings surprise for weeks after the announcement; the most durable violation of semi-strong efficiency
* **CAPE**: Shiller's cyclically adjusted price-earnings ratio, price divided by a ten-year moving average of real earnings
* **Excess volatility**: The finding that prices vary more than the bound $$\sigma(P) \le \sigma(P^{\ast})$$ permits, where $$P^{\ast}$$ is the perfect-foresight price
* **Campbell-Shiller identity**: $$dp\_t \approx r\_{t+1} - \Delta d\_{t+1} + \rho dp\_{t+1}$$, which forces returns or dividend growth to be predictable if valuation ratios move
* **Stambaugh bias**: Upward bias in a predictive-regression slope when the predictor is persistent and its innovations correlate with returns
* **Performativity**: The use of a model changing the market it describes
* **Barnesian**: When use improves the model's fit
* **Counterperformative**: When use undermines the conditions the model requires

***

## Readings

### Required

* Fama, E. (1970). "Efficient Capital Markets: A Review of Theory and Empirical Work." *Journal of Finance* 25(2): 383-417. *The paper that made efficiency a testable proposition by defining it relative to an information set and stating the joint-hypothesis problem plainly; read §§I-II for the framework and skim the rest as a survey of what was known in 1970.*
* Grossman, S. and J. Stiglitz (1980). "On the Impossibility of Informationally Efficient Markets." *American Economic Review* 70(3): 393-408. *Short, and the whole argument is in it; read it for the free-entry condition and the non-existence result, which are what §7.2 reconstructs.*

### Recommended

* Asquith, P., M. Mikhail and A. Au (2005). "Information Content of Equity Analyst Reports." *Journal of Financial Economics* 75(2): 245-282. *Gives §7.2's cost of information some institutional content by looking at who actually produces it. Read it for the finding that the report's text moves prices beyond the recommendation and the earnings forecast.*
* Shiller, R. (1981). "Do Stock Prices Move Too Much to Be Justified by Subsequent Changes in Dividends?" *American Economic Review* 71(3): 421-436. *The volatility bound and its violation; read it alongside the standard objections, since the modern form of the result is the Campbell-Shiller identity rather than the original bound.*
* Fama, E. (1991). "Efficient Capital Markets II." *Journal of Finance* 46(5): 1575-1617. *Fama's own twenty-year reassessment, and the clearest statement anywhere of why he reads predictability as time-varying expected returns rather than as mispricing.*
* Cochrane, J. (2011). "Presidential Address: Discount Rates." *Journal of Finance* 66(4): 1047-1108. *The statement of record on predictability. Its organizing claim — that all valuation-ratio variation is discount-rate news, in every asset class — is the single most useful thing to take from this chapter.*
* Cochrane, J. (2008). "The Dog That Did Not Bark: A Defense of Return Predictability." *Review of Financial Studies* 21(4): 1533-1575. *The mechanical result behind §7.4: it is the failure of dividend growth to be predictable that forces returns to be predictable, so the two cannot both be dismissed as statistical artifacts. The 2011 address states the conclusion; this paper is the argument.*
* De Long, J. B., A. Shleifer, L. Summers and R. Waldmann (1990). "Noise Trader Risk in Financial Markets." *Journal of Political Economy* 98(4): 703-738. *The canonical formal answer to §7.5's question. Sentiment risk is created by the noise traders themselves, cannot be hedged, and bounds the size of the rational position — which is why mispricing survives in equilibrium rather than merely for a while.*
* Shleifer, A. and R. Vishny (1997). "The Limits of Arbitrage." *Journal of Finance* 52(1): 35-55. *The delegation version of the same argument, and the one §7.5 needs to reach Chapter 16's constrained capital: arbitrage capital contracts exactly when the opportunity improves.*
* Allen, F. and G. Gorton (1993). "Churning Bubbles." *Review of Economic Studies* 60(4): 813-836. *The source for Box 7.1. A bubble constructed from an agency relation with no irrational agent anywhere in it — a clean counterexample to the claim that persistent mispricing requires mistaken beliefs.*
* MacKenzie, D. (2006). *An Engine, Not a Camera: How Financial Models Shape Markets*. MIT Press. *The performativity argument with the archival work behind it; Chapters 5-6 on option pricing and Chapter 7 on 1987 are the ones §7.6 draws on.*
* MacKenzie, D. (2010). "Models as Coordination Devices." In M. Akrich, Y. Barthe, F. Muniesa and P. Mustar (eds.), *Débordements: Mélanges offerts à Michel Callon*. Paris: Presses des Mines, 299-302. *Four pages that extend §7.6 in the direction its last paragraphs lean: a model's value as a shared language is separable from its accuracy, so a model known to describe the world badly can be retained precisely because everyone speaks it.*
* Beunza, D. and D. Stark (2008). "Reflexive Modeling: The Social Calculus of the Arbitrageur." Working paper (SSRN). *The mechanism of §7.6 observed rather than inferred, on a live merger-arbitrage desk: the same model that lets a trader interpret an uncertain world also tells him what everyone else is looking at.*
* Beunza, D., I. Hardie and D. MacKenzie (2006). "A Price Is a Social Thing: Towards a Material Sociology of Arbitrage." *Organization Studies* 27(5): 721-745. *Four trading-floor ethnographies synthesized, and the source of §7.6's claim that arbitrage rests on a theory of the similarity between assets. The connection to §7.5 and to Chapter 15 §15.5 runs through that formulation.*
* Bond, P., A. Edmans and I. Goldstein (2012). "The Real Effects of Financial Markets." *Annual Review of Financial Economics* 4: 339-360. *The survey behind §7.6's economic-register paragraph: prices are inputs to the decisions they are forecasting, so the feedback loop the performativity literature describes has a corporate-investment counterpart with measurable outcomes.*

***

## Discussion Questions

1. **The joint hypothesis, applied.** A researcher reports that firms with high advertising expenditure earn abnormal returns of 4 percent per year relative to a five-factor benchmark, over forty years. Write down the two hypotheses her test is jointly testing. Then describe two research designs — one narrowing the window, one exploiting cross-sectional variation in the ease of arbitrage — that would shift weight between them, and say precisely what result from each would move your belief and in which direction. Is there any design that separates them fully?
2. **What the noise traders are.** Section 7.2's equilibrium requires $$\sigma\_z > 0$$, and the model is silent about who supplies the noise. Name three real sources of price-insensitive order flow in modern equity markets, and for each say whether it should be growing or shrinking as a share of volume. If price-insensitive flow is growing, the model predicts a larger information industry with unchanged price accuracy. Is that what we observe, and what evidence would you look at?
3. **Does performativity undermine efficiency or explain it?** One reading of §7.6 is that performativity is corrosive: if a model's fit improves because everyone uses it, the fit is not evidence that the model is true, and "prices reflect information" loses its meaning. Another reading is that performativity is a mechanism *for* efficiency: shared conventions are how a decentralized market coordinates on a common valuation, and coordination is what makes prices informative in the first place. Take a position. Then apply your criterion to a case where the two readings clearly diverge — the pre-1987 flat implied-volatility surface, or the reconstitution of a major index.
4. **Same reduced form, different world.** A time-varying risk premium and extrapolative expectations both predict that high prices forecast low returns. Suppose you had unlimited data on quantities as well as prices — every holder's position, every trade, every flow. Describe the pattern in the *holdings* data that each theory predicts. Which is more likely to be observed, and what would it take to convince you that both mechanisms operate in different states of the world?
5. **When does the anomaly stop being an anomaly?** Post-earnings-announcement drift has been documented since 1968 and has decayed. Value has been documented since 1992 and was absent from 2007 to 2020 (Chapter 6 §6.2). One reading is that publication attracts capital and removes mispricing; another is that a risk premium can go through long dry spells without being any less real. Design a test that distinguishes them using only the time series of the strategy's returns and the growth of capital tracking it. What is the identification problem you cannot solve, and does Chapter 6's post-publication decay evidence solve it?
6. **Information production as an industry.** Section 7.2 treats the acquisition of information as a cost parameter $$\xi$$ paid by an anonymous fraction of traders. In practice most of what reaches the market is produced by sell-side analysts who are paid by their employers rather than by the users of their research, and by firms deciding what to disclose and when. Rewrite $$\xi$$ as a description of that industry: who bears the cost, who captures the return, and what the model's free-entry condition corresponds to when entry means hiring an analyst rather than buying a signal. Then say what the model predicts should happen to price informativeness when a regulator mandates broader disclosure — and why the answer is not obviously an increase.

***

## Problems

**Problem 1 — A Grossman-Stiglitz equilibrium.** A risky claim pays $$x = y + \varepsilon$$ with $$\sigma\_y^2 = 9$$ and $$\sigma\_\varepsilon^2 = 1$$. Traders have absolute risk aversion $$\tau = 1$$; per-capita supply has variance $$\sigma\_z^2 = 1$$; information costs $$\xi = 0.5$$.

(a) Compute $$\mathrm{Var}\[x \mid p]$$ in equilibrium, the signal-to-noise ratio $$\Psi^{\ast}$$, the informed fraction $$n^{\ast}$$, and price informativeness. (b) Now quadruple the supply variance to $$\sigma\_z^2 = 4$$, holding everything else fixed. What happens to $$\Psi^{\ast}$$, to $$n^{\ast}$$, and to price informativeness? Explain the result in one sentence. (c) With $$\sigma\_z^2 = 4$$, what cost $$\xi$$ restores an interior equilibrium, and what are $$n^{\ast}$$ and price informativeness at $$\xi = 0.8$$? (d) At what cost does information production stop entirely? Show that your answer does not depend on $$\sigma\_z^2$$, and explain why.

**Problem 2 — An event study.** A firm's market model, fitted on the prior 250 trading days, gives $$\hat\alpha = 0.03$$ percent per day, $$\hat\beta = 0.9$$, and residual standard deviation $$1.2$$ percent per day. The firm announces an acquisition before the open on day 0.

| Day           | −2    | −1    | 0      | +1     | +2     |
| ------------- | ----- | ----- | ------ | ------ | ------ |
| Firm return   | 0.10% | 1.40% | −6.20% | −0.90% | 0.20%  |
| Market return | 0.20% | 0.60% | −0.30% | 0.50%  | −0.10% |

(a) Compute the abnormal return each day and $$CAR(-2,+2)$$, $$CAR(-1,0)$$, and $$CAR(+1,+2)$$. (b) Test $$CAR(-1,0)$$ against zero. What do you conclude about the market's assessment of the acquisition? (c) The day −1 abnormal return is large and positive. Give two interpretations, and say what additional data would distinguish them. (d) A colleague proposes extending the window to $$(-30, +30)$$ to capture the full effect. State the cost of doing so in terms of the joint-hypothesis problem, with a rough magnitude: at a benchmark expected return of 8 percent per year, how large is the expected return over a 61-day window, and how does that compare with the abnormal returns you computed?

**Problem 3 — Reading a long-horizon regression.** An analyst regresses the subsequent ten-year annualized real return on the S\&P 500 on the log CAPE, using overlapping annual observations from 1900 to 2015. She obtains a slope of $$-6.2$$ percentage points per unit of log CAPE, an $$R^2$$ of 0.38, and a $$t$$-statistic of 5.1 computed with conventional standard errors. She concludes that the market is inefficient and that the strategy should be traded.

(a) State three separate reasons the $$t$$-statistic overstates the evidence. Be specific about which one is the Stambaugh bias and which one concerns the effective number of independent observations. (b) Grant that the coefficient is real. Explain why it is not, by itself, evidence of inefficiency, and state what auxiliary hypothesis would have to be added to make it so. (c) Using the Campbell-Shiller identity with $$\rho = 0.96$$ and a dividend-price persistence of $$\phi = 0.94$$, and given a one-year return-forecasting slope of $$\beta\_r = 0.10$$, compute the implied dividend-growth slope $$\beta\_d$$ and the two long-run coefficients $$\beta\_r/(1-\rho\phi)$$ and $$\beta\_d/(1-\rho\phi)$$. Interpret. (d) What would have to be true of dividend growth for the same valuation-ratio variation to be consistent with *constant* expected returns?

**Problem 4 — The volatility bound.** Let $$P\_t$$ be the price and $$P^{\ast}\_t$$ the perfect-foresight price, the discounted stream of realized future dividends, so that $$P\_t = E\_t\[P^{\ast}\_t]$$ under the constant-discount-rate model.

(a) Show that $$P^{\ast}\_t = P\_t + u\_t$$ with $$E\_t\[u\_t] = 0$$, and that if the model is correct then $$\mathrm{Cov}(P\_t, u\_t) = 0$$. Derive the bound $$\sigma(P) \le \sigma(P^{\ast})$$. (b) Suppose $$\sigma(P)/\sigma(P^{\ast}) = 2.5$$ in a sample. By what factor is the bound violated in variance terms? (c) State the two auxiliary assumptions the bound requires and, for each, describe a plausible failure that would generate an apparent violation with no inefficiency at all. (d) Explain why the Campbell-Shiller identity of Problem 3 is a more robust statement of the same finding, and what it gives up in exchange for that robustness.

**Problem 5 ★ — The value of information and the corner.** Return to the setup of Problem 1(b): $$\sigma\_y^2 = 9$$, $$\sigma\_\varepsilon^2 = 1$$, $$\tau = 1$$, $$\sigma\_z^2 = 4$$, $$\xi = 0.5$$.

(a) The interior solution gives $$n^{\ast} > 1$$. Evaluate $$\mathrm{Var}\[x\mid p]$$ at the corner $$n = 1$$ and verify that the free-entry condition fails in the direction implying that every trader wants to be informed. (b) In this corner equilibrium, is the price fully revealing? Compute price informativeness and explain why it is bounded away from one even when everyone is informed. (c) Now let $$\sigma\_z^2 \to 0$$ with $$\xi = 0.5$$ held fixed. Show that neither $$n = 0$$ nor any $$n > 0$$ can be an equilibrium, and state in one sentence what the model is telling us about the concept of an informationally efficient price. (d) Chapter 17 §17.6 argues that the passive share of assets is self-limiting. Restate that argument in the notation of this problem, being explicit about which parameter the shift to passive is moving and in which direction, and say what the model does and does not tell you about the *level* at which the active share stabilizes.

***

## Selected Solutions

*Solutions to Problems 1 and 3 follow. Solutions to the remainder are in the instructor materials.*

**Problem 1.**

(a) The free-entry condition gives $$\mathrm{Var}\[x\mid p] = e^{2\tau\xi}\sigma\_\varepsilon^2 = e^{1} = 2.7183$$. Then

$$
\Psi^{\ast} = \frac{\sigma\_y^2}{(e^{2\tau\xi}-1)\sigma\_\varepsilon^2} - 1 = \frac{9}{1.7183} - 1 = 5.2378 - 1 = 4.2378
$$

and

$$
n^{\ast} = \frac{\tau\sigma\_z\sigma\_\varepsilon^2}{\sigma\_y}\sqrt{\Psi^{\ast}} = \frac{1 \times 1 \times 1}{3}\sqrt{4.2378} = \frac{2.0586}{3} = 0.686
$$

Price informativeness is $$\Psi^{\ast}/(1+\Psi^{\ast}) = 4.2378/5.2378 = 0.809$$. Roughly 69 percent of traders pay for information, and the price reveals about 81 percent of the variance of the learnable component. Check: $$\mathrm{Var}\[x\mid p] = 1 + 9/5.2378 = 2.718$$, as required.

(b) $$\Psi^{\ast}$$ is unchanged at 4.2378, because it depends only on $$\xi$$, $$\tau$$, $$\sigma\_y^2$$ and $$\sigma\_\varepsilon^2$$. Since $$n^{\ast} \propto \sigma\_z$$, the informed fraction doubles to $$2 \times 0.686 = 1.372$$, which exceeds one, so the equilibrium is at the corner $$n^{\ast} = 1$$. Price informativeness at that corner is $$\Psi = \sigma\_y^2/(\tau^2\sigma\_z^2\sigma\_\varepsilon^4) = 9/4 = 2.25$$, giving $$2.25/3.25 = 0.692$$. **More noise trading raises the number of informed traders one-for-one and, until the corner binds, leaves price accuracy exactly unchanged; past the corner, extra noise makes prices strictly less informative because there are no more traders left to hire.**

(c) An interior equilibrium requires $$n^{\ast} \le 1$$, that is $$\sqrt{\Psi^{\ast}} \le \sigma\_y/(\tau\sigma\_z\sigma\_\varepsilon^2) = 3/2$$, so $$\Psi^{\ast} \le 2.25$$, so $$e^{2\xi} \ge 1 + 9/3.25 = 3.769$$, so $$\xi \ge \tfrac{1}{2}\ln 3.769 = 0.663$$. At $$\xi = 0.8$$: $$e^{1.6} = 4.9530$$, $$\Psi^{\ast} = 9/3.9530 - 1 = 1.2767$$, $$n^{\ast} = (2/3)\sqrt{1.2767} = 0.753$$, and price informativeness is $$1.2767/2.2767 = 0.561$$.

(d) Information production stops when $$\Psi^{\ast} \le 0$$, that is when $$\sigma\_y^2 \le (e^{2\tau\xi}-1)\sigma\_\varepsilon^2$$, giving $$e^{2\xi} \ge 10$$ and $$\xi \ge \tfrac{1}{2}\ln 10 = 1.151$$. It does not depend on $$\sigma\_z^2$$ because the shutdown condition compares the value of information *when the price is uninformative* — which is $$\sigma\_y^2 + \sigma\_\varepsilon^2$$ against $$\sigma\_\varepsilon^2$$, a comparison in which supply noise plays no part — with its cost. Noise determines how many informed traders an equilibrium supports, never whether information is worth buying to the first trader.

**Problem 3.**

(a) Three reasons. First, **overlapping observations**: 116 annual observations of ten-year returns contain roughly 11 non-overlapping windows, so the effective sample is an order of magnitude smaller than $$T$$ and conventional standard errors are badly understated. Second, the **Stambaugh bias**: log CAPE is highly persistent and its innovations are strongly negatively correlated with returns, which biases the estimated slope away from zero in finite samples and shrinks the estimated standard error. Third, **specification search**: CAPE is one of many valuation ratios that have been tried on one history of US returns, and the survivor's coefficient is upward-biased for the same reason Chapter 6 §6.4's factor zoo is.

(b) A predictable expected return is exactly what a time-varying risk premium looks like. The regression measures $$E\_t\[r\_{t+H}]$$; it says nothing about whether that expectation is a required return or an error. To convert it into evidence of inefficiency one must add an auxiliary hypothesis that the required return is constant — or, more defensibly, evidence from outside the return series, such as the survey expectations of Chapter 15 §15.3.

(c) The identity gives $$\beta\_r - \beta\_d + \rho\phi = 1$$, so

$$
\beta\_d = \beta\_r + \rho\phi - 1 = 0.10 + 0.9024 - 1 = 0.0024
$$

With $$1 - \rho\phi = 0.0976$$, the long-run coefficients are $$\beta\_r/(1-\rho\phi) = 0.10/0.0976 = 1.025$$ and $$\beta\_d/(1-\rho\phi) = 0.0024/0.0976 = 0.025$$. They differ by exactly one, as the identity requires. The reading: about 102 percent of the variation in the dividend-price ratio is news about future returns and about 2 percent is news about future dividend growth. Valuation ratios move because discount rates move, not because anyone is forecasting cash flows.

(d) For constant expected returns, $$\beta\_r$$ would have to be zero, and the identity would then force $$\beta\_d = \rho\phi - 1 = -0.0976$$: a high dividend-price ratio would have to forecast *low* subsequent dividend growth, strongly enough to justify the low price. In the data the dividend-price ratio has essentially no forecasting power for dividend growth, and what little it has is often of the wrong sign. This is why the constant-discount-rate model is rejected without any appeal to behavior — it fails an accounting identity plus one regression.

***

## Data Exercise: CAPE, Dividend Yields, and What They Forecast

All parts run on free data.

**Part A — The classic scatter (free data: Shiller).** Download Robert Shiller's long-run monthly US series from his Yale website: S\&P Composite price, dividends, earnings, and the consumer price index, beginning 1871, together with his CAPE column.

1. Construct the annualized subsequent ten-year *real total return* on the index for every month for which ten years of subsequent data exist. Plot it against CAPE at the start of the period, with CAPE on a log scale. This is the classic scatter; label the points from 1929, 1966, 1982, 2000, and 2009 so you can see where the famous episodes sit.
2. Regress the subsequent ten-year annualized real return on $$\ln(\text{CAPE})$$ and report the slope, the $$R^2$$, and both a conventional and a Newey-West standard error with a lag length appropriate to the overlap. Report the ratio of the two standard errors and comment.
3. Compute the fitted value at today's CAPE. Then compute the width of a 90 percent prediction interval around it using the *effective* number of non-overlapping windows rather than the number of monthly observations. State plainly what the regression does and does not license you to say about the next decade.

**Part B — Dividend-yield predictability by subsample.**

1. Build annual real total returns and the annual dividend-price ratio from the same file. Regress the one-year return on the lagged dividend-price ratio over the full sample, and report the slope, $$t$$-statistic, and $$R^2$$.
2. Repeat separately for 1871-1945, 1946-1990, and 1991 to the present. Present the three as a table alongside the full sample. The instability is the finding, not a nuisance; describe it.
3. Repeat the exercise with *dividend growth* on the left-hand side instead of returns. Compare the two sets of coefficients against the Campbell-Shiller identity with $$\rho = 0.96$$ and your estimated persistence $$\phi$$. How close does the identity come to holding in each subsample?
4. Run Goyal and Welch's out-of-sample test: at each year $$t$$, forecast $$t+1$$'s return using only data through $$t$$, once with the dividend-yield regression and once with the historical mean, and compare cumulative squared forecast errors. Plot the difference over time. In which decades does the regression win?

**Part C — Excess volatility, reconstructed.**

1. Using the full dividend series and a constant real discount rate of your choosing (justify it from the sample's average real return), construct the perfect-foresight price $$P^{\ast}\_t$$ for every year for which enough subsequent dividends are available, terminating the sum with the realized price at the end of the sample.
2. Plot $$P\_t$$ and $$P^{\ast}\_t$$ on the same detrended axes and report both standard deviations. By what factor is Shiller's bound violated in your construction?
3. Recompute $$P^{\ast}\_t$$ under a *time-varying* discount rate given by your fitted regression from Part B. How much of the volatility gap closes? This is the rational reading of the Fama-Shiller disagreement, implemented.

**Part D ★ (if you have WRDS).** Replicate the post-earnings-announcement drift of §7.3 on CRSP and Compustat/IBES. Sort quarterly announcements into deciles by standardized unexpected earnings, compute size-and-book-to-market-adjusted cumulative abnormal returns from day +1 to day +60, and report the top-minus-bottom decile spread by decade from the 1970s to the present. Overlay an estimate of assets in quantitative equity strategies. State what your decay pattern can and cannot establish about whether the drift was mispricing.
