> For the complete documentation index, see [llms.txt](https://laurence-wilse-samson.gitbook.io/textbooks/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://laurence-wilse-samson.gitbook.io/textbooks/financial-economics-claims-prices-holders/part-ii-asset-pricing/chapter_06_factor_models.md).

# Chapter 6: Factor Models and the Cross-Section of Returns

*Part II: Asset Pricing — Financial Economics: Claims, Prices, and Holders*

***

## Opening Episode: The Death of Beta

The paper appeared in the June 1992 *Journal of Finance* under a title so plain it gave nothing away: "The Cross-Section of Expected Stock Returns." Eugene Fama and Kenneth French took every nonfinancial firm on the NYSE, AMEX and NASDAQ from 1963 to 1990, sorted them, and asked which of the variables the literature had proposed actually forecast average returns once the others were held fixed. Their answer was that two did — a firm's market capitalization and the ratio of its book equity to its market equity — and that one did not. The one that did not was beta.

This was not a marginal result. Chapter 4 §4.7 recorded that the empirical security market line had been known to be too flat since Black, Jensen and Scholes (1972) and Fama and MacBeth (1973). Fama and French went further: in their sample, once you controlled for size, the residual relation between beta and average return was not merely flat but statistically indistinguishable from nothing at all. The single number that two decades of textbooks, cost-of-capital calculations and performance evaluations had been built on could be dropped from the regression without loss.

The press understood the stakes before the profession had finished arguing about them. The *New York Times* ran Eric Berg's account of the working paper on 18 February 1992 under a headline about a study shaking confidence in the volatile-stock theory; the trade press followed through the spring. "Is Beta Dead?" had in fact been a magazine headline once already — *Institutional Investor* used it in July 1980, when the anomalies literature was young — but 1992 was when it stuck, because this time the author was Fama.

That last fact is the interesting one. Fama did not treat the result as a defeat, and it is worth being precise about why, because students routinely file it wrong. His position, then and since, was that what had died was the *Sharpe-Lintner CAPM* — beta as the sole variable explaining average returns — and that this was a finding about a model of expected returns, not about whether prices reflect information. Chapter 7 develops the logic in full; the short version is the joint-hypothesis problem. A test of market efficiency is always a test of efficiency *plus* a model of equilibrium expected returns, so any rejection can be assigned to either. Fama assigned it to the model. Within a year he and French had proposed a replacement — a three-factor equilibrium risk model, published in the *Journal of Financial Economics* in 1993 — in which size and book-to-market are not anomalies at all but priced state variables that the CAPM had left out. The 1992 paper, on this reading, is not the demolition of rational asset pricing. It is rational asset pricing doing maintenance on itself.

Whether the maintenance succeeded is the subject of this chapter, and the honest answer after thirty years is: partly, and less than its authors hoped. A pattern in average returns admits three readings, and the profession has never fully separated them.

It may be **compensation for risk**: the sorted portfolios load on something investors genuinely fear, and the spread in returns is the price of bearing it. It may be **mispricing**: the pattern is real, and it persists because the capital that would trade it away is limited, constrained, or absent — the mechanism Chapter 4 §4.9 already used to explain the flat line. Or it may be **an artifact of searching**: with thousands of candidate variables and one history of returns, some will fit by construction, and the ones that fit are the ones that get published.

The three are not mutually exclusive, they imply very different things about whether a strategy will keep working, and — this is the chapter's organizing difficulty — they are extremely hard to tell apart with the evidence available. Sections 6.1 to 6.3 build the machinery; §§6.4 and 6.5 confront what happens when it is run thousands of times; §6.6 takes it out of equities, where the same sorts price currencies and commodities. Section 6.8 asks the question this book always asks: whose capital stands behind these returns, and what happens when it moves.

***

## 6.1 The APT and the Logic of Factor Pricing

Chapter 4 §4.6 closed with an observation offered as a throwaway: nothing in the SDF derivation of the security market line required the single factor to be the market. Replace $$R\_M$$ with a vector of factors, $$m = a - \sum\_k b\_k f\_k$$, and the same three lines deliver a multi-beta pricing relation. This section supplies the economics that makes such a relation more than an algebraic possibility.

Stephen Ross's **arbitrage pricing theory** (1976) begins not with preferences but with a statistical assumption about returns. Suppose asset returns are generated by

$$
r\_i = E\[r\_i] + \sum\_{k=1}^{K} \beta\_{i,k} \big(f\_k - E\[f\_k]\big) + \varepsilon\_i
$$

where $$f\_k$$ is the realization of factor $$k$$ and $$f\_k - E\[f\_k]$$ its unexpected component, $$\beta\_{i,k}$$ is asset $$i$$'s loading on it, and $$\varepsilon\_i$$ is an idiosyncratic disturbance with $$E\[\varepsilon\_i] = 0$$, uncorrelated with the factors and only weakly correlated across assets. This is a **factor structure**: it says that the common movement in a large cross-section of returns is spanned by $$K$$ sources, with $$K$$ small relative to the number of assets.

Now form a portfolio with weights $$w$$ and let the number of assets grow, holding each weight small — a **well-diversified portfolio**, in which no single position is large. Its residual is $$\sum\_i w\_i \varepsilon\_i$$, and Chapter 4 §4.1's arithmetic applies unchanged: the residual variance falls at rate $$1/N$$ and vanishes in the limit. A well-diversified portfolio's return is therefore, to a close approximation, *deterministic given the factors*:

$$
r\_p = E\[r\_p] + \sum\_{k} \beta\_{p,k} \big(f\_k - E\[f\_k]\big)
$$

Everything follows from that sentence. Two well-diversified portfolios with identical factor loadings have identical realized returns in every state; if their expected returns differed, buying one and shorting the other would be a riskless money machine requiring no capital. More generally, expected returns must lie on a hyperplane in loading space:

$$
\boxed{E\[r\_i] - r\_f = \sum\_{k=1}^{K} \beta\_{i,k}\lambda\_k}
$$

with $$\lambda\_k$$ the **price of risk** for factor $$k$$ — the expected excess return on a portfolio with unit loading on $$f\_k$$ and zero loading on everything else. That portfolio is a **factor-mimicking portfolio**, and when the factor is itself constructed as a traded excess return, $$\lambda\_k = E\[f\_k]$$ directly. This is why the empirical literature builds its factors as portfolio returns rather than as macroeconomic series: doing so makes the price of risk something you can read off an average rather than something you must estimate from a cross-section.

A worked case makes the force of the argument visible. Take a single-factor economy with three well-diversified portfolios:

| Portfolio | Loading $$\beta$$ | Expected return |
| --------- | ----------------- | --------------- |
| A         | 0.5               | 6%              |
| B         | 1.5               | 12%             |
| C         | 1.0               | 8%              |

*Source: Author's construction.*

Portfolios A and B fix the line: the slope is $$(12 - 6)/(1.5 - 0.5) = 6$$ percent per unit of loading, and the zero-beta intercept is $$6 - 0.5(6) = 3$$ percent. A portfolio with loading 1.0 must therefore be worth $$3 + 6 = 9$$ percent. Portfolio C offers 8. So hold half of A and half of B — loading exactly 1.0, expected return 9 percent — and short C. The combined position requires no net capital, carries no factor exposure, and has no residual risk because all three legs are well diversified. It pays 1 percent of the gross position, for certain, every year. At $100 million a side that is $1 million a year out of nothing, and the position scales: an arbitrageur will take it until C's price rises and its expected return falls to 9.

What has been assumed here is remarkably little. No utility function, no market clearing, no mean-variance investor, no observable market portfolio and therefore no exposure to Roll's critique (Chapter 4 §4.7). The APT delivers a multi-beta security market line from a statistical assumption plus the impossibility of a free lunch. That is its power.

Its weakness is the exact complement. **The theory does not say what the factors are.** It does not say how many there are, what economic quantity they correspond to, or even the sign of $$\lambda\_k$$ — the CAPM at least tells you the price of market risk is $$\mu\_M - r\_f$$, a number you can compute. The APT says that *if* returns have a factor structure *then* expected returns are linear in the loadings, and leaves the antecedent to be filled in empirically. A theory that licenses any factor found in the data and disciplines none of them will be taken up enthusiastically. Section 6.4 is the bill.

Between the theory and the zoo there was a middle period worth recording, because it read the APT's silence about the factors as an instruction rather than a permission. Chen, Roll and Ross (1986) argued that the factors ought to be macroeconomic — innovations in industrial production, in expected and in unexpected inflation, in the term spread and in the default spread — on the ground that these are the state variables a discount factor should plausibly depend on, and they tested whether loadings on them were priced in the cross-section of US equities. The approach has one large merit and one large defect. The merit is that the factors are named in advance and are not built out of the returns they are asked to explain, so the exercise is a test rather than a fit. The defect is statistical: macroeconomic series are measured at low frequency, with error, and subject to revision, so the estimated loadings are noisy and the tests have little power against the alternatives that matter. That is why the literature migrated toward factors constructed as portfolio returns, and why §6.4's census counts characteristics rather than macroeconomic risks — a substitution made on statistical grounds whose economic cost is rarely stated.

### ★ How near is "near"?

*Starred. The precision of the arbitrage argument, and what it can be held to.*

The argument above is exact only in the limit. For a finite cross-section the conclusion is weaker, and the weakness matters. Ross's bound, sharpened by Huberman (1982), says that the sum of squared pricing errors across assets is bounded: $$\sum\_i \alpha\_i^2 < \infty$$, where $$\alpha\_i$$ is asset $$i$$'s deviation from the hyperplane. Only finitely many assets can be badly mispriced, but *any given asset* may be one of them, and the bound places no restriction on how large a single $$\alpha\_i$$ can be. The APT is a statement about the cross-section as a whole, not a valuation model for a security.

Two further gaps are worth naming. First, the factors are identified only up to rotation: if $$K$$ portfolios span the common variation, so does any nonsingular linear combination of them, so "the factors" is never a well-posed question — only "a spanning set" is. Second, as Shanken (1982) pressed, an approximate relation with an unspecified approximation error is difficult to reject, which means the APT is more nearly a framework than a hypothesis. Chapters 5 and 20 take the opposite route, deriving $$m$$ from something economic and accepting the sharper rejections that follow.

***

## 6.2 The Canonical Cross-Section

The literature the APT licensed converged, over thirty years, on a short list of characteristics that sort stocks into portfolios with reliably different average returns. Five have survived hard enough scrutiny to be worth teaching. Table 6.1 states them, Figure 6.1 gives their realized record decade by decade, and the discussion does the work.

**Table 6.1: The canonical equity factors**

| Factor                  | Sort variable                              | Construction                      | Approximate long-run premium          | Source paper                                       |
| ----------------------- | ------------------------------------------ | --------------------------------- | ------------------------------------- | -------------------------------------------------- |
| **SMB** (size)          | Market capitalization at end of June       | Small minus big                   | \~2-3% per year; near zero since 1980 | Banz (1981)                                        |
| **HML** (value)         | Book equity / market equity                | High minus low                    | \~4% per year; roughly flat 2007-2020 | Rosenberg-Reid-Lanstein (1985); Fama-French (1992) |
| **UMD** (momentum)      | Return over months $$t{-}12$$ to $$t{-}2$$ | Up minus down, rebalanced monthly | \~8% per year; severe crashes         | Jegadeesh-Titman (1993)                            |
| **RMW** (profitability) | Operating profits / book equity            | Robust minus weak                 | \~3% per year                         | Novy-Marx (2013)                                   |
| **CMA** (investment)    | Growth in total assets                     | Conservative minus aggressive     | \~3% per year                         | Cooper-Gulen-Schill (2008); Fama-French (2015)     |

*Source: Approximate magnitudes from the original papers and Kenneth French's data library, rounded and stated in simple annual terms. These are order-of-magnitude figures only: realized premia depend heavily on the sample period, on value- versus equal-weighting, and on whether NYSE or all-stock breakpoints are used. Section 6.4 is about why such numbers should be read with suspicion.*

![Figure 6.1: Factor premia by decade](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-fa0162a46c57557c1ed26e785a28755e48586329%2Ffig_06_01_factor_premia_by_decade.png?alt=media)

**Figure 6.1: Factor premia by decade.** SMB, HML, UMD, RMW and CMA average returns and t-statistics, decade by decade — the exact table the data exercise Part A step 1 asks for, and the visual form of §6.2's "near zero since 1980" and "roughly flat 2007-2020". Each cell carries the annualized average return, twelve times the mean monthly return, with the monthly-mean t-statistic in parentheses below it; the t-statistic is on the monthly mean and is not annualized. Decades are calendar decades on the French monthly key — the 1930s column is 193001-193912 — so the first and last columns are partial and carry their spans, and RMW and CMA's 1960s cell covers 196307-196912. Cells are shaded by the premium, saturating at plus or minus ten percent a year; the colour key runs along the foot of the figure. *Source: Kenneth R. French data library: SMB and HML from 1926-07, UMD from 1927-01, RMW and CMA from 1963-07, all monthly through 2026-06; author's calculations.*

**Construction, once, because it generalizes.** Fama and French (1993) build SMB and HML from a two-by-three double sort. Each June, stocks are split at the NYSE median market capitalization into "small" and "big", and independently into three book-to-market groups at the 30th and 70th NYSE percentiles: "value", "neutral", "growth". The intersection gives six value-weighted portfolios. SMB is the average return on the three small portfolios minus the average on the three big ones; HML is the average on the two value portfolios minus the average on the two growth ones. Each leg averages *across* the other characteristic, so SMB is a size spread holding book-to-market roughly fixed and HML is a value spread holding size roughly fixed. RMW and CMA are built the same way with profitability and asset growth in place of book-to-market. UMD replaces the annual rebalance with a monthly one, sorting on the return from twelve months ago to two months ago — the one-month gap excludes the short-horizon reversal that would otherwise contaminate the signal.

**Size** was the first anomaly and is the weakest survivor. Banz (1981) found that small firms earned more than their betas justified; the effect was large in the 1926-1980 data and has been close to zero in the United States since roughly the year it was published — a fact that recurs in §6.4 with an explanation attached. Much of the historical premium sat in January, in the smallest microcaps, and in stocks sensitive to delisting-bias corrections — and those are the same stocks that are expensive to trade, which is why Chapter 11 §11.4 reads the size premium and the liquidity premium as measurements taken on nearly the same cross-section. Its best modern defense is as a conditioning variable rather than a standalone premium: the profitability, investment and value premia are all substantially larger among small firms.

**Value** carries the field's deepest disagreement, and it is exactly the risk-versus-mispricing one. The risk reading, Fama and French's, is that a high book-to-market ratio identifies firms whose market value has been beaten down — distressed, levered, holding assets that are hard to redeploy — and that such firms do badly precisely when investors can least afford it. Zhang (2005) supplies the production-side mechanism: value firms carry unproductive capital they cannot easily disinvest, so their cash flows are more exposed to bad aggregate times. The mispricing reading, Lakonishok, Shleifer and Vishny's (1994), is that investors extrapolate recent growth too far, overpaying for glamour and underpaying for the dull, so the value spread is the correction of a systematic forecasting error. Both predict the same sort. They differ on what happens next: risk premia persist, corrected errors do not.

A third possibility is neither. Book equity capitalizes factories and expenses research and brands. As the corporate capital stock has shifted toward intangibles, book-to-market has become a progressively worse measure of what it was meant to proxy — one candidate explanation for the value premium's long absence between the financial crisis and 2020, and for its sharp return in 2021-2022. That is a measurement story rather than an economic one, and it is the kind of story a factor model has no way of telling on its own.

![Figure 6.2: Cumulative factor returns, 1927-present](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-2fc5ef3c8480c471d8e4a90aa6aaeddda55f4c34%2Ffig_06_02_cumulative_factor_returns.png?alt=media)

**Figure 6.2: Cumulative factor returns, 1927-present.** A dollar in each factor, on a log scale so that equal vertical distances are equal proportional gains and a straight line is a constant compound rate. Three readings the summary statistics of Table 6.1 cannot give. The premia are not steady: every one of these lines has multi-decade stretches of nothing, and the size factor's entire cumulative gain was earned before 1984. The shaded band is the value premium's absence, and drawn rather than tabulated it is not a pause but a drawdown — HML gives back about thirty percent over nineteen years, which is longer than most careers and longer than the sample that established the premium in the first place. And the market line is the one to measure the others against: over the full century it compounds faster than any of the long-short portfolios, which is worth remembering in a chapter about the cross-section. A factor premium is a spread between two portfolios, and a spread can be real, large, and still smaller than simply owning the market. *Source: Kenneth R. French data library, the three-factor, five-factor and momentum files, monthly; author's calculations.*

**Momentum** is the awkward one. Jegadeesh and Titman (1993) showed that buying stocks that have risen over the past six to twelve months and shorting those that have fallen earns roughly one percent a month, and the result has held across subsequent decades, across forty-odd national equity markets, and — as §6.6 shows — across asset classes that share no investors and no accounting. It is the most robust pattern in the empirical literature and has no comfortable seat in any equilibrium model. Fama and French excluded it from the three-factor model in 1993 and from the five-factor model in 2015 while conceding it is the largest failure of both; Carhart (1997) simply added it, and the four-factor model is what most practitioners use.

The discomfort is structural. Size and book-to-market are *stable* firm characteristics that could plausibly proxy for persistent exposure to some state variable. Momentum is a property of the recent price path: membership in the winner portfolio turns over within months, so any risk story must explain why the *same firm* is riskier this quarter and not next. The behavioral readings — gradual diffusion of information, underreaction to news, reluctance to realize losses (Chapter 15) — sit more naturally, and post-earnings-announcement drift is a close relative pointing the same way. Momentum also carries turnover of several hundred percent a year, and therefore a premium substantially smaller after realistic trading costs than the headline number.

And it crashes. Daniel and Moskowitz (2016) document that the momentum portfolio's return distribution is severely left-skewed, with losses concentrated in sharp market rebounds following steep declines. The mechanism is a conditional beta: after a crash the loser leg is loaded with high-beta, high-leverage survivors, so the short side of a winner-minus-loser position has enormous market exposure exactly when the market turns. In two months of 1932 the strategy lost roughly ninety percent of its value; between March and May of 2009, as the market bottomed and the most damaged financial stocks multiplied off their lows, it lost on the order of seventy percent. A factor whose payoff resembles a written call on the market is not obviously a compensation-for-risk story of the usual kind, and not obviously not one either — writing crash insurance earns a premium in every model in this book.

Figure 6.5 takes the 2009 episode apart leg by leg.

![Figure 6.5: The momentum crash of 2009](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-f7febc4cc74d6dbd5988616cf358b5adae28dbfd%2Ffig_06_05_the_momentum_crash.png?alt=media)

**Figure 6.5: The momentum crash of 2009.** March to May 2009, the three months in which the market bottomed and turned. The winner decile did nothing wrong: it rose 7 percent, roughly in line with a rising market. The loser decile — the most damaged financial and cyclical names, priced for bankruptcy in February — rose 159 percent. A strategy long the first and short the second lost 74 percent in three months, and every dollar of that loss came from the short leg. This is why §6.2 calls momentum's payoff a written call. After a crash the loser portfolio is a basket of high-beta, high-leverage survivors, so the short side carries enormous market exposure exactly when the market turns, and the strategy's beta is not a constant to be estimated but a state variable that goes sharply negative in precisely the state that hurts. Note where the loss sits in the cycle: not at the bottom, but in the recovery. A risk measure calibrated on the drawdown would have shown momentum performing well through February and would have said nothing at all about March. *Source: Kenneth R. French data library, Portfolios Formed on Prior 2-12 Returns and the three-factor file, monthly; author's calculations. The strategy drawn is the extreme-decile winner-minus-loser portfolio, which is what Section 6.2's magnitudes refer to.*

**Profitability and investment** joined the canon last and arrived with a theory attached, which is why they are treated together. Novy-Marx (2013) showed that operating profitability forecasts returns positively *after* controlling for book-to-market — the two work best in combination, since a profitable firm at a high book-to-market is cheap relative to its earnings rather than merely cheap. Cooper, Gulen and Schill (2008) and a related literature showed that firms expanding their asset base aggressively subsequently underperform. Fama and French (2015) folded both into the five-factor model, with a consequence they reported honestly: once RMW and CMA are included, HML becomes largely **redundant** in US data, its five-factor alpha indistinguishable from zero. The value factor, having killed beta, was in turn absorbed.

The rationalization of RMW and CMA is not an APT argument at all. It comes from the firm's side and is developed in **Chapter 22 §22.4**. A firm's market value is the present value of its cash flows; hold value fixed and raise expected profitability, and the discount rate applied to those cash flows must be higher; hold profitability fixed and raise investment, and the discount rate must be lower, since firms invest more when capital is cheap. Those are RMW and CMA read off a valuation identity plus optimal investment — the *investment CAPM* of Hou, Xue and Zhang (2015) and the q-theoretic tradition. It is the best answer this literature has produced to the "which factors?" question the APT left open, and notice what kind of answer it is: not a claim about what investors fear, but a claim about what firms do.

***

## 6.3 How the Evidence Is Made: Sorts and Fama-MacBeth

Every empirical claim in §6.2 rests on one of two procedures. Both are simple enough to state here; their full mechanics, including standard errors, the treatment of unbalanced panels, and the small-sample corrections, are in **Appendix A**.

**Portfolio sorts.** Rank every stock on a characteristic, cut the ranking into groups — deciles are conventional, quintiles common — form a value-weighted portfolio from each group, hold for a period, rebalance, and report the average return on the top group minus the bottom. Report also the alpha of that spread against whatever factor model you are trying to beat.

Three features explain the method's dominance. It imposes no functional form: if the relation between the characteristic and expected return is monotone but wildly nonlinear, the sort finds it and a linear regression may not. It is robust to outliers, because only ranks matter. And it produces a *tradable portfolio*, which means the resulting number is an implementable return rather than a regression coefficient — and can therefore be confronted with transaction costs, capacity, and the crowding of §6.8.

**Double sorts** exist because characteristics are correlated. Value stocks are on average smaller than growth stocks, so a univariate book-to-market sort is partly a size sort, and a naive reader cannot tell which characteristic is doing the work. An **independent** double sort assigns each stock to a size group and, separately, to a value group, and reports the value spread within each size group — this is the Fama-French construction of §6.2, and it answers "does value work holding size fixed?". A **conditional** (sequential) sort first splits on size, then forms value groups *within* each size bucket; this guarantees balanced cells when the characteristics are strongly dependent, at the cost of making the second sort's interpretation conditional on the first. Which to use depends on which characteristic you regard as the control. Both beat a univariate sort, and neither can handle more than about three dimensions before the cells empty — a limitation that §6.5's methods exist to remove.

**Fama-MacBeth regressions.** The second procedure, introduced by Fama and MacBeth (1973) and still the workhorse, estimates the prices of risk directly. It runs in two passes.

The **first pass** is a set of time-series regressions, one per test asset, of excess returns on the factors:

$$
r\_{i,t} - r\_{f,t} = a\_i + \sum\_{k=1}^{K} \beta\_{i,k} f\_{k,t} + \varepsilon\_{i,t}
$$

which yields the loadings $$\hat\beta\_{i,k}$$. These are estimated on a prior window if the design is to be implementable in real time, or on the full sample if the object is a description.

The **second pass** runs a *separate cross-sectional regression in each period* $$t$$, of that period's excess returns on the loadings estimated in the first pass:

$$
r\_{i,t} - r\_{f,t} = \hat\alpha\_t + \sum\_{k=1}^{K} \hat\lambda\_{k,t}\hat\beta\_{i,k} + u\_{i,t}
$$

This delivers a time series of fitted intercepts $$\hat\alpha\_t$$ and fitted prices of risk $$\hat\lambda\_{k,t}$$ — the same hatted objects Chapter 4 §4.9 used to describe the flattened security market line, now estimated period by period. The estimates of interest are their time-series averages, and the standard errors come from the time-series variation:

$$
\hat\lambda\_k = \frac{1}{T}\sum\_{t=1}^{T}\hat\lambda\_{k,t}, \qquad \mathrm{se}(\hat\lambda\_k) = \frac{s(\hat\lambda\_{k,t})}{\sqrt{T}}
$$

The trick is worth appreciating. Returns in any given month are enormously cross-sectionally correlated — everything moves with the market — which makes standard errors from a single pooled regression badly wrong. By estimating the cross-sectional relation once per period and then treating the resulting sequence as a sample of $$T$$ independent draws, Fama and MacBeth sidestep the cross-sectional dependence entirely, at the cost of assuming the coefficients are not autocorrelated through time. Under the CAPM, $$\hat\alpha$$ should average to zero and $$\hat\lambda\_M$$ should average to the market premium; Chapter 4 §4.9 reported what actually happens, which is a positive intercept and a slope well below it.

Two caveats travel with the method and neither is optional. The loadings entering the second pass are *estimated*, so the second-pass coefficients suffer errors-in-variables attenuation, biasing $$\hat\lambda$$ toward zero and the intercept upward; Shanken (1992) supplies the correction, and it is in Appendix A. And the standard defense — use portfolios rather than individual stocks as test assets, since portfolio betas are estimated far more precisely — has a cost that took the profession thirty years to take seriously: sorting assets into portfolios on the basis of the characteristic under test destroys much of the cross-sectional spread in *other* dimensions and can manufacture apparent factor structure where none exists. Problem 2 works the second pass on a small panel by hand.

***

## 6.4 The Factor Zoo

By 2012 the literature had published, by Harvey, Liu and Zhu's count, 316 distinct factors claimed to explain the cross-section, with the discovery rate accelerating; the number is now well past four hundred. Since the number of genuinely independent sources of common variation in equity returns is plainly not four hundred, the diagnosis is straightforward: the literature has run an enormous number of tests against a single history of returns while reporting each as though it were the only test conducted.

The statistics are unforgiving. Under the conventional threshold of $$|t| > 1.96$$, a single test of a true null rejects five percent of the time. Run 316 such tests and you expect **15.8** false discoveries, and the probability that at least one true null clears the bar is, to four decimal places, 1.0000. If the hurdle is meant to control the family-wise error rate at five percent, Bonferroni's correction requires a per-test significance level of $$0.05/316$$, which is a two-sided cutoff of $$|t| > 3.78$$.

**Table 6.2: What a** $$t$$**-statistic has to clear**

| Number of tests $$J$$ | Bonferroni per-test level | Two-sided cutoff $$\lvert t\rvert$$ | Expected false positives at 5% |
| --------------------- | ------------------------- | ----------------------------------- | ------------------------------ |
| 1                     | 0.050                     | 1.96                                | 0.05                           |
| 10                    | 0.005                     | 2.81                                | 0.5                            |
| 100                   | 0.0005                    | 3.48                                | 5.0                            |
| 316                   | 0.000158                  | 3.78                                | 15.8                           |
| 400                   | 0.000125                  | 3.84                                | 20.0                           |

*Source: Author's calculation. Cutoffs are two-sided normal quantiles at the Bonferroni per-test level, a family-wise size of 0.05 divided by the number of tests; the count of 316 tests is Harvey, Liu and Zhu's (2016) tally of published factors through 2012.*

Harvey, Liu and Zhu recommend a working hurdle of about $$|t| > 3.0$$, softer than Bonferroni because they use a false-discovery-rate criterion and because they attempt to adjust for the tests that were run and never published. Their conclusion is the one to carry away: **most published cross-sectional findings, evaluated at a hurdle appropriate to the number of tests behind them, would not be declared significant.** The problem is worse than the arithmetic suggests, because the 316 are the survivors. The denominator is not the number of published factors but the number of specifications ever estimated, including the sorts that were tried, produced nothing, and were quietly abandoned — a quantity nobody can observe.

The second line of attack is replication, and it produced the finding that matters most for this book. McLean and Pontiff (2016) took 97 characteristics whose predictive power had been documented in published papers, and re-computed each one's return in the period *after* the original sample ended. They found two distinct decays. Predictor returns were about **26 percent lower** in the post-sample period before publication — the signature of statistical overfitting, since the original researcher chose the sample and the specification that worked. And they were about **58 percent lower** after publication. The gap between the two figures, more than half the total decay, is not a statistical artifact. It is the effect of publication itself: after a paper appears, trading volume and short interest in the affected stocks rise, and the returns fall. Section 6.8 takes that seriously as the chapter's central holder result.

![Figure 6.3: Post-publication decay](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-be8abb32ffae4fd462a9f1dae388b3f225b5e0b6%2Ffig_06_03_post_publication_decay.png?alt=media)

**Figure 6.3: Post-publication decay.** McLean and Pontiff's two numbers, drawn against the in-sample estimate they are measured from. Ninety-seven published predictors, each recomputed in the period after its own original sample ended: 26 percent lower out of sample before publication, 58 percent lower after it. The arrows are the two decays, and the arithmetic of the second one is the chapter's argument. Of the 58 points of total decline, 26 are gone before anybody outside the author's office could have read the paper — that is overfitting, and it is a statement about statistics. The remaining 32 arrive after publication, which is more than half the total and cannot be a statistical artifact, because the statistics were already fixed when the paper went to press. Nothing about the firms changes on publication day. What changes is who holds them: volume rises, short interest rises, and the return falls. §6.8 reads that as a constraint's shadow price dying when the constraint is relaxed. *Source: McLean and Pontiff (2016), as reported in §6.4.*

Hou, Xue and Zhang (2020) applied a uniform methodology — value-weighted returns and NYSE breakpoints throughout, which removes the microcap tilt that inflates equal-weighted results — to 452 published anomalies, and found that roughly 65 percent failed to clear even a conventional $$|t| > 1.96$$. Chen and Zimmermann (2022) push back from an open-source replication of a comparable set, arguing that a large majority of published predictors do reproduce in-sample and that the decay observed afterwards is closer to what a rational-learning model would predict than to what pure data-mining would predict.

What survives, on a fair reading of the evidence together: the market factor; a value-like dimension and a profitability dimension, each with an economic story attached; investment; momentum, whose robustness across markets and asset classes is what saves it from the multiple-testing objection, since nobody chose forty countries and six asset classes as a specification search; and, in a diminished form, size. That is roughly the five-factor model plus momentum. Nearly everything else is either a repackaging of these or a candidate for the graveyard. One dimension is absent from that list rather than rejected by it: liquidity, both an asset's own trading costs and its loading on market-wide liquidity shocks, has as strong a claim to be priced as anything in the surviving set, and it is left out here only because its mechanism is a question about market structure rather than about a characteristic sort — **Chapter 11 §11.5** owns the priced-liquidity evidence and **Chapter 19** owns the funding side that makes the exposure systematic.

This is a live literature, and honest teaching requires marking the line, as Chapter 20 §20.5 does for the demand system.

**Table 6.3: The state of the evidence on the cross-section**

| Claim                                                                                     | Status                                                                                                                                                                    |
| ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Value, profitability, investment and momentum sort average returns                        | **Established.** Thirty years, several methods, and — for momentum — forty national markets and several asset classes                                                     |
| Momentum has no seat in an equilibrium model                                              | **Established** as a fact about the literature; the behavioral readings fit more naturally, and none of them is settled                                                   |
| Size is a standalone premium in US equities                                               | **Contested.** Large before 1981 and near zero since; its best modern defense is as a conditioning variable                                                               |
| Most published factors would fail a hurdle appropriate to the number of tests behind them | **Widely accepted.** Harvey, Liu and Zhu's working cutoff is $$\lvert t\rvert > 3.0$$; Hou, Xue and Zhang fail roughly 65 percent of 452 anomalies under a uniform method |
| Predictor returns decay out of sample, and decay further after publication                | **Established** as a magnitude; **contested** as to how much is overfitting and how much is arriving capital                                                              |
| HML is redundant once RMW and CMA are included                                            | **Established** in US data and no wider than that — a statement about one factor set, not about value                                                                     |

*Source: Author's assessment of the literature discussed in §§6.2-6.4.*

> **Box 6.1 — One Person's Portfolios**
>
> Table 6.2's arithmetic treats the 316 tests as 316 draws. They are not, and the reason is worth more attention than it usually gets. Nearly every number in §6.2, and a large share of the published cross-section since 1993, is computed from the same source: the portfolios and factor returns Kenneth French posts, free, at Dartmouth. The library is the single most-used piece of infrastructure in empirical asset pricing, and using it means inheriting a set of construction choices rather than making them. NYSE breakpoints rather than all-stock breakpoints, which keeps the microcap tail from setting the cut points. Annual rebalancing in June. Book equity from the fiscal year ending in the previous calendar year, market equity from that December for the ratio and from June for the weighting. Momentum measured from twelve months back to two months back, with the one-month gap that excludes short-run reversal. Value-weighted returns within each cell.
>
> Every one of those is defensible and none is forced. Appendix A §A.2 works through what each does to the resulting spread. The point here is that the choice among them was made once, by one researcher, in the early 1990s, and has since been inherited a few thousand times rather than re-litigated. That has three consequences for how the evidence should be read.
>
> The tests share more than a sample period. Two studies that use different characteristics but the same pre-built portfolios are running two tests against one construction, so the effective number of independent tests is smaller than the count, and the effective sample smaller than the number of papers implies. Bonferroni's cutoff is computed as though the opposite were true. This does not make the multiple-testing problem smaller; it makes it harder to size, and it is a second problem rather than a restatement of the first.
>
> Robustness checks that vary the specification but not the portfolios test less than they appear to. Which is exactly why Hou, Xue and Zhang's re-estimation of 452 anomalies under one uniform method carries the weight it does: they re-made the construction instead of re-using it, and roughly 65 percent of the anomalies did not survive the remaking. A finding that depends on the library's conventions rather than on the characteristic is a finding about the library.
>
> And the files move. They are rebuilt when CRSP and Compustat revise, and the revisions reach back decades, so two researchers running the same regression on the same nominal sample can disagree because they downloaded in different years. Appendix B §B.2 states the practical rule that follows: record the file date, because the vintage is part of the result.
>
> None of this is an argument against the library, which is a public good of the first order — it is why this literature is replicable at all, why Chapter 6's data exercise costs nothing to run, and why a masters student can reproduce Fama and French in an afternoon. It is an argument for saying which vintage you used, and for re-making the construction yourself when the construction is what the result turns on.

***

## 6.5 Machine Learning and the High-Dimensional Cross-Section

The zoo poses a problem that portfolio sorts cannot solve. If there are hundreds of candidate characteristics, and the relation between them and expected returns may be nonlinear and interactive, then the sort — which handles one variable well, two adequately, and three badly — is the wrong instrument. So is the linear cross-sectional regression, which with several hundred correlated regressors and a few hundred monthly observations will fit the sample beautifully and forecast nothing.

Gu, Kelly and Xiu (2020) ran the comparison properly. Taking roughly thirty thousand US stocks from 1957 to 2016, they assembled 94 firm characteristics, industry indicators, and eight macroeconomic series, interacted them, and forecast individual monthly returns with a ladder of methods: ordinary least squares, penalized linear models, dimension-reduction methods, random forests, gradient-boosted trees, and shallow neural networks. Everything was evaluated **out of sample**, on data the model never saw, with the sample split chronologically so that no future information leaks backwards.

Three results define the field.

First, **the ranking is decisive and it favors the flexible methods.** Unregularized OLS with the full predictor set produces a *negative* out-of-sample $$R^2$$ — worse than forecasting every stock's return with the historical mean. Penalized linear models repair the damage and get to roughly zero. Trees and neural networks reach an out-of-sample monthly $$R^2$$ of about 0.4 percent for individual stocks. That number sounds trivially small and is not: monthly stock returns are nearly unforecastable, and 0.4 percent of their variance, applied across thousands of names, translates into a value-weighted long-short decile portfolio with an annualized Sharpe ratio in the neighborhood of 1.3 — roughly double what the same exercise delivers using linear methods. Depth helps only up to a point: performance peaks at three to five hidden layers and then deteriorates, which is what one should expect from a signal-to-noise ratio this low.

Second, **the gains come from interactions, not from new information.** The variables the machines lean on are not exotic. Price trends — momentum at various horizons, short-term reversal — come first, followed by liquidity measures such as bid-ask spread, dollar volume and market capitalization, followed by volatility. These are the characteristics the sorting literature already knew about. What the flexible methods add is the ability to say that momentum matters *differently* in illiquid stocks than in liquid ones, and that the effect of one characteristic bends at particular values of another. The zoo, on this reading, is not four hundred independent phenomena; it is a modest number of phenomena observed through four hundred correlated, nonlinearly related proxies.

Third, **the out-of-sample discipline is the substance, not a formality.** The entire contribution of this literature is a protocol: split the sample chronologically, tune every hyperparameter on a validation window that precedes the test window, and report performance only on data the model has never touched. That protocol is a direct institutional response to §6.4's problem. It does not eliminate the multiple-testing concern — the *researcher* has still seen the out-of-sample period in every previous paper — but it converts an unfalsifiable claim into a measurable one.

### ★ IPCA and shrinkage

*Starred. Two responses to high dimension that keep the pricing framework intact.*

**Instrumented principal components analysis** (Kelly, Pruitt and Su, 2019) attacks the identification problem rather than the prediction problem. Its question is the old one: are characteristics proxies for *covariances*, as the APT requires, or do they predict returns directly, in which case something other than risk is at work? IPCA estimates latent factors whose loadings are functions of observable characteristics, $$\beta\_{i,t} = c\_{i,t}'\Gamma$$, so a characteristic earns its place by mapping into a loading. The test is then sharp: allow the characteristics to enter the loadings, and ask whether they still need to enter expected returns separately. Kelly, Pruitt and Su find that around five latent factors, with characteristic-driven loadings, absorb most of the cross-section, and that the residual characteristic effects — the "alphas" — are largely insignificant. Characteristics are covariances, on this evidence, though the covariances are with factors nobody has named.

**Shrinkage** is the zoo's other remedy, and Kozak, Nagel and Santosh (2020) draw the sharpest lesson from it. They estimate an SDF directly from a large set of characteristic-based portfolios, penalizing the SDF coefficients rather than selecting among the characteristics. Heavy shrinkage toward the leading principal components works well out of sample; *sparse* models — the handful of factors a discipline that prizes parsimony would prefer — do badly. There is no near-sparse SDF hiding in the data. The information in the cross-section is genuinely spread across many characteristics, which is uncomfortable for anyone who wanted the zoo to collapse to five animals, and is the strongest available argument for treating factor models as compression devices rather than as economic theories.

***

## 6.6 The Cross-Section Beyond Equities

Nothing in §6.1 was about stocks. The APT's argument runs on any set of assets with a factor structure and diversifiable residuals, and the sorting machinery of §6.3 runs on any characteristic that can be measured for a cross-section of instruments. Applying both outside equities does two things: it supplies an out-of-sample test that no amount of equity data mining can manufacture, and it converts a set of phenomena usually taught as market-specific curiosities into a single pricing question.

**Currencies.** The relevant characteristic is the interest rate. Sort the currencies of the developed and larger emerging markets each month by their short-term interest rate differential against the dollar, group them into portfolios, and hold each portfolio's currencies against the dollar. Lustig and Verdelhan (2007) and Lustig, Roussanov and Verdelhan (2011) showed that the resulting portfolios have monotonically increasing average excess returns, and — the important part — that two factors price them. The first is the **dollar factor**: the average excess return of all foreign currencies against the dollar, which carries no cross-sectional spread but captures the common movement. The second is **carry**, the high-interest-rate portfolio minus the low-interest-rate portfolio, which is exactly an HML-style long-short spread built on interest rates rather than book-to-market. Currency **momentum** works as well, on the same twelve-month formation the equity literature uses, and is largely uncorrelated with carry (Menkhoff, Sarno, Schmeling and Schrimpf, 2012).

The SDF framing makes the object clear. In a complete-markets two-country setting, the log excess return on holding foreign currency equals the difference between the domestic and foreign log stochastic discount factors. A currency earns a positive expected excess return precisely when its country's SDF is *less* volatile than the domestic one — when investing there provides less insurance. Carry portfolios are then a way of sorting countries by their exposure to a global source of risk, and the empirical counterpart is that the carry factor loads heavily on global FX volatility and on measures of global funding conditions. That is a pricing statement about a cross-section, and it is this chapter's.

It is not a treatment of the carry trade. **The mechanics of the trade, the evidence on uncovered interest parity and the forward premium puzzle, and the crash-risk and skewness story that goes with them are developed in&#x20;*****International Finance*****, Chapter 11 §11.5**, with the return and Sharpe-ratio tables. What is added here is the factor-portfolio construction and the pricing perspective: that currency returns are organized by a small number of factors extracted from cross-sectional sorts, in the same way and using the same machinery as equities.

**Commodities.** Two characteristics organize the commodity cross-section. The first is the shape of the futures curve — the **basis**, or equivalently carry: sort commodities by the difference between the nearby and deferred futures prices, and the backwardated ones (nearby above deferred) outperform the contangoed ones. The second is momentum, again on past returns of roughly a year. Gorton and Rouwenhorst (2006) established that a diversified commodity futures portfolio earns an equity-like premium with low correlation to equities and bonds; the subsequent literature showed that essentially all of the cross-sectional variation is captured by basis and momentum sorts, and that the basis signal is a noisy read on inventory — low inventories produce backwardation *and* high expected returns, because scarce inventory means the physical holder's convenience yield is high and the futures buyer is being paid to supply storage capacity the market does not have.

**Why commodity carry works is a question about storage, and it belongs to Chapter 8 §8.2**, where the cost-of-carry relation, the theory of storage, convenience yield, and backwardation and contango are developed as pricing theory. This section takes the futures curve as given and treats its slope as a sortable characteristic.

Koijen, Moskowitz, Pedersen and Vrugt (2018) close the circle by defining carry uniformly across asset classes — the return an asset earns if its price does not change — and showing that carry-sorted portfolios earn positive average returns in global equities, bonds, currencies, commodities, credit, and index options. Asness, Moskowitz and Pedersen (2013) do the same for the other two workhorses, documenting value and momentum in eight markets and asset classes, and reporting the structure that makes the result hard to dismiss as coincidence: value and momentum are *negatively* correlated with each other within every market, and each is *positively* correlated with itself across markets. A combined value-and-momentum portfolio, diversified across asset classes, therefore has a far higher Sharpe ratio than any of its parts.

**Table 6.4: The same sorts, different assets**

| Asset class      | Carry-type sort            | Momentum sort          | Value-type sort         |
| ---------------- | -------------------------- | ---------------------- | ----------------------- |
| Equities         | Dividend yield             | Past 12-2 month return | Book / market           |
| Government bonds | Term-structure slope       | Past 12-2 month return | Real yield level        |
| Currencies       | Interest rate differential | Past 12-2 month return | Deviation from PPP      |
| Commodities      | Futures basis              | Past 12-2 month return | Reversal over \~5 years |

*Source: Constructed from the sorting definitions in Asness, Moskowitz and Pedersen (2013) and Koijen, Moskowitz, Pedersen and Vrugt (2018). The table states the signals, not their realized premia.*

Read the table for what it says about §6.4's problem. Value and momentum in currencies and commodities were not found by searching a database of firm characteristics; the same two signals, defined analogously, work on assets that share no accounting standards, no investor base, and no regulator. That is the strongest available argument that something real is being measured. It is also, read the other way, an argument that whatever is being measured is not firm-specific risk — because there are no firms.

***

## 6.7 Climate Risk Premia and the Greenium

The newest sort is on carbon. If investors demand compensation for holding assets exposed to climate transition risk, brown firms should have higher expected returns than green ones; if instead a growing share of capital holds green assets for nonpecuniary reasons, the same conclusion follows from the demand side, and Chapter 16 §16.6 derives it there. Either way the prediction is a **greenium**: green assets priced high, expected returns low.

The evidence is genuinely mixed and the reason is instructive. Bolton and Kacperczyk (2021) report a carbon premium — firms with higher total emissions and faster emissions growth earn higher subsequent returns, after controlling for the standard factors. Subsequent work has contested it hard, showing that the result depends on using *total* rather than intensity-scaled emissions (which reintroduces a size effect), on vendor-estimated rather than firm-disclosed emissions data, and on the treatment of the look-ahead bias created by emissions disclosures that arrive with a lag. In bond markets the estimate is smaller and more stable: matched-pair studies of green bonds against otherwise identical conventional bonds of the same issuer find yield differentials of at most a few basis points, in the predicted direction.

The subtlety that organizes all of it is the distinction between **expected and realized** returns, and it is worth stating carefully because it is where most of the popular discussion goes wrong. Pástor, Stambaugh and Taylor's (2021) equilibrium with nonpecuniary preferences implies that green assets have *lower* expected returns — that is what a greenium means. But green assets can and did deliver *higher* realized returns over the 2010s, precisely because climate concern strengthened unexpectedly and drove green prices up. A price increase caused by a fall in the required return is not evidence of a high required return; it is evidence of the opposite. Their follow-up work decomposes green stocks' strong realized performance and attributes much of it to exactly this channel. Anyone selling green investing as an outperformance strategy is quoting the transition path as though it were the destination.

Two implications for the rest of the chapter. First, this is the cleanest live case of the risk-versus-demand ambiguity, because both readings — a transition risk premium, and a nonpecuniary preference — predict the same sign in expected returns and are distinguished only by mechanism. Second, the identification problem is the one §6.4 warned about in an acute form: the sample is short, the emissions data are recent and revised, and the number of plausible specifications is large. The mechanism is not in serious doubt; the magnitude is. **Mandates, divestment arithmetic, and the equilibrium with nonpecuniary preferences are developed in Chapter 16 §16.6** and are not restated here.

***

## 6.8 Who Holds the Factors, and What Their Constraints Do to Their Prices

Every result in this chapter is a statement about an average return. This section asks who was earning it, what happened to the return when more of them arrived, and what the capacity of the trade has to do with the answer.

Return to §6.4's central number. McLean and Pontiff found that predictor returns fall about 26 percent out of sample before publication and about 58 percent after it. The first decay is a statement about statistics. The second is a statement about *capital*: more than half the total decline occurs after a paper appears, accompanied by rising volume, rising short interest, and rising correlation among the affected stocks. Nothing about the firms changes on publication day. What changes is who holds them.

Read that alongside Chapter 4 §4.9. There, a constraint on holders *created* a premium: mutual funds that cannot borrow bid up high-beta stocks, and someone had to be paid to take the other side. Here, the arrival of holders *destroys* one. Both say that the constraints of the marginal investor set the price, and that a factor premium is not a property of a claim but a property of an equilibrium between a claim and the people who can hold it.

**The capacity logic is Berk and Green's, applied one level up.** Chapter 17 §17.3 develops the mechanism for a single fund: gross alpha declines in assets under management, $$\alpha^{\text{gross}}(A) = a - bA$$, competitive capital drives net alpha to zero, and the fund settles at $$A^{\ast} = (a-f)/b$$, with the manager collecting the rent and the investor earning the benchmark. The same algebra applies to a *strategy* rather than a fund — the case Chapter 17 flagged as the one that matters. If decreasing returns bind at the level of the trade, because price impact depends on how much total capital is pushing the same stocks the same way rather than on how much any one manager is pushing, then $$A$$ is industry-wide capital in the factor and $$A^{\ast}$$ is the factor's capacity. Suppose a strategy earns 5 percent gross alpha at small scale, each additional billion dollars deployed erodes it by 5 basis points, and managers charge 1 percent. Then $$A^{\ast} = (5 - 1)/0.05 = 80$$ billion dollars, at which point the trade still earns 1 percent gross and exactly nothing net. Publication changes neither $$a$$ nor $$b$$. It changes how many managers know the value of $$a$$, and capital arrives until the identity binds.

Two consequences follow that a fund-level model does not deliver. Capacity is a *shared* resource, so every entrant imposes a negative externality on incumbents that none of them prices. And because the constraint binds on the aggregate, no individual manager's exit restores the premium — which is why factor returns decay and stay decayed rather than oscillating around a long-run mean.

**The externality has a name and a date.** In the second week of August 2007, quantitative equity market-neutral funds — long-short portfolios built on essentially the signals of §6.2, hedged to zero market exposure — suffered losses their risk models had assigned negligible probability. Nothing had happened to the signals. What appears to have happened, in Khandani and Lo's reconstruction, is that a large multi-strategy fund facing losses elsewhere in its book liquidated its liquid market-neutral positions to raise cash. Those were the same positions everyone else held, because everyone had sorted on the same characteristics using the same data. The liquidation pushed the crowded names against every fund running a similar book; those funds hit risk limits; some deleveraged; the deleveraging moved prices further. Losses accumulated over 7 to 9 August and a large fraction reversed on 10 August once the forced selling stopped — the reversal being the proof that the move was liquidity, not information.

No factor model can see this coming, and the reason is structural. Section 6.1's arbitrageur is a price-taker with infinite patience whose only interaction with other arbitrageurs is that they collectively enforce the pricing relation. The August 2007 arbitrageur is one of a few dozen holders of a shared position, financed with borrowed money, facing a risk limit set by someone else. Crowding is an exposure that appears in no covariance matrix estimated from returns, because it is a fact about *holdings* rather than prices — until the moment it becomes a fact about prices. Chapter 16 §16.5 states the mechanism generally; Chapter 19 develops the balance sheets that transmit it.

**So who are these holders?** The composition has changed twice. Through the 1990s, factor exposure was the private property of a small quantitative hedge-fund industry, financed with prime-broker leverage and charging performance fees. From the mid-2000s the same exposures were repackaged as **smart beta**: index-tracking funds and ETFs delivering value, momentum, quality and low-volatility tilts for a few basis points, unlevered and daily-liquid. That reorganization, described institutionally in Chapters 16 and 17, moved a great deal of capital into the factors and changed the character of the holders. A hedge fund can wait out a drawdown for as long as its lockups permit; a factor ETF's holders redeem after two bad years, and its manager has no discretion to stop tracking. The strategy's patience is now the patience of its least patient holder. The third layer is delegation: pensions and insurers (Chapter 16 §16.2) buy factor exposure through mandates specified as a benchmark plus a tracking-error budget, and Chapter 16 §16.3's agency problem then applies in full. A manager judged quarterly against a factor benchmark cannot take the other side of a crowded trade even when the expected return says she should, because the tracking error is what she is punished for. The holders best placed to relieve the crowding are contractually prohibited from relieving it.

Which returns the chapter to the question it opened with. If a factor premium is compensation for risk, it should survive the arrival of new capital, because the risk survives. If it is mispricing, it should not, and post-publication decay is what its disappearance looks like. If it is data-mining, there was nothing there to arrive for. The evidence will not sort every factor into the right bin and may never do so. But the diagnostic that discriminates best is not a statistical one: watch the holders. A risk premium is paid to whoever bears the risk, and that person's identity can change without the premium changing. A constraint's shadow price dies when the constraint is relaxed — and the constraint, in every case in this chapter, is a fact about somebody's balance sheet.

In Chapter 1 §1.2's terms, a factor premium that decays after publication is a flow and not news — capital arriving at a characteristic nobody re-valued — while a factor that loses money over three days on signals that did not change, as the quantitative equity book did in August 2007, is a constraint binding on the holders who had crowded into it.

***

## Elsewhere in the Series

* **The carry trade itself** — its mechanics, the evidence on uncovered interest parity and the forward premium puzzle, and the crash-risk and skewness interpretation — is *International Finance*, **Chapter 11 §11.5**, with return and Sharpe-ratio tables. Section 6.6 builds the currency factor portfolios and asks the pricing question; it does not re-derive the parity conditions.
* Within this book: **Chapter 8 §8.2** owns futures pricing — cost of carry, the theory of storage, convenience yield, backwardation and contango — which is what makes the commodity basis a meaningful sort variable. **Chapter 22 §22.4** owns the q-theoretic and investment-CAPM rationalization of the profitability and investment factors. **Chapter 16 §16.6** owns mandates, divestment arithmetic, and nonpecuniary demand. **Appendix A** owns the full mechanics of portfolio sorts, Fama-MacBeth regressions with the Shanken correction, and the machine-learning protocols of §6.5.

***

## Summary

1. **The 1992 episode was a rejection of a model, not of a research program.** Fama and French showed that beta has no reliable relation to average returns once size and book-to-market are controlled. Fama's own reading — that the Sharpe-Lintner CAPM died and market efficiency did not — follows from the joint-hypothesis problem (Chapter 7), and the 1993 three-factor model was the replacement.
2. **The APT derives a multi-beta pricing relation from near-arbitrage.** Given a factor structure $$r\_i = E\[r\_i] + \sum\_k \beta\_{i,k}(f\_k - E\[f\_k]) + \varepsilon\_i$$, residual risk vanishes in well-diversified portfolios, so any deviation from $$E\[r\_i] - r\_f = \sum\_k \beta\_{i,k}\lambda\_k$$ is a riskless money machine. No preferences, no market clearing, no observable market portfolio — and, decisively, no statement about which factors or how many.
3. **The contrast with the CAPM is a trade of content for robustness.** The CAPM names its factor and pins $$\lambda\_M = \mu\_M - r\_f$$, at the cost of assumptions that fail. The APT survives almost any assumption and pins nothing. The factor zoo is the price of that license.
4. **Five canonical equity factors.** Size (\~2-3% per year historically, near zero since 1980), value (\~4%, absent 2007-2020), momentum (\~8%, with crashes), profitability (\~3%), investment (\~3%) — all approximate, all sample-dependent. Fama and French (1993) built SMB and HML from a 2×3 double sort; the 2015 five-factor model adds RMW and CMA and makes HML largely redundant in US data.
5. **Momentum has no comfortable seat.** It is the most robust pattern in the data and the least explicable: a property of the price path rather than of the firm, excluded from both Fama-French models, added by Carhart (1997), and prone to crashes — roughly 90 percent lost in two months of 1932 and on the order of 70 percent between March and May 2009, when the loser leg's high beta met a market rebound.
6. **Profitability and investment come with a theory, and it is a corporate one.** Holding value fixed, higher expected profitability implies a higher discount rate; holding profitability fixed, higher investment implies a lower one. That is RMW and CMA read off a valuation identity plus optimal investment — the investment CAPM, developed in **Chapter 22 §22.4**.
7. **Two procedures make the evidence.** Portfolio sorts impose no functional form, are robust to outliers, and yield a tradable portfolio; double sorts control a correlated characteristic. Fama-MacBeth runs time-series regressions for $$\hat\beta\_{i,k}$$, then a cross-sectional regression each period for $$\hat\alpha\_t$$ and $$\hat\lambda\_{k,t}$$, and takes time-series means and standard errors — which handles cross-sectional correlation but not errors-in-variables in $$\hat\beta$$. Full mechanics: **Appendix A**.
8. **The zoo is a multiple-testing problem.** 316 published factors through 2012 imply 15.8 expected false discoveries at a 5 percent threshold and a Bonferroni cutoff of $$|t| > 3.78$$; Harvey, Liu and Zhu recommend about 3.0. McLean and Pontiff find predictor returns 26 percent lower out of sample and 58 percent lower after publication; Hou, Xue and Zhang find roughly 65 percent of 452 anomalies fail to replicate under uniform value-weighted methodology.
9. **Machine learning changes the method, not the phenomena.** Gu, Kelly and Xiu find trees and shallow neural networks reach an out-of-sample monthly $$R^2$$ near 0.4 percent where unpenalized OLS goes negative, translating into a long-short Sharpe ratio near 1.3 — with the gains coming from interactions among predictors the sorting literature already knew (price trends, liquidity, volatility). ★ IPCA maps characteristics into loadings and finds about five latent factors absorb the cross-section; ★ shrinkage works and sparsity does not, so there is no near-sparse SDF.
10. **The same sorts price currencies and commodities.** Interest-rate sorts of currencies produce dollar and carry factors; futures-basis and momentum sorts organize commodities; carry, value and momentum appear across eight markets and asset classes with a correlation structure — negative within markets, positive across them — that is hard to attribute to data-mining. Boundaries: the carry trade and UIP are *International Finance* Ch 11 §11.5; storage and cost of carry are Chapter 8.
11. **The greenium turns on expected versus realized returns.** Equilibrium with nonpecuniary preferences implies green assets have *lower* expected returns; strengthening climate concern can nonetheless deliver *higher* realized returns while the repricing happens. The carbon-premium evidence is contested on emissions-data construction; the green-bond differential is small, stable, and correctly signed.
12. **Factor returns are a property of an equilibrium between a claim and its holders.** Post-publication decay is capital arriving. Berk-Green capacity applied at the strategy level gives $$A^{\ast} = (a-f)/b$$ for the whole trade, with an externality no entrant prices. August 2007 is what the externality looks like when it binds. And the holder base has shifted from levered hedge funds to daily-liquidity smart-beta vehicles and tracking-error-constrained mandates — which means the strategies' patience is now the patience of their least patient holders.

***

## Key Terms

* **Arbitrage pricing theory (APT)**: Ross's (1976) result that a factor structure plus the absence of near-arbitrage forces expected returns to be approximately linear in factor loadings
* **Factor structure**: The assumption that common variation in a large cross-section of returns is spanned by a small number of factors, leaving idiosyncratic residuals
* **Well-diversified portfolio**: A portfolio in which every weight is small, so residual variance vanishes as the number of holdings grows
* **Factor-mimicking portfolio**: A portfolio with unit loading on one factor and zero on the rest; its expected excess return is that factor's price of risk $$\lambda\_k$$
* **SMB, HML, UMD, RMW, CMA**: The traded factor returns for size, value, momentum, profitability and investment — small minus big, high minus low book-to-market, up minus down, robust minus weak, conservative minus aggressive
* **Book-to-market ratio**: Book equity divided by market equity; the value characteristic, and an increasingly imperfect one as intangible capital grows
* **Portfolio sort**: Ranking assets on a characteristic, forming portfolios from the ranking, and reporting the top-minus-bottom average return
* **Double sort**: A two-dimensional sort, independent or conditional, used to measure one characteristic's effect holding another fixed
* **Fama-MacBeth regression**: The two-pass estimator — time-series regressions for loadings, then period-by-period cross-sectional regressions for $$\hat\alpha\_t$$ and $$\hat\lambda\_{k,t}$$, with inference from the time series of the estimates
* **Errors-in-variables problem**: The attenuation of second-pass coefficients caused by using estimated betas as regressors; corrected by Shanken (1992)
* **Factor zoo**: The several hundred published characteristics claimed to predict the cross-section, most of which do not survive a multiple-testing-adjusted hurdle
* **Multiple-testing hurdle**: The raised $$t$$-statistic threshold appropriate when many hypotheses are tested against one dataset; about 3.0 on Harvey, Liu and Zhu's recommendation
* **Post-publication decay**: The fall in a predictor's return after its publication, attributable to arbitrage capital rather than to statistics
* **Out-of-sample** $$R^2$$: Forecast accuracy on data never used in estimation or tuning; the discipline that distinguishes the machine-learning literature from what preceded it
* **IPCA**: Instrumented principal components; latent factors whose loadings are functions of observable characteristics, testing whether characteristics are covariances
* **Carry**: The return an asset earns if its price does not change — the interest differential in currencies, the futures basis in commodities, the yield in bonds
* **Dollar factor**: The average excess return on foreign currencies against the dollar; the level factor in the currency cross-section
* **Greenium**: The lower expected return on green assets implied by nonpecuniary demand or transition-risk pricing; distinct from their realized return during a repricing
* **Crowding**: Shared exposure to the same positions across levered holders; invisible in a return covariance matrix because it is a fact about holdings
* **Factor capacity**: The industry-wide capital at which a strategy's net alpha reaches zero; Berk-Green's $$A^{\ast} = (a-f)/b$$ applied to a trade rather than a fund

***

## Readings

### Required

* Fama, E. and K. French (1992). "The Cross-Section of Expected Stock Returns." *Journal of Finance* 47(2): 427-465. *The paper of the opening episode: read §II for the bivariate sorts that flatten beta, and the conclusion for the authors' own insistence that this is a finding about a model of expected returns.*
* Fama, E. and K. French (1993). "Common Risk Factors in the Returns on Stocks and Bonds." *Journal of Financial Economics* 33(1): 3-56. *The replacement model and, more usefully for a student, the definitive statement of how SMB and HML are actually built — the 2×3 sort of §6.2 in the authors' own words.*

### Recommended

* Ross, S. (1976). "The Arbitrage Theory of Capital Asset Pricing." *Journal of Economic Theory* 13(3): 341-360. *The original near-arbitrage argument; read it for how little is assumed and how much less is delivered than a CAPM-trained reader expects.*
* Chen, N.-F., R. Roll and S. Ross (1986). "Economic Forces and the Stock Market." *Journal of Business* 59(3): 383-403. *The macro-factor tradition behind §6.1: factors named in advance from the macroeconomy rather than extracted from the returns they price. Read it for what an APT test looks like when the antecedent is filled in honestly.*
* Fama, E. and K. French (1996). "Multifactor Explanations of Asset Pricing Anomalies." *Journal of Finance* 51(1): 55-84. *The follow-up to the two Required papers, and the one that answers §6.2's "risk or mispricing" objection directly by testing the three-factor model against the anomaly set instead of merely constructing it.*
* Jegadeesh, N. and S. Titman (1993). "Returns to Buying Winners and Selling Losers." *Journal of Finance* 48(1): 65-91. *The momentum paper; note the care taken over formation and holding periods, which is what made the result survive.*
* Hansen, L. P. and S. Richard (1987). "The Role of Conditioning Information in Deducing Testable Restrictions Implied by Dynamic Asset Pricing Models." *Econometrica* 55(3): 587-613. *Sits underneath §6.3's Fama-MacBeth machinery: a conditional model implies restrictions that unconditional tests do not recover, which is why the standard cross-sectional regression is weaker evidence than it appears.*
* Harvey, C., Y. Liu and H. Zhu (2016). "…and the Cross-Section of Expected Returns." *Review of Financial Studies* 29(1): 5-68. *Counts the factors, applies multiple-testing corrections, and proposes the raised hurdle; the framing of the whole replication debate.*
* McLean, R. D. and J. Pontiff (2016). "Does Academic Research Destroy Stock Return Predictability?" *Journal of Finance* 71(1): 5-32. *The 26-versus-58 percent decomposition; the single most important empirical result in this chapter for the book's thesis, because the second number is capital and not statistics.*
* Shin, H. S. (2003). "Disclosures and Asset Returns." *Econometrica* 71(1): 105-133. *A third explanation for why characteristics predict returns, alongside risk and data mining: what firms disclose, and what they are permitted to withhold, generates cross-sectional variation in expected returns without either.*
* Amihud, Y., H. Mendelson and L. H. Pedersen (2005). "Liquidity and Asset Prices." *Foundations and Trends in Finance* 1(4): 269-364. *The survey behind §6.4's one line on liquidity, covering both the trading-cost level and the systematic-liquidity loading; the natural bridge to Chapter 11 §11.5.*
* Gu, S., B. Kelly and D. Xiu (2020). "Empirical Asset Pricing via Machine Learning." *Review of Financial Studies* 33(5): 2223-2273. *The methods comparison of §6.5; read the variable-importance results, which say the gains come from interactions among familiar predictors.*
* Asness, C., T. Moskowitz and L. Pedersen (2013). "Value and Momentum Everywhere." *Journal of Finance* 68(3): 929-985. *Value and momentum across eight markets and asset classes, with the correlation structure — negative within, positive across — that is the strongest single argument against the data-mining reading.*
* Daniel, K. and T. Moskowitz (2016). "Momentum Crashes." *Journal of Financial Economics* 122(2): 221-247. *Why momentum's payoff resembles a written option on the market, and what happened in 1932 and 2009.*
* Kozak, S., S. Nagel and S. Santosh (2020). "Shrinking the Cross-Section." *Journal of Financial Economics* 135(2): 271-292. *Shrinkage beats selection, and there is no near-sparse SDF — the most consequential negative result in the zoo literature.*
* Asquith, P., M. Mikhail and A. Au (2005). "Information Content of Equity Analyst Reports." *Journal of Financial Economics* 75(2): 245-282. *Who actually produces the signals a characteristic sort is built from, and what the market does with them; the supply-side companion to §6.8's crowding-and-capacity argument.*

***

## Discussion Questions

1. **Risk, mispricing, or the search itself.** Take the value premium. State the sharpest version of each of the three readings, and for each one name an observable that the other two do not predict. Then explain why the three are hard to separate even in principle, and say whether the fact that value also appears in currencies and commodities (§6.6) helps with any of the three, all of them, or only one.
2. **Should publication kill a factor?** McLean and Pontiff find that predictor returns fall by more than half after publication. One reading is that academic research improves market efficiency and should be celebrated. Another is that authors are unpaid research staff for the funds that trade their results. A third is that the decay is evidence the finding was never a risk premium in the first place. Which reading do you take, and what would you have to observe to change your mind? Does your answer imply anything about whether journals should require pre-registration of the specification search?
3. **What the APT cannot be blamed for.** A researcher publishes a factor built on the ratio of a firm's cash holdings to its advertising expenditure, with a $$t$$-statistic of 2.6, and cites the APT as its theoretical justification. Is the citation legitimate? Construct the strongest defense of the paper and the strongest objection, and say where in §6.1's argument the disagreement actually lies.
4. **Momentum's seat at the table.** Fama and French exclude momentum from both the three- and five-factor models while conceding it is the largest failure of each. Give the best case for that exclusion and the best case against it. If you were designing the benchmark against which a delegated equity manager's performance is measured (Chapter 17 §17.3), would you include momentum? Note that your answer determines whether a manager who runs a momentum strategy is judged to have skill.
5. **Size and liquidity outside the developed markets.** Section 6.6 extends the cross-section to bonds, currencies and commodities but stays inside developed markets. In an emerging equity market, small firms are also the least liquid firms, and the two characteristics are close to collinear. Say what that does to the interpretation of a measured size premium there, and design a sort that would separate the two. Then ask the harder question: if the premium is compensation for illiquidity, who is the marginal holder collecting it, and what happens to the estimate when foreign investors — who face a different redemption horizon than domestic ones — enter or leave? (Chapter 11 §11.5 supplies the liquidity machinery.)
6. **Whose patience?** Section 6.8 argues that a factor strategy's ability to earn its premium depends on the redemption behavior of its holders. Compare a levered hedge fund with three-year lockups, a daily-liquidity factor ETF, and a pension mandate with a two percent tracking-error budget. For each, say what happens after a three-year drawdown in the factor, and what that implies about who *should* hold factor exposure. Is there a holder for whom the crowding externality of §6.8 is not a problem?

***

## Problems

**Problem 1 — Constructing the APT arbitrage.** In a single-factor economy, three well-diversified portfolios have the following loadings and expected returns: P has $$\beta\_P = 0.4$$ and $$E\[r\_P] = 7$$ percent; Q has $$\beta\_Q = 1.2$$ and $$E\[r\_Q] = 15$$ percent; S has $$\beta\_S = 0.8$$ and $$E\[r\_S] = 10$$ percent.

(a) Using P and Q, find the implied price of factor risk $$\lambda$$ and the implied zero-beta intercept. (b) What expected return does the resulting pricing line assign to a portfolio with loading 0.8? Is S cheap or dear? (c) Construct an explicit zero-investment, zero-factor-exposure position using all three portfolios. State the weights and verify that both the net investment and the net loading are zero. (d) Compute the annual profit on a $250 million gross position on each side. Why is this profit riskless here, and what exactly in the setup would have to fail for it not to be? (e) The arbitrage is executed. Which portfolio's price moves, in which direction, and by how much must its expected return change for the opportunity to close?

**Problem 2 — A Fama-MacBeth premium from a small panel.** Four test portfolios have full-sample loadings on a single factor of $$\hat\beta = (0.5, 1.0, 1.5, 2.0)$$. Their monthly excess returns, in percent, over five months are:

**Table 6.5: A five-month panel of excess returns (percent per month)**

| Month | Portfolio 1 ($$\hat\beta = 0.5$$) | Portfolio 2 ($$\hat\beta = 1.0$$) | Portfolio 3 ($$\hat\beta = 1.5$$) | Portfolio 4 ($$\hat\beta = 2.0$$) |
| ----- | --------------------------------- | --------------------------------- | --------------------------------- | --------------------------------- |
| 1     | 1.2                               | 0.9                               | 1.4                               | 2.7                               |
| 2     | 0.0                               | −1.1                              | −0.2                              | −1.3                              |
| 3     | 1.1                               | 1.5                               | 2.5                               | 4.1                               |
| 4     | 0.4                               | −0.1                              | 0.4                               | −0.1                              |
| 5     | 1.1                               | 0.7                               | 1.3                               | 2.9                               |

*Source: Author's construction.*

(a) Run the second-pass cross-sectional regression $$r\_{i,t} = \hat\alpha\_t + \hat\lambda\_t\hat\beta\_i + u\_{i,t}$$ separately for each of the five months, and report $$\hat\alpha\_t$$ and $$\hat\lambda\_t$$. (With four assets and a single regressor the OLS slope is $$\sum\_i (\hat\beta\_i - \bar\beta)(r\_{i,t} - \bar r\_t)/\sum\_i(\hat\beta\_i - \bar\beta)^2$$, and $$\sum\_i(\hat\beta\_i-\bar\beta)^2 = 1.25$$ throughout.) (b) Compute the Fama-MacBeth estimates $$\hat\lambda$$ and $$\hat\alpha$$ as time-series means, their standard errors, and their $$t$$-statistics. (c) Annualize $$\hat\lambda$$. Would you report this as evidence that the factor is priced? Answer using the $$t$$-statistic and using §6.4's hurdle, and say which of the two objections is the more serious here. (d) Now compute each portfolio's *average* excess return over the five months and run a single cross-sectional regression of those four averages on the four betas. Compare the coefficients with your answers in (b). Explain why the two agree exactly, and state the condition on the panel under which they would not.

**Problem 3 — The multiple-testing hurdle.** A journal receives submissions proposing new return predictors. Assume every test statistic is standard normal under the null of no predictability, and that tests are independent.

(a) If 316 predictors are tested and every null is true, how many will clear $$|t| > 1.96$$ in expectation? What is the probability that at least one does? (b) What per-test significance level controls the family-wise error rate at 5 percent under Bonferroni, and what two-sided $$t$$-cutoff does it imply? (c) Repeat (b) for 100 and for 400 tests. Comment on how slowly the cutoff rises in the number of tests, and why that is a mathematically comforting and practically useless fact. (d) A referee objects that the relevant $$J$$ is not the number of *published* factors but the number of specifications ever estimated. Explain why this makes the true hurdle unknowable, and propose one institutional arrangement that would make it knowable. (e) A factor has a true annual premium of 3 percent with annual volatility of 12 percent. How many years of data are needed for its expected $$t$$-statistic to reach 1.96? To reach 3.0? Comment on what your answer implies about the discovery of genuinely small but real premia.

**Problem 4 — Factor capacity and post-publication decay.** A published factor strategy generates gross alpha $$\alpha^{\text{gross}}(A) = a - bA$$, where $$A$$ is the total capital in the trade across all managers, measured in billions of dollars. Take $$a = 6$$ percent and $$b = 0.04$$ percentage points per billion. Managers charge a fee $$f$$ and capital enters until net alpha is zero.

(a) With $$f = 1$$ percent, find the equilibrium capital $$A^{\ast}$$, the gross alpha at that point, and the total fees collected per year. (b) Fees fall to 0.3 percent as the strategy is repackaged as a smart-beta ETF. Recompute $$A^{\ast}$$ and the fees collected. Who gained and who lost? (c) McLean and Pontiff report post-publication returns 58 percent below in-sample returns. If the in-sample estimate was the pre-entry gross alpha $$a$$, what level of $$A$$ does a 58 percent decline correspond to under this model? Is that a plausible amount of capital? (d) Explain why an incumbent manager cannot restore her alpha by reducing her own position, and identify the externality. Which of the holder types in §6.8 is most likely to be the marginal entrant, and why does that matter for how fast $$A$$ reaches $$A^{\ast}$$?

**Problem 5 ★ — Redundancy and the rotation problem.** Consider a four-factor model $$\lbrace f\_1, f\_2, f\_3, f\_4\rbrace$$ in which all four factors are traded excess returns, and suppose the true SDF is $$m = a - \sum\_k b\_k f\_k$$.

(a) Show that if $$f\_4$$ can be written as a linear combination of $$f\_1$$, $$f\_2$$ and $$f\_3$$ plus a residual that is uncorrelated with all returns, then $$f\_4$$ contributes nothing to pricing. Relate this to Fama and French's (2015) finding that HML becomes redundant once RMW and CMA are included. (b) Let $$G$$ be any invertible $$4\times 4$$ matrix and define rotated factors $$g = Gf$$. Show that the rotated model prices exactly the same set of assets. What does this imply for the question "which four factors are the right ones?" (c) In light of (b), explain what content remains in the claim that a particular factor is "priced". What additional structure — of the kind Chapters 5, 20 or 22 supply — would be needed to make the claim more than a statement about spanning? (d) Kozak, Nagel and Santosh find that shrinking toward the leading principal components of a large set of characteristic portfolios beats selecting a sparse subset. Interpret that result in light of (b). Is it an argument that the factor zoo is real, or an argument that the question of which factors are real is malformed?

**Problem 6 — Momentum's conditional beta.** A winner-minus-loser portfolio is constructed monthly. In normal states the winner and loser legs both have market beta 1.0, so the spread has beta 0. Following a market decline of 30 percent or more, the loser leg is populated by distressed, highly levered firms and its beta rises to 1.8, while the winner leg's beta stays at 1.0.

(a) What is the momentum portfolio's market beta in the post-crash state? (b) The market then rebounds by 40 percent. What is momentum's return from the beta channel alone? (c) Sketch momentum's payoff as a function of the market return across the two states. Which option position does it resemble, and what does that imply about the sign of its expected return under any model in which crash insurance is priced? (d) A manager proposes hedging the crash by scaling the position down when realized volatility is high. Explain the mechanism by which this could work, and name one cost of it that §6.8 would emphasize.

***

## Selected Solutions

*Solutions to Problems 1 and 2 follow. Solutions to the remainder are in the instructor materials.*

**Problem 1.**

(a) The line through P and Q has slope $$\lambda = (0.15 - 0.07)/(1.2 - 0.4) = 0.08/0.8 = 10$$ percent per unit of loading. Its intercept is $$0.07 - 0.4(0.10) = 3$$ percent. So the pricing line is $$E\[r] = 3 + 10\times\beta$$, in percent.

(b) At $$\beta = 0.8$$ the line gives $$3 + 8 = 11$$ percent. S offers 10 percent, so S is **dear** — its expected return is one percentage point too low, meaning its price is too high.

(c) A portfolio of P and Q with weight $$w$$ in Q has loading $$0.4 + 0.8w$$; setting this to 0.8 gives $$w = 0.5$$. So hold 50 percent P and 50 percent Q — loading 0.8, expected return $$0.5(7) + 0.5(15) = 11$$ percent — and short an equal dollar amount of S. Net investment: $$+0.5 + 0.5 - 1 = 0$$. Net loading: $$0.5(0.4) + 0.5(1.2) - 0.8 = 0.2 + 0.6 - 0.8 = 0$$. Expected return: $$11 - 10 = +1$$ percent of the gross position.

(d) On $250 million a side, the position pays $2.5 million a year. It is riskless because all three portfolios are well diversified, so their residuals have vanished; the only remaining source of return variation is the single factor, and the position's exposure to it is exactly zero. What must fail for the profit to be risky: the residuals must not in fact be negligible — which is exactly what happens if the "portfolios" are individual securities, or if the residuals are correlated across the three legs, or if the factor structure has a fourth factor the analyst has not modeled and on which the position's loading is not zero.

(e) The short leg is where the mispricing sits, so **S's price falls** as arbitrageurs sell it. Its expected return must rise from 10 percent to 11 percent — a one-percentage-point increase — for the position's expected profit to reach zero. Equivalently, in a one-period setting with a fixed expected payoff, S's price must fall by about 0.9 percent, since $$1.10/1.11 - 1 \approx -0.009$$.

**Problem 2.**

(a) With $$\bar\beta = 1.25$$ and $$\sum\_i(\hat\beta\_i - \bar\beta)^2 = (-0.75)^2 + (-0.25)^2 + 0.25^2 + 0.75^2 = 1.25$$, each month's slope is the cross-product divided by 1.25 and the intercept is $$\bar r\_t - \hat\lambda\_t\bar\beta$$. The results:

**Table 6.6: Month-by-month Fama-MacBeth estimates in Problem 2**

| Month | $$\hat\alpha\_t$$ (%) | $$\hat\lambda\_t$$ (%) |
| ----- | --------------------- | ---------------------- |
| 1     | 0.3                   | 1.0                    |
| 2     | 0.1                   | −0.6                   |
| 3     | −0.2                  | 2.0                    |
| 4     | 0.4                   | −0.2                   |
| 5     | 0.0                   | 1.2                    |

*Source: Author's calculation from Table 6.5.*

(b) $$\hat\lambda = (1.0 - 0.6 + 2.0 - 0.2 + 1.2)/5 = 3.4/5 = 0.68$$ percent per month. The sample standard deviation of the five monthly estimates is 1.064, so $$\mathrm{se}(\hat\lambda) = 1.064/\sqrt{5} = 0.476$$ and $$t = 0.68/0.476 = 1.43$$. For the intercept, $$\hat\alpha = 0.6/5 = 0.12$$ percent per month, with sample standard deviation 0.239, standard error 0.107, and $$t = 1.12$$.

(c) Annualized, $$\hat\lambda = 12 \times 0.68 = 8.2$$ percent — an economically large premium. It is not statistically distinguishable from zero at any conventional level, since $$t = 1.43$$, and it comes nowhere near §6.4's multiple-testing hurdle of 3.0 or Bonferroni's 3.78. But the sample-size objection is the serious one here, not the multiple-testing objection: with $$T = 5$$ the standard error is enormous by construction, and no hurdle is informative about a single estimate this noisy. The multiple-testing hurdle is the binding constraint when $$T$$ is large and the number of *specifications* is what has been inflated — a different failure mode, and the one §6.4 is about.

(d) The four average excess returns are 0.76, 0.38, 1.08 and 1.66 percent. Regressing these on the betas gives an intercept of 0.12 and a slope of 0.68 — exactly the answers in (b). The agreement is not a coincidence: OLS is a linear operator, the same regressor matrix is used in every month, and the average of the per-month coefficient vectors is therefore the coefficient vector of the averaged dependent variable. The equality would fail if the loadings varied across months — if $$\hat\beta\_{i,t}$$ were re-estimated on a rolling window, as any implementable design requires — in which case the two procedures answer different questions and the Fama-MacBeth version is the one that gives correct standard errors.

***

## Data Exercise: Value and Momentum by Decade, and the Crash of 2009

**Part A — The premia, decade by decade (free data: Ken French's data library).** From Kenneth French's data library at Dartmouth, download the monthly **Fama/French 3 Factors**, the monthly **Momentum Factor (Mom)**, and the monthly **Fama/French 5 Factors (2x3)**. Use the longest sample each file offers.

1. For `HML`, `SMB` and `Mom`, compute the average monthly return, its standard deviation, its $$t$$-statistic, and the annualized Sharpe ratio, over the full sample and then separately by decade. Present one table with decades as rows and factors as columns.
2. Identify the decades in which each factor's average return was negative. For HML, confirm the long flat stretch after 2007 and locate the month in which it ends. State the full-sample $$t$$-statistic for HML computed on data through 2006 and again through the end of your sample, and comment on what the change does to your confidence in the premium.
3. Repeat step 1 for `RMW` and `CMA` over their available sample. Then regress `HML` on the other four five-factor returns and report the intercept with its $$t$$-statistic. This is Fama and French's (2015) redundancy result; state whether it holds in your sample.

**Part B — The momentum crash of 2009.**

4. Extract the monthly `Mom` series for 2008 and 2009 and print it. Identify the three worst consecutive months and compute the cumulative return over them. Compare with the cumulative return of `Mkt-RF` over the same three months.
5. Using French's ten momentum decile portfolios (**Portfolios Formed on Prior 2-12 Returns**), compute the cumulative return of the top decile and of the bottom decile over March-May 2009 separately. Which leg produced the loss? Verify that this is the mechanism §6.2 describes rather than a failure of the winner portfolio.
6. Estimate the momentum factor's market beta on a 24-month rolling window across 1927-present. Plot it against the trailing two-year market return. Does the beta fall sharply after market declines, as Daniel and Moskowitz predict? Mark 1932 and 2009 on the plot.
7. Construct a volatility-scaled momentum strategy: each month, scale the `Mom` position by (target volatility) / (realized volatility over the prior six months), capping leverage at 2. Report the Sharpe ratio, the worst three-month drawdown, and the skewness, against the unscaled factor. Write one paragraph on what §6.8 would say about running this strategy with other people's redeemable money.

**Part C ★ (if you have WRDS).** Rebuild HML and UMD from CRSP and Compustat directly, following the Fama-French (1993) 2×3 construction described in §6.2, and verify that your series correlates above 0.95 with French's. Then split the value premium by the share of each stock's market capitalization held by institutions (Thomson Reuters 13F) and by the stock's Amihud illiquidity measure. Section 6.8 predicts the premium is larger where arbitrage capital is scarcer and trading is more costly; test it, and report whether the post-2007 disappearance of the value premium is concentrated in the liquid, heavily institutionally held part of the cross-section — the part where §6.8's capital would have arrived first.
