> For the complete documentation index, see [llms.txt](https://laurence-wilse-samson.gitbook.io/textbooks/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://laurence-wilse-samson.gitbook.io/textbooks/financial-economics-claims-prices-holders/appendices/appendix_b_data_sources.md).

# Appendix B: Data Sources and Measurement

*Appendices — Financial Economics: Claims, Prices, and Holders*

***

Every empirical claim in this book rests on a file somebody built, under conventions somebody chose, for purposes that were usually not this book's. This appendix is the record of those files: what each measures, what it silently omits, whether you can get it without a budget, and where in the book it does its work.

## B.0 How to read this appendix

Each source gets the same five fields in the same order, so that the appendix reads as a reference rather than an essay: **what it measures**, **what it misses**, **free or licensed**, **access path**, and **where this book uses it**. Major sources get five paragraphs; minor ones get the five fields compressed into one. Nothing appears here that the book does not use, and nothing the book uses should be absent — if you find a number in a chapter whose provenance you cannot trace to an entry below, that is a defect, and the author would like to hear about it.

**The free-data-first rule.** The baseline version of every data exercise in this book runs on data anyone can download without a subscription, an affiliation, or a signed license. This is not an aesthetic preference: a textbook whose exercises require a $50,000 site license teaches that empirical finance is something done by people at wealthy institutions, and that lesson is both false and corrosive. The free path is therefore the main path — Kenneth French's library, FRED, the Financial Accounts (Z.1), EDGAR, and Shiller's long-run series carry the great majority of what Chapters 2 through 20 ask.

**The ★ convention.** A section, exercise, or extension marked ★ requires either PhD-track machinery or licensed data. Where the star marks data, the entry names the free substitute. The rule for the author is strict: no ★ exercise may establish a result the reader needs in order to follow the argument. Licensed data buy precision, longer histories, and security-level detail; they do not buy the mechanism, and if a chapter's mechanism is visible only to a WRDS subscriber, the chapter is written wrong.

**"Free" is a property that changes, and it changes in one direction.** Data posted openly get moved behind registration; registration becomes membership; membership becomes an enterprise agreement. One source here — the AAII Sentiment Survey, §B.4 — has moved back and forth within the life of the literature that uses it. Every entry states its access status as of drafting. If an access path below is dead when you read it, the appendix is not wrong about what the data are; it is out of date about who may have them.

**Cite the producer, not the mirror.** FRED and the World Bank are convenience layers over other people's statistics, and a source line reading "FRED" says where you clicked, not what you measured. The book's convention names the producing agency and the series, with the delivery mechanism second: "ICE BofA option-adjusted spreads, via FRED."

**Vintages.** Financial statistics are revised. The Z.1 rewrites its own history four times a year; Compustat restates; CRSP reissues; French's library silently inherits both. A result that cannot name its vintage is an anecdote with decimal places. Section B.7 prints the workflow that stamps every generated table with its release date, and lists the exhibits still awaiting it.

**Redistribution.** The book ships code, not data. Licensed sources forbid redistribution, which is why no repository here contains a CRSP extract, and why the best evidence in Chapter 15 — the account-level brokerage records behind Odean's results — cannot be an exercise at all. Free sources ship as a URL and a fetch script, so a reader gets the current vintage.

![Figure B.1: What each source covers](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-0d2dcfaa4b9d7fd3fc61fbf997f3ecaccc596157%2Ffig_B_01_source_coverage.png?alt=media)

**Figure B.1: What each source covers.** Table B.1 read as coverage rather than as a list: eighteen of its sources against the seven things this book needs to measure, with each cell marked free or licensed. Two readings. Every column has at least one free row, which turns the free-data-first rule from an aspiration into a checkable fact about the sources — there is no mechanism in this book that is visible only to a subscriber, and the firm-fundamentals column is carried by the SEC's own XBRL data sets rather than by Compustat. And the licensed cells cluster in two columns, security-level history and intraday or options data, which is where money buys precision and length rather than mechanism: a licensed source buys a longer history, a cleaner identifier, or a finer time stamp, and none of those is what makes an argument in this book work. The rows are also sparser than a reader expects. Most sources do one thing, which is why almost every exercise in the book joins two of them, and why §B.7's vintage discipline matters as much as it does. *Source: Author's construction from Table B.1.*

**Table B.1: The sources used in this book, free or licensed**

| Source                                           | §   | Free or licensed                       | Unit of observation          | Frequency                    | Coverage begins                 |
| ------------------------------------------------ | --- | -------------------------------------- | ---------------------------- | ---------------------------- | ------------------------------- |
| Financial Accounts of the US (Z.1)               | B.1 | Free                                   | Sector × instrument          | Quarterly                    | 1945 annual, 1952 quarterly     |
| Distributional Financial Accounts (DFA)          | B.1 | Free                                   | Wealth group × instrument    | Quarterly                    | 1989                            |
| Survey of Consumer Finances (SCF)                | B.1 | Free                                   | Family                       | Triennial                    | 1983 (modern series)            |
| FRED and ALFRED                                  | B.1 | Free                                   | Series                       | Varies                       | Varies                          |
| BEA national accounts (NIPA)                     | B.1 | Free                                   | National aggregate           | M / Q / A                    | 1929 annual, 1947 quarterly     |
| CRSP                                             | B.2 | **Licensed**                           | Security × day or month      | Daily, monthly               | December 1925                   |
| Compustat                                        | B.2 | **Licensed**                           | Firm × fiscal period         | Annual, quarterly            | 1950 annual                     |
| Kenneth French data library                      | B.2 | Free                                   | Portfolio × day or month     | Daily, monthly, annual       | July 1926                       |
| Shiller long-run series (`ie_data`)              | B.2 | Free                                   | Market aggregate × month     | Monthly                      | January 1871                    |
| IBES analyst forecasts                           | B.2 | **Licensed**                           | Analyst × firm × forecast    | Event, monthly               | 1976 (US)                       |
| TRACE, enhanced academic file                    | B.2 | **Licensed**                           | Trade                        | Intraday                     | July 2002                       |
| TRACE, FINRA end-of-day summaries                | B.2 | Free                                   | Bond × day                   | Daily                        | 2002                            |
| OptionMetrics IvyDB US                           | B.2 | **Licensed**                           | Option × day                 | Daily                        | January 1996                    |
| WRDS (access platform)                           | B.2 | **Licensed**                           | n/a                          | n/a                          | n/a                             |
| Form 13F, via EDGAR                              | B.3 | Free                                   | Manager × security × quarter | Quarterly                    | 1999 electronic                 |
| Thomson/Refinitiv s34 and s12                    | B.3 | **Licensed**                           | Manager or fund × security   | Quarterly                    | 1980                            |
| Forms N-PORT and N-CEN, via EDGAR                | B.3 | Free                                   | Fund × security              | Monthly, published quarterly | 2019                            |
| Form 5500 (DOL/EBSA)                             | B.3 | Free                                   | Plan × year                  | Annual                       | 1999 research files             |
| NAIC Schedule D                                  | B.3 | **Licensed**                           | Insurer × security           | Annual                       | Varies by vendor                |
| Investment Company Institute (ICI)               | B.3 | Free at aggregate level                | Industry aggregate           | Weekly, monthly, annual      | 1984                            |
| NY Fed primary dealer statistics (FR 2004)       | B.3 | Free                                   | Dealer sector aggregate      | Weekly                       | 1998 in current form            |
| CFTC Commitments of Traders                      | B.3 | Free                                   | Trader category × contract   | Weekly                       | 2000 weekly, 2006 disaggregated |
| Treasury International Capital (TIC)             | B.3 | Free                                   | Country × instrument         | Monthly, annual benchmark    | 1974                            |
| Office of Financial Research monitors            | B.3 | Free                                   | Sector aggregate             | Weekly, quarterly            | 2013                            |
| Treasury CMT yields and Fed H.15                 | B.4 | Free                                   | Maturity × day               | Daily                        | 1953                            |
| Gürkaynak-Sack-Wright zero curve                 | B.4 | Free                                   | Maturity × day               | Daily                        | 1961 nominal, 1999 TIPS         |
| ACM term premium (NY Fed)                        | B.4 | Free                                   | Maturity × day               | Daily                        | 1961                            |
| ICE BofA index OAS, via FRED                     | B.4 | **Free only for a rolling \~3 years**  | Index × day                  | Daily                        | 1996-2023 history withdrawn     |
| Federal Reserve H.4.1 and SOMA holdings          | B.4 | Free                                   | Security or aggregate × week | Weekly                       | 2003 security-level             |
| Bank of England statistical database             | B.4 | Free                                   | Series                       | Daily, monthly               | Varies                          |
| Pástor-Stambaugh liquidity series                | B.4 | Free                                   | Market × month               | Monthly                      | 1962                            |
| He-Kelly-Manela intermediary factor              | B.4 | Free                                   | Sector × month               | Monthly, quarterly           | 1970                            |
| Baker-Wurgler sentiment index                    | B.4 | Free                                   | Market × month               | Monthly                      | 1965                            |
| Yale ICF confidence indices                      | B.4 | Free, registration                     | Investor type × month        | Monthly                      | 1989                            |
| AAII Sentiment Survey                            | B.4 | Free current, registration for history | Survey week                  | Weekly                       | July 1987                       |
| Goyal-Welch predictor dataset                    | B.4 | Free                                   | Market × month or year       | Monthly, annual              | 1871 annual, 1927 monthly       |
| Moody's and S\&P default studies                 | B.5 | Free summaries, **licensed** microdata | Rating cohort × year         | Annual                       | 1920 Moody's, 1981 S\&P         |
| Call Reports (FFIEC)                             | B.5 | Free                                   | Bank × quarter               | Quarterly                    | 1976 machine-readable           |
| FR Y-9C holding company reports                  | B.5 | Free                                   | Holding company × quarter    | Quarterly                    | 1986                            |
| Supervisory stress test disclosures              | B.5 | Free                                   | Bank × scenario × year       | Annual                       | 2009                            |
| Leveraged loan compilers (LCD, LSTA)             | B.5 | **Licensed**                           | Loan or deal                 | Continuous                   | 1997                            |
| EDGAR                                            | B.6 | Free                                   | Filing                       | Event                        | 1993, universal from 1996       |
| Jay Ritter's IPO data                            | B.6 | Free                                   | IPO, or year                 | Annual tables                | 1960 counts, 1980 pricing       |
| SIFMA issuance statistics                        | B.6 | Free summaries                         | Market × period              | Monthly, quarterly           | 1996                            |
| Treasury Monthly Statement of the Public Debt    | B.6 | Free                                   | Security × month             | Monthly                      | 1953                            |
| World Bank Global Financial Development Database | B.6 | Free                                   | Country × year               | Annual                       | 1960                            |
| BIS credit and debt securities statistics        | B.6 | Free                                   | Country × quarter            | Quarterly                    | 1952                            |

*Source: Author's compilation. Coverage-start dates are the earliest data in the current vintage of each product and are not always the dates a given researcher can use; see the individual entries.*

**Table B.2: The licensed sources and their free substitutes**

| Licensed source               | Free substitute                                                                                     | What the substitution costs                                                                                 |
| ----------------------------- | --------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| CRSP returns                  | Kenneth French library portfolios and factors; Shiller for the market since 1871                    | Portfolio-level only: no individual securities, no identifiers, no custom sorts                             |
| Compustat fundamentals        | SEC Financial Statement Data Sets (XBRL) via EDGAR                                                  | Coverage begins about 2009, items are as-filed rather than standardized, and you do the normalizing         |
| CRSP-Compustat merged link    | EDGAR CIK plus a ticker or CUSIP crosswalk you build                                                | Link errors become your errors; no validated date ranges                                                    |
| IBES forecasts                | 8-K Item 2.02 filing timestamps; naive seasonal-random-walk earnings expectations                   | No analyst expectation, so "surprise" changes meaning; see §B.2                                             |
| TRACE enhanced file           | FINRA end-of-day bond activity summaries; MSRB EMMA for municipals                                  | No trade-level record, no dealer side, no uncapped sizes                                                    |
| OptionMetrics IvyDB           | CBOE free index series (VIX, SKEW, PUT, BXM) and current exchange option chains                     | No historical surface, no per-option greeks, no cross-section of strikes                                    |
| Thomson/Refinitiv s34 and s12 | Form 13F and Form N-PORT parsed from EDGAR                                                          | You do the parsing, the manager mapping, and the split adjustment                                           |
| NAIC Schedule D               | Z.1 tables L.114 and L.115 for insurer aggregates; public insurers' 10-K schedules                  | Sector totals rather than insurer-by-security positions                                                     |
| Moody's DRD, S\&P CreditPro   | The annual default studies' published cohort tables                                                 | Aggregate transition and default rates only; no issuer-level histories                                      |
| Leveraged loan compilers      | Shared National Credit review; Senior Loan Officer Opinion Survey; Fed *Financial Stability Report* | Aggregate and supervisory, published with a lag, no deal-level terms                                        |
| WRDS itself                   | Everything in this appendix flagged free                                                            | Your own extraction, cleaning, and linking code — which is, not incidentally, where most of the learning is |

*Source: Author's compilation.*

***

## B.1 The aggregate accounts

### The Financial Accounts of the United States (Z.1)

**What it measures.** The Z.1 is Chapter 2's map in published form: for every major sector of the US economy, what that sector owns and what it owes, instrument by instrument, in matched stock and flow form, quarterly. Its distinguishing property — Copeland's design decision (Copeland 1952) — is that every financial claim is recorded twice, once as somebody's asset and once as somebody's liability, and the two are forced to reconcile. That constraint is what lets a reader ask "who holds Treasury securities?" and get an answer that sums to the amount outstanding.

Tables are identified by a prefix letter. Within each family, tables numbered in the 100s are organized **by sector** and tables in the 200s **by instrument** — the same data, transposed. This is the most useful single fact about the release's structure and answers most questions of the form "which table do I want?": what a pension fund holds is an L.1xx question, who holds municipal bonds is an L.2xx question.

**Table B.3: The Z.1 table families**

| Prefix        | What it reports                                                                                    | Example                                                        |
| ------------- | -------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- |
| **B**         | Balance sheets, including nonfinancial assets and net worth                                        | B.101, Balance Sheet of Households and Nonprofit Organizations |
| **L**         | Levels of financial assets and liabilities, by sector (L.1xx) or by instrument (L.2xx)             | L.101 households; L.210 Treasury securities                    |
| **F**         | Flows: transactions during the quarter, matched one-for-one to the L tables                        | F.101, household transactions                                  |
| **R**         | Reconciliation: change in level equals transactions plus revaluations plus other volume changes    | R.101                                                          |
| **S**         | Integrated macroeconomic accounts, joining the Z.1 to the BEA's national accounts sector by sector | S.3.a, households and nonprofit organizations                  |
| **D**         | Debt outstanding and debt growth by sector                                                         | D.1, D.3                                                       |
| Supplementary | Detail not carried in the main tables, including some household and nonprofit splits               | Varies by release                                              |

*Source: Financial Accounts of the United States (Z.1), Federal Reserve Board; family definitions from the Financial Accounts Guide.*

The R family is the one students skip and should not. A level changes for two entirely different reasons — somebody bought something, or the thing repriced — and the accounts separate them. Chapter 2's exercise Part B turns on exactly this: the R tables are what let you say how much of the post-1990s rise in household net worth was saving and how much was revaluation. The answer is not close.

**What it misses.** Four things, each load-bearing somewhere in this book.

First, **the household sector is a residual**. For most instruments the Z.1 does not measure what households hold; it measures the total outstanding and what every other identifiable sector holds, and assigns the difference to "households and nonprofit organizations" — which therefore contains domestic hedge funds, private partnerships, personal trusts, and nonprofits, plus everyone else's measurement error. When the household row of L.210 jumped after 2021, practitioners read most of it as hedge fund positions in the cash-futures basis trade, checked against the OFR and CFTC data (§B.3) that see leveraged funds directly. The rule: whenever the household row shows a large direct position in an instrument no ordinary household owns, the residual is doing the work.

Second, **instruments in zero net supply mostly do not appear**. Derivatives net to zero across holders, so the master map has no derivatives row and the notional and gross-market-value data come from elsewhere (BIS, §B.6, and clearinghouse disclosures).

Third, **sector definitions are legal-institutional, not economic**. "Private depository institutions" is a charter category; a dealer subsidiary inside a bank holding company sits in the depository sector or in security brokers and dealers (L.130) depending on where the entity is booked, and the answer has changed as firms reorganized.

Fourth, **the accounts carry a discrepancy and do not hide it**. Sector discrepancies (measured saving against measured net lending) and instrument discrepancies (assets reported against liabilities issued) are printed as lines. A reader who forces holdings to sum exactly to outstanding has silently allocated the discrepancy, which is a choice and should be stated.

**Free or licensed.** **Free**, entirely, with no registration.

**Access path.** The release PDF is at `federalreserve.gov/releases/z1/`, published about ten weeks after quarter end. The full data package — CSVs plus a dictionary — downloads as one ZIP from the same page, and the Data Download Program (`federalreserve.gov/datadownload/`) serves individual series, which is what a pipeline should use.

Series codes look forbidding and are not. `LM153064105.Q` decomposes as: `LM` = level at market value (`FL` = level, `FA` = seasonally adjusted flow, `FU` = unadjusted flow, `FR` = revaluation, `FV` = other volume change); `15` = the household and nonprofit sector; then the instrument code, whose final digit is `5` for an asset and `3` for a liability; then the frequency. That series is direct household holdings of corporate equities — the $25 trillion cell of Chapter 2's Table 2.5. The *Financial Accounts Guide* at `federalreserve.gov/apps/fof/` indexes every table and series and explains the estimation method behind each; Chapter 2's Required reading tells you to read its household-sector description first, and that instruction is not decorative.

**Renumbering between vintages is real and will break your code.** Tables have been renumbered more than once, most consequentially in the 2013 restructuring that renamed the flow of funds accounts the Financial Accounts. Series codes are far more stable than table numbers, which is the argument for building pipelines on codes and citing table numbers only in prose, with the release date attached. Every table number quoted in this book is as of the current vintage; verify before publication, which is what §B.7 is for.

**Where this book uses it.** Chapter 2 entire — Tables 2.3 through 2.6 are Z.1 tables condensed, and its flagship exercise rebuilds them from source — then Chapters 9, 10, 13, 14, 16, 17, 19, and 20, as Table B.4 sets out. It is the most-used source in the book by a wide margin.

**Table B.4: Z.1 tables cited in this book**

| Table          | Contents                                                                | Used in           |
| -------------- | ----------------------------------------------------------------------- | ----------------- |
| B.101, F.101   | Household and nonprofit balance sheet and transactions                  | Chs 2, 14         |
| B.103, F.103   | Nonfinancial corporate business balance sheet and transactions          | Chs 2, 22, 25     |
| L.106, L.108   | Federal government; domestic financial sectors                          | Chs 2, 9, 19      |
| L.109 to L.113 | Monetary authority; depository institutions and credit unions           | Chs 2, 19         |
| L.114, L.115   | Property-casualty and life insurance companies                          | Chs 2, 16         |
| L.116 to L.120 | Private and public pension entitlements, with the DB/DC split           | Chs 2, 14, 16     |
| L.121 to L.124 | Money market funds; mutual funds; closed-end funds and ETFs             | Chs 2, 16, 17     |
| L.130          | Security brokers and dealers                                            | Chs 11, 19        |
| L.204, L.205   | Checkable and time-and-savings deposits                                 | Chs 2, 19         |
| L.210 to L.213 | Treasury; agency and GSE-backed; municipal; corporate and foreign bonds | Chs 2, 9, 10, 13  |
| L.223, L.224   | Corporate equities; mutual fund shares                                  | Chs 2, 14, 17, 20 |

*Source: Financial Accounts of the United States (Z.1), Federal Reserve Board. Table numbers are as of the current vintage and have been renumbered in past releases.*

### Distributional Financial Accounts (DFA)

**What it measures.** The Z.1's household balance sheet distributed across wealth groups — the top 1 percent, the next 9, the next 40, the bottom 50 — quarterly since 1989, with parallel breakdowns by income, age, education, and race (Batty et al. 2019). It is the only quarterly, aggregate-consistent picture of who inside the household sector owns what.

**What it misses.** The limitation this book states three times because readers keep forgetting it: the DFA does not measure the distribution. It takes the Z.1 aggregate and splits it using shares estimated from the Survey of Consumer Finances, interpolated between triennial waves and extrapolated to the current quarter. What the SCF cannot see — the extreme top tail, closely held business valuation, assets held through structures — the DFA cannot see either, and the quarterly frequency is interpolation, not observation. Movements over a few quarters are an aggregate revaluation applied to fixed shares.

**Free or licensed.** **Free**, at `federalreserve.gov/releases/z1/dataviz/dfa/`.

**Where this book uses it.** Chapter 2's data exercise Part D; Chapter 14's Tables 14.2 and 14.4 and its argument that models calibrated to the average household are calibrated to the wrong household; Chapter 5's limited-participation discussion.

### Survey of Consumer Finances (SCF)

**What it measures.** The triennial household wealth survey conducted for the Federal Reserve Board — the best US measurement of the joint distribution of assets, debts, income, and participation. Roughly four to six thousand families per wave in the modern series, which begins in 1983 and is designed for comparability from 1989.

**What it misses, and how to use it correctly.** Three properties determine whether your numbers are right.

The sample is **dual-frame**: an area-probability sample for national representativeness plus a list sample drawn from tax records that deliberately oversamples the wealthy, because wealth is concentrated enough that a simple random sample cannot measure it. Weights are therefore not decoration — an unweighted SCF statistic is meaningless. Replicate weights are provided for sampling variance.

Missing values are **multiply imputed**, and the public file contains five *implicates*: five complete copies of each family differing in the imputed cells. The file therefore has five times as many rows as families, and a researcher who ignores this reports a sample five times too large and standard errors far too small. Rubin's procedure: average the statistic across implicates, then take total variance as the average within-implicate sampling variance plus 1.2 times the variance of the five estimates across implicates, the 1.2 being one plus one-fifth.

The top tail is **truncated by design** — the frame excludes the Forbes 400 — so SCF top shares are a lower bound and should not be compared with capitalization-based estimates without adjustment.

For most purposes the Board's *Changes in U.S. Family Finances* Bulletin article and the SCF Chartbook are the right starting point, and are what Chapter 14's tables check against. Go to the microdata when you need a statistic the Bulletin does not print.

**Free or licensed.** **Free**: microdata, summary extract, codebooks, and replicate weights at `federalreserve.gov/econres/scfindex.htm`.

**Where this book uses it.** Chapter 14 throughout; Chapter 5's participation evidence; and, indirectly, everywhere the DFA appears.

### FRED and ALFRED

**What it measures.** Nothing itself. FRED is a mirror and convenience layer maintained by the Federal Reserve Bank of St. Louis — several hundred thousand series from roughly a hundred producing agencies, with a good API. Its value is that it removes the friction between you and the producer; its danger is the same thing, since it is very easy to plot a series without learning what the producing agency thinks it measures. Definitions are abridged and discontinued series stay up. **ALFRED**, the archival version, is the underrated half: it serves *vintages*, the value of a series as it stood on a historical date, which anything involving data known in real time requires.

**Free or licensed.** **Free**, with an API key issued instantly.

**Where this book uses it.** As the delivery mechanism for Treasury yields, ICE BofA spreads, consumption and GDP, the mortgage rate, and the household debt-service ratio in Chapters 5, 9, 10, 13, 14, and 16 — with *Source:* lines naming the producer and adding "via FRED."

### BEA national accounts (NIPA)

**What it measures.** The national income and product accounts, of which this book needs three things: consumption by major type of product (Table 2.3.5, with nondurables and services the standard measure for consumption-based asset pricing), the deflators (Table 1.1.4), and GDP.

**What it misses, and what it imputes.** The consumption series is not a series of transactions: it contains large imputations — the rental value of owner-occupied housing, financial services furnished without payment — that behave differently from purchased components. There is also a **timing convention** Chapter 5's exercise makes an issue of: consumption is a flow over a period while a return is point-to-point, so pairing them means dating the period's consumption at its beginning, end, or middle, and the estimate of risk aversion moves with the choice (Campbell 2003). And the accounts are **revised**, annually each September and comprehensively every few years, reaching back decades.

**Free or licensed.** **Free**, at `bea.gov`, with an API; main series mirrored on FRED.

**Where this book uses it.** Chapter 5's equity premium calibration; the deflators behind every real series in the book; the finance-share-of-GDP figures in Chapter 2's Table 2.6 and Chapter 27.

***

## B.2 Security-level and firm-level data

### CRSP

**What it measures.** The Center for Research in Security Prices at Chicago Booth maintains the security-level panel on which essentially all US empirical asset pricing rests: prices, returns, shares outstanding, volume, dividends, splits, exchange and share-class codes, and delisting information for US exchange-listed equities. Monthly data begin in December 1925; the daily file traditionally begins in July 1962, though current vintages extend it earlier. Coverage adds AMEX in July 1962, Nasdaq in December 1972, and Arca in the 2000s. Companion products cover Treasury securities — the Fama-Bliss files the term-structure literature uses — mutual funds, and indices.

**What it misses.** Everything not listed on a US exchange: private firms, over-the-counter securities outside the covered tiers, non-US listings, the pre-1926 market. And, more subtly, nothing at all while appearing to, which is the trap. Four points do the damage.

**Survivorship bias is a property of your query, not of the database.** CRSP retains delisted securities forever. Bias enters when a researcher merges with a current index list, requires data present at the end of the sample, or uses an "active companies" screen. A 1995 sample consisting of firms you can still look up today is survivorship-biased data built out of unbiased data.

**Delisting returns are the classic silent error.** A stock leaving the exchange has a final partial-period return and then a delisting return, and the month's total return compounds the two:

$$
1 + r\_{i,t} = (1 + r^{\text{partial}}\_{i,t})(1 + r^{\text{dl}} \_{i,t})
$$

Where the delisting return is missing — disproportionately for performance-related delistings, CRSP codes 500 and 520 through 584 — dropping the observation drops the worst outcomes, and the bias is largest in exactly the small, distressed, high book-to-market corner where the interesting anomalies live. Shumway (1997) documents this and proposes imputing minus 30 percent for missing performance-related delistings; Shumway and Warther (1999) find a larger figure for Nasdaq. Report the convention you used, and note that CRSP restructured its stock file in the early 2020s with a different treatment of delisting returns — check the current data guide rather than inherited code.

**Share and exchange codes define your universe.** "US common stocks" conventionally means share codes 10 and 11 on NYSE, AMEX, and Nasdaq, a screen that excludes ADRs, closed-end funds, REITs, and units. Results are not always robust to relaxing it.

**Prices are closing prices, or negative.** A negative price is the negated bid-ask midpoint, recorded when no trade occurred at the close. Take absolute values deliberately, and remember that a quote midpoint is not a transaction price — the whole subject of Chapter 11.

**Free or licensed.** **Licensed**, almost always through WRDS. There is no free version and no legitimate redistribution.

**Access path.** WRDS web queries, or the WRDS cloud PostgreSQL server via the `wrds` Python package, which is what a reproducible pipeline uses.

**Substitution.** French's library, below, is built from CRSP and Compustat and posted with permission at the *portfolio* level, so anything expressible as a return on a sorted portfolio is free, with conventions documented. Anything requiring individual securities is not, and the honest thing is to say so: the ★ exercises in Chapters 6, 11, and 17 that need firm-level returns cannot be replicated on free data, and each names the portfolio-level exercise showing the same mechanism.

**Where this book uses it.** ★ extensions in Chapters 4 through 7, 11, 12, 17, and 19; and, at one remove, everything in the French library.

### Compustat

**What it measures.** Standardized accounting fundamentals for North American public firms, from S\&P Global Market Intelligence: balance sheet, income statement, cash flow statement, and segments, annually from 1950 and quarterly from the early 1960s, normalized into consistent items across filers and over time. Firms are keyed by `GVKEY`.

**What it misses.** Private firms, and the years before a firm entered the database — which is where the trouble starts.

**Backfill and inclusion bias.** When a firm is added, several years of history are often added with it, so a screen requiring two or three prior years selects on subsequent survival. Banz and Breen (1986) named the problem; Kothari, Shanken and Sloan (1995) showed how much measured book-to-market effect it can manufacture.

**Restatements and look-ahead.** The standard files carry accounting data as most recently reported, not as originally filed, so a firm that restated 2019 earnings in 2021 has restated numbers in the 2019 row — and a strategy formed on that row trades on the future. Point-in-Time products carry as-first-reported vintages; use them for backtests.

**The reporting lag.** Accounting data are known only after filing. The Fama-French convention — match the fiscal year ending in calendar year *t* − 1 to returns from July of *t* through June of *t* + 1 — builds in a six-month cushion, and is why French's portfolios rebalance in June.

**Definitions move under you.** Book equity across the 2001 goodwill rules and the 2019 lease-capitalization rules is not the same construct, and neither is leverage. Davis, Fama and French (2000) document the definition the factor literature uses; French's library documents its own; state yours.

**The merge is where results die.** Linking to CRSP runs through the CRSP-Compustat Merged link table, mapping `GVKEY` to `PERMNO` with link types and validity ranges; restricting to primary links of types `LU` and `LC` within valid dates is the standard discipline. A naive merge on ticker or CUSIP duplicates firm-years, attaches the wrong history across a reverse merger, and produces a cross-section no referee can reproduce.

**Free or licensed.** **Licensed** (WRDS).

**Substitution.** The SEC's Financial Statement Data Sets (§B.6) are quarterly ZIPs of every numeric XBRL fact in every filing, free, from roughly 2009. They are not standardized — you will reconcile `StockholdersEquity` against its including-noncontrolling-interests sibling yourself — but they are as-filed, which makes them *better* than standard Compustat for point-in-time work, and building your own book-equity definition teaches more about book equity than reading `CEQ` ever will.

**Where this book uses it.** ★ extensions in Chapters 6, 12, 22, 23, and 25; and, at one remove, the sorting variables behind every French portfolio.

### Kenneth French's data library

**What it measures.** The free workhorse of Chapters 4 through 7 and of Appendix A's worked examples: monthly, daily, and annual returns on the factors (market minus riskless, SMB, HML, RMW, CMA, momentum), univariate and bivariate sorted portfolios (the 25 size and book-to-market portfolios are the canonical test assets), industry portfolios in five to forty-nine groupings, international sets, and the underused breakpoint files that record where each sort's cut points fell. Factors begin in July 1926.

**What it misses.** Individual securities. The library is a portfolio-level projection of licensed data, which is precisely why it can be free.

**The conventions are the content.** NYSE breakpoints rather than all-stock breakpoints, which keeps the microcap tail from dominating; annual June rebalancing; book equity from the fiscal year ending in *t* − 1; market equity from December of *t* − 1 for the ratio and June of *t* for weighting; momentum from *t* − 12 to *t* − 2 with the gap that excludes short-run reversal. Appendix A §A.2 works through what each choice does, and Box 6.1 makes the deeper point: a literature in which most tests run on one person's pre-built portfolios under one set of conventions has a common-infrastructure problem as well as a multiple-testing problem.

**Practical traps.** The CSVs stack several panels — value-weighted returns, equal-weighted returns, firm counts, average size — in one file with blank-line separators, so a naive `read_csv` silently yields garbage. Returns are in percent, missing values are −99.99 or −999, and files are rebuilt as CRSP and Compustat revise, so results shift across vintages: record the file date.

**Free or licensed.** **Free**, at `mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html`, no registration.

**Where this book uses it.** The baseline exercise of Chapters 4 through 7; Appendix A's worked examples; the benchmark in Chapters 17 and 18.

### Robert Shiller's long-run series

**What it measures.** The `ie_data` workbook — monthly S\&P Composite prices, dividends, earnings, the CPI, a long-term interest rate, and the cyclically adjusted price-earnings ratio — from January 1871. It is the free source for anything needing a century and a half of US equity history.

**What it is, which is not what it looks like.** Pre-1926 prices and dividends come from the Cowles Commission's back-compilation of newspaper quotations (Cowles 1939). Three properties follow. The monthly price is an **average of daily closes**, not a month-end close, so returns computed from it are not returns anyone could have earned and the averaging induces autocorrelation that masquerades as predictability. Dividends and earnings are **interpolated** from lower-frequency data, so their monthly variation is an artifact. The CPI is **spliced** to a pre-1913 wholesale index. The CAPE has its own caveat: its denominator is a ten-year real average of *reported* earnings, whose relation to economic earnings shifted with the write-down and impairment rules (Siegel 2016), so cross-regime comparisons are partly comparisons of accounting regimes.

**Free or licensed.** **Free**, at `shillerdata.com` (moved from the Yale page older citations give).

**Where this book uses it.** Chapter 5's long-run equity premium; §7.4's predictability and volatility-bound material; Chapter 22's valuation anchors; Appendix A §A.5.

### IBES analyst forecasts

**What it measures.** Individual and consensus analyst forecasts of earnings, revenue, and long-term growth, with actuals on the same basis, for US firms from 1976. **Licensed** via WRDS, and used here only in starred exercises.

**What it misses, and the two traps.** The IBES "actual" is a street number — adjusted, excluding what the analysts agreed to exclude — and it is not the GAAP figure in Compustat, so a surprise computed as IBES actual minus IBES consensus is coherent while one mixing an IBES consensus with a Compustat EPS is not. Second, the adjusted files restate historical forecasts into current share terms, and the rounding introduces spurious dispersion and spurious zero-surprises in low-priced or heavily split stocks (Payne and Thomas 2003); use unadjusted files and do your own adjustment.

**Free substitute.** There is none for the forecasts, and this is one of the few places where the free path genuinely stops. What is free is the *timing*: 8-K filings under Item 2.02 carry earnings announcements with SEC timestamps, so a reader without IBES can still run a correctly dated event study (Appendix A §A.4) against a naive seasonal-random-walk expectation. The exercise changes meaning — the response to earnings news rather than to news relative to analysts — and Chapter 7 says so.

### TRACE

**What it measures.** FINRA's Trade Reporting and Compliance Engine: post-trade reports for TRACE-eligible fixed income securities — price, size, time, side — filed by dealers within minutes of execution. It is the most important measurement innovation in the corporate bond literature, because before its phase-in beginning July 2002 there was essentially no public record of what corporate bonds traded at. Coverage of public corporate bonds was substantially complete by early 2005; 144A trades were added in 2014; Treasury reporting to FINRA began in 2017 with aggregate weekly dissemination from 2020.

**What it misses.** Three things that shape what can be measured.

**Trade sizes are capped in the public dissemination** at $5 million for investment grade and $1 million for high yield, displayed as "5MM+" and "1MM+" — precisely the institutional trades that move markets are the ones whose size the public feed hides. The enhanced academic file lifts the caps but arrives with a substantial delay, currently on the order of eighteen months; verify before designing a project around recent data.

**The tape needs cleaning before it means anything.** Cancellations, corrections, and reversals appear as separate records that must be matched to the trades they modify, and interdealer trades are reported by both sides, so naive volume roughly double-counts the dealer-to-dealer segment. Dick-Nielsen's (2009, 2014) filters are the standard, and the difference in measured volume and spreads is large.

**A bond is not a stock.** Most bonds do not trade on most days, so "the price" is the last trade of an unknown size at an unknown time from an unknown side, and Bessembinder et al. (2009) show how much the price convention — last trade, size-weighted average, institutional trades only — changes measured returns and risk. Chapter 11's illiquidity material is partly about why this is unavoidable rather than sloppy.

**Free or licensed.** Split, which is why this entry matters for the book's rule. The **academic enhanced file is licensed**, through WRDS. FINRA's **end-of-day summary files and bond activity reports are free** at `finra.org/finra-data/fixed-income`, giving daily high, low, and last prices and aggregate volume by security; for municipals, the MSRB's EMMA publishes trade-level data free.

**Where this book uses it.** Chapter 10's spread decomposition and Chapter 11's transaction-cost estimates in ★ form; the free FINRA summaries carry Chapter 10's baseline exercise; Chapter 19's March 2020 dealer-inventory narrative.

### OptionMetrics IvyDB US

**What it measures.** End-of-day bid and ask quotes, volumes, open interest, implied volatilities, and greeks for exchange-listed US equity and index options from January 1996, plus an interpolated standardized volatility surface. **Licensed** via WRDS.

**What it misses.** Quotes rather than trades; end-of-day snapshots near but not exactly at the underlying's close, so nonsynchronicity contaminates computed hedge ratios; zero-bid and stale contracts that must be screened; and a surface that is a *model output* — nobody traded at the 30-day, 50-delta point. Report your screens.

**Free substitute.** CBOE publishes VIX, VVIX, SKEW, and the buy-write and put-write index histories free, covering what Chapter 8's baseline exercise needs about the level and shape of index implied volatility; current option chains from exchanges and broker APIs support a cross-sectional smile exercise on today's data, though not a history.

**Where this book uses it.** Chapter 8's ★ smile and variance-risk-premium exercises; Chapter 26's tail-hedging discussion.

### WRDS

**What it measures.** Nothing. Wharton Research Data Services is an access platform — a hosted, uniformly structured layer over dozens of separately licensed databases — plus WRDS-built products that genuinely add value: the CRSP-Compustat Merged link table, the SEC Analytics Suite (which parses EDGAR filings, including 13F holdings, rather than relying on a vendor file), the financial-ratios library, and the beta suite.

**What it misses.** **A WRDS login is not a license to everything on WRDS**: institutions subscribe database by database, so the menu you see is your library's purchase order, and having CRSP but not OptionMetrics is entirely normal. And **its convenience is not neutral**: pre-built merges, default screens, and the fact that running a query is easier than reading a data guide have produced a generation of papers whose sample construction is a WRDS default nobody examined.

**Free or licensed.** **Licensed**, at institutional level, with per-database entitlements; some universities extend access only to students in specific courses. If you are reading this book independently, assume you do not have it.

**Access path.** Web query forms, or the cloud PostgreSQL server via the `wrds` Python package and its R and Stata equivalents.

**Where this book uses it.** As the delivery mechanism for every ★ exercise. Its terms are also why the repository holds fetch-and-build scripts rather than data: redistribution of licensed extracts is prohibited, and a textbook shipping a CRSP snapshot would be a textbook with a legal problem.

**Table B.5: Five ways a firm-level panel lies**

| Failure                | Mechanism                                                                               | Where it bites hardest                                          | Fix                                                                                           |
| ---------------------- | --------------------------------------------------------------------------------------- | --------------------------------------------------------------- | --------------------------------------------------------------------------------------------- |
| Survivorship           | Sample built from entities that still exist, or that have data at the end of the sample | Long-horizon returns; fund performance; anything about distress | Build the universe as of each formation date, never from a current list                       |
| Delisting              | Missing final return on performance-related delistings                                  | Small, distressed, high book-to-market portfolios               | Compound the delisting return; impute the literature's convention where missing and report it |
| Backfill and inclusion | History added when a firm enters the database; screens requiring prior years            | Value and profitability sorts in the early sample               | Point-in-time files, or a screen that does not condition on later existence                   |
| Restatement look-ahead | Standard files carry as-restated, not as-reported, accounting data                      | Any strategy backtest formed on accounting variables            | Point-in-time vintages, or a lag long enough to cover restatement                             |
| Merge error            | Ticker or CUSIP joins across reorganizations; wrong link type or date range             | Everything downstream, invisibly                                | Validated link tables with type and date range enforced; count rows before and after          |

*Source: Author's compilation, drawing on Banz and Breen (1986), Kothari, Shanken and Sloan (1995), and Shumway (1997).*

***

## B.3 Holdings and ownership

This section measures the second half of the book's title. Chapters 16, 17, and 20 are possible only because US institutional holdings are substantially public — an accident of the 1975 amendments to the securities laws that other jurisdictions did not replicate.

### Form 13F

**What it measures.** Under Section 13(f) of the Securities Exchange Act and Rule 13f-1, an institutional investment manager exercising discretion over $100 million or more in "Section 13(f) securities" files Form 13F-HR within 45 days of quarter end, listing each position: issuer, class, CUSIP, market value, shares. Section 13(f) securities are exchange-traded equities, closed-end funds, ETFs, certain equity options, and some convertible debt, enumerated on a list the SEC publishes quarterly. The result is a quarterly census of US institutional equity ownership, free to anyone.

**What it misses.** Everything that makes it a partial view — which Box 20.1 states and every demand-system estimate in Chapter 20 inherits. **Long positions only**: a manager long $2 billion and short $1.9 billion appears as a $2 billion long-only investor. **Reportable securities only**: no cash, bonds, foreign listings, swaps, or futures, so much of a multi-strategy manager's exposure is invisible. **Manager-level aggregation**: one filing can cover hundreds of funds with different mandates, and an index complex's passive and active books arrive as one number. **A 45-day lag and no intra-quarter path**: round-trip trades and window dressing are gone. **Confidential treatment**: managers may have positions withheld temporarily, and Agarwal et al. (2013) show these differ systematically from disclosed ones, so the disclosed panel is not a random subsample. And **the threshold has never been indexed** — $100 million was set in 1978 — so a series of the number of filers is mostly a chart of the price level. The Commission proposed raising it to $3.5 billion in 2020 and later withdrew the proposal; verify the current threshold before quoting it.

**Free or licensed.** **Free**, via EDGAR. This is the book's clearest demonstration that free data are not second-class data: the institutional-demand literature of Chapters 16, 17, and 20 rests on a public filing.

**Access path.** EDGAR full-text search and the daily index files; filings carry a machine-readable XML information table from roughly the early 2010s and plain text before that. The `data.sec.gov` submissions API enumerates a filer's history efficiently. Respect the fair-access limits — ten requests per second, with a declared user agent — or be blocked.

**Where this book uses it.** Chapter 16's institutional-ownership evidence; Chapter 17's measurement of the passive complex; Chapter 20, where holdings are literally the dependent variable; Box 20.1.

### Thomson/Refinitiv s34 and s12

Vendor-cleaned institutional 13F holdings (s34) and mutual fund holdings (s12), with manager-type classifications and history from 1980. **Licensed** via WRDS. The attraction is that somebody else did the parsing and identifier mapping; the problem is documented defects, including share-adjustment errors and coverage breaks around 2010 when the vendor's collection changed, which have produced spurious findings about institutional ownership in papers spanning the break (Ben-David et al. 2021). The free path is now better: parse EDGAR yourself, or use the WRDS SEC Analytics Suite, which does the same parse under license. If you use s34, check your series for the break.

### Forms N-PORT and N-CEN

**Free**, via EDGAR, and underused. Registered funds file monthly portfolio holdings on Form N-PORT, with the third month of each quarter made public, and annual census data on N-CEN. Together they give security-level holdings for the entire US registered fund complex at higher frequency than 13F and including the fixed income that 13F omits. Public data begin in 2019 — short for a panel, ample for the cross-sectional exercises of Chapters 17 and 18. Form N-PX, which since the 2022 rulemaking carries managers' say-on-pay voting records, serves Chapter 24.

### Form 5500

**Free**, from the Department of Labor's EBSA in annual research files. Every private-sector ERISA plan reports participants, assets, contributions, distributions, service providers and fees, with actuarial detail for defined benefit plans in Schedules SB and MB and financial detail in Schedule H — the plan-level ground truth behind Chapter 16's pension material and the DB/DC cross-check in Chapter 2's Table 2.6. Two limitations, both flagged against Chapter 16's table: scope is **private-sector only**, so several trillion dollars of state and local plans are absent and must come from Z.1 L.120; and the lag is roughly two years, since a plan year plus a seven-month deadline plus extensions plus processing adds up.

### NAIC Schedule D

Statutory annual statements filed by US insurers, of which Schedule D Part 1 lists every long-term bond at CUSIP level and Part 2 the equities — the security-level view of the sector Chapter 10 calls the marginal holder of corporate credit. **Licensed**: compilations come from the NAIC and commercial vendors, and arrangements differ across universities, so check what your library actually has. **Free substitutes** carry the aggregate story: Z.1 tables L.114 and L.115, the NAIC's Capital Markets Bureau reports, and the investment schedules in public insurers' 10-K filings. What you lose is the security-level cross-section that identifies regulatory-driven demand — the design behind Becker and Opp's (2013) study of the 2009-2010 change in capital treatment of structured securities.

### Investment Company Institute

**Free at the aggregate level**: weekly estimated flows, monthly *Trends in Mutual Fund Investing*, quarterly retirement tables, and the annual *Investment Company Fact Book*. It is the cleanest free source for the fund and retirement rows of Chapter 2's Tables 2.4 and 2.6 and for the IRA totals Chapter 2's exercise Part C needs, since IRAs sit awkwardly in the Z.1. Flow figures are estimates covering most but not all of the industry, and are revised; the member-level data behind them are not public.

**Which ICI release is free is a live distinction, and it caught this book.** The member statistical reports — including the "Combined Active and Index Mutual Funds and ETFs" workbook, the obvious source for the active-against-index split Chapter 17 is about — are **gated**: the spreadsheet returns an HTTP 403 and the landing page asks for a login. What is genuinely open is the *Fact Book*'s own data tables, posted one spreadsheet per table at `icifactbook.org/xls/26-fb-table-NN.xlsx`, roughly sixty-nine tables in the current edition, with no key and no registration. Between them, table 42 (active and index mutual funds, total net assets) and table 11 (exchange-traded funds by type) carry the domestic-equity active/index split annually from 1993, which is the series behind Chapter 17's Figure 17.1. Two cautions. The filenames carry the edition year, so they move each spring and a stable reference means a committed snapshot, not a URL. And ICI publishes ETF assets *by objective* and *by strategy* but never by both, so any domestic-equity index total built this way silently includes actively managed ETFs; the honest treatment is to bound it, which is what Figure 17.1's shaded band does. This is §B.0's warning about "free" in its narrowest form: not a source that closed, but two releases from the same producer on opposite sides of the line.

### NY Fed primary dealer statistics

**Free**, from the New York Fed: weekly aggregate positions, financing, and transaction volumes of the primary dealers by instrument and maturity bucket — the closest direct measurement of dealer balance-sheet capacity, and the empirical anchor of Chapter 19's intermediary story and Chapter 11's account of who supplies immediacy. The reporting form (FR 2004) was substantially revised in 2013, adding and redefining categories, so series break there and splicing needs a footnote.

### CFTC Commitments of Traders

**Free**, weekly: aggregate futures and options positions by trader category as of Tuesday, released Friday. Three report families coexist — the legacy commercial/non-commercial split, the disaggregated report (producer, swap dealer, managed money, other) from 2006, and the Traders in Financial Futures report. Two cautions: categories were **reclassified** when the disaggregated reports began, so long series splice across a definitional break; and classification is at the *entity* level and self-reported, so a bank's hedging and speculative books land in one bucket. Used in Chapter 8's commodity material and, with the OFR data below, as the cross-check on the household-residual Treasury position in Chapters 2 and 9.

### Treasury International Capital, and the Office of Financial Research

Two free sources carrying the rest-of-world column of Chapter 2's master table and the leveraged-fund correction to its household row. **TIC** reports cross-border holdings and transactions in US securities monthly, with an annual benchmark survey that revises the monthly estimates substantially and splits foreign Treasury holdings into official and private — the split Box 2.2 and Chapter 9's Table 9.3 both need. Its weakness is **custodial attribution**: securities are recorded against the custodian's country, so holdings intermediated through financial centers are misassigned badly enough that country-level figures are lower bounds for several large holders. The **OFR** publishes free monitors built on data nobody else sees, including hedge fund aggregates from Form PF and a short-term funding monitor covering repo — the published route to the leveraged-fund positions the Z.1 buries in the household residual.

**Table B.6: What each holdings source sees**

| Source             | Holder sectors covered                                   | Security-level?                    | Lag                         | Free?        |
| ------------------ | -------------------------------------------------------- | ---------------------------------- | --------------------------- | ------------ |
| Z.1 (L.2xx tables) | All sectors, by construction                             | No — instrument totals             | \~10 weeks                  | Free         |
| Form 13F           | Managers over the threshold; US-listed equity only       | Yes                                | 45 days                     | Free         |
| Form N-PORT        | Registered funds (mutual funds, ETFs), all asset classes | Yes                                | \~60 days, quarterly public | Free         |
| Form 5500          | Private-sector ERISA plans                               | No — plan aggregates and schedules | \~2 years                   | Free         |
| NAIC Schedule D    | US insurers                                              | Yes                                | \~1 year                    | **Licensed** |
| TIC                | Rest of world, by custodial country                      | Partly, in the benchmark surveys   | 2 months, annual benchmark  | Free         |
| NY Fed FR 2004     | Primary dealers, in aggregate                            | No                                 | \~1 week                    | Free         |
| CFTC COT           | Futures trader categories                                | No — contract aggregates           | 3 days                      | Free         |
| OFR monitors       | Hedge funds, repo market participants                    | No                                 | Varies                      | Free         |

*Source: Author's compilation.*

***

## B.4 Prices, rates, and spreads

### Treasury constant-maturity yields, and what they are not

**Free.** The Treasury publishes a daily par yield curve and the Federal Reserve republishes it in the H.15; both are mirrored on FRED as the `DGS` series. The construction is a spline fitted to indicative bid yields on the most recently issued securities at a fixed afternoon hour, interpolated to constant maturities. Two consequences. A constant-maturity yield is **not a bond you can buy** — it is a fitted par yield on a hypothetical ten-year security, and a strategy "earning" it cannot be implemented. And because the inputs are on-the-run issues, the curve inherits their liquidity premium, the wedge Chapter 9 spends a section on. For zero-coupon yields and forward rates — what Chapter 9's theory concerns — use the **Gürkaynak-Sack-Wright** curve, posted free by the Board with daily parameters from 1961 and a TIPS companion from 1999 (Gürkaynak, Sack and Wright 2007). The licensed alternative is the CRSP Fama-Bliss file; the free one is longer and better documented.

### ACM term premium

**Free**, from the New York Fed: daily decompositions of Treasury yields into expectations and term premium, from Adrian, Crump and Moench (2013). The standing warning, to be printed whenever it is plotted: **this is a model output, not an observation**. Term premia are the residual of a no-arbitrage affine model estimated on a particular sample with a particular number of factors, and rival decompositions — the Board's Kim-Wright estimates, survey-based measures — differ at some dates by enough to reverse a narrative. Plot two, or say which one you chose.

### ICE BofA index option-adjusted spreads

**No longer free at any useful history, and this is a change since the book was drafted.** FRED's rating-bucket series — `BAMLC0A4CBBB` and its siblings — now return only a rolling window of roughly three years. The truncation was verified three ways: the keyless CSV endpoint, the keyed API, and ALFRED, where the whole family was reset to a 2023 start together with its vintage record, so there is no archived vintage to request and no discontinued twin holding the past. A reader following this appendix as written in 1996 terms will not find what it promised, and the exhibits that needed the history have been rebuilt on substitutes that are labelled as such — Moody's seasoned Aaa and Baa yields less the ten-year Treasury, which reach back to 1919 and are *not* option-adjusted. Where the spread level itself matters rather than its shape, no free substitute exists at all above investment grade. Three construction facts still govern interpretation of the series where it is available. The **option-adjusted spread** is a model output, computed with the provider's volatility and prepayment assumptions, so the series is a joint statement about prices and about that model. Index membership **rebalances monthly** against rating and size rules, so a bond downgraded out of investment grade leaves at month end — part of every measured spread change around a downgrade wave is survivorship at rebalance rather than repricing. And a **spread is not a return**: Chapter 10's credit-risk-premium arithmetic needs returns, expected losses, and recoveries, of which the spread is only the first ingredient.

### Federal Reserve H.4.1 and SOMA holdings

**Free.** The weekly H.4.1 gives the Federal Reserve's balance sheet; the New York Fed publishes System Open Market Account holdings at CUSIP level weekly. Together they are the "Federal Reserve" cell of Chapter 2's Table 2.5 and the central-bank line in Chapter 9's Treasury decomposition — and they let a reader do something rare in this book: measure one very large holder's position exactly, at security level, with no lag and no license.

### Bank of England statistical database

**Free**, at `bankofengland.co.uk/boeapps/database/`, together with the separately posted nominal and real gilt curves — the source for Chapter 16's September 2022 liability-driven investment episode, alongside the pension-sector positions in the Bank's *Financial Stability Report*.

### Author-posted research series

Five series this book uses that exist because their authors keep posting them. All are **free**, and all carry the same caveat: a series maintained by a person rather than an institution can stop, move, or change definition without notice, and several have.

**Pástor-Stambaugh liquidity** (Pástor and Stambaugh 2003): monthly aggregate liquidity from 1962, in three distinct forms — level, innovation, and traded factor — that papers cite interchangeably and should not; the history is revised as its CRSP inputs revise. **He-Kelly-Manela intermediary capital** (He, Kelly and Manela 2017): the equity capital ratio of primary dealers' holding companies and the traded factor built from it, the object at the center of Chapter 19; availability is flagged for re-checking before publication. **Baker-Wurgler sentiment** (Baker and Wurgler 2006, 2007), on Wurgler's NYU page in raw and orthogonalized forms — the component set has changed since the original paper, NYSE turnover having been dropped, so the current file is not the 2006 series extended. **Yale ICF Stock Market Confidence Indices**: four indices, separately for individual and institutional investors, monthly as six-month moving averages, free to fetch and confirmed at byte level — but Yale states the material may not be reproduced or distributed without written permission, which is a constraint on publishing a figure drawn from them rather than on reading them, and is why this book draws none; Chapter 15 uses crash confidence, whose status as an overstated subjective probability is the point of the exercise. **Goyal-Welch predictors** (Welch and Goyal 2008), updated and posted by Amit Goyal — dividend and earnings ratios, book-to-market, term and default spreads, net equity expansion, the bill rate — which is why §7.4's exercise runs with no license at all.

### AAII Sentiment Survey

**Free, and confirmed so on 7 September 2026.** The American Association of Individual Investors has published a weekly bullish/neutral/bearish survey of its members since July 1987. The historical workbook at `aaii.com/files/surveys/sentiment.xls` returns 2,039 weeks from July 1987 with no registration. The test that settles it is worth stating, because the failure mode is silent: check the `Content-Type`, not the status code. A gated response comes back `200 text/html` and will poison a spreadsheet reader without raising anything; the live file answers `application/vnd.ms-excel` with an OLE2 magic number. Figure 15.4 is built on it. This remains the appendix's worked example of its own warning in §B.0 — the terms here have moved in both directions over the years, and a future reader should re-run the test rather than trust this paragraph.

***

## B.5 Credit, default, and ratings

### Moody's and S\&P annual default studies

**What they measure.** Cumulative default rates and rating transition matrices by initial rating and cohort year, plus recovery rates by seniority — Moody's back to 1920, S\&P to 1981. They are the foundation of Chapter 10's gap between physical and risk-neutral default probabilities and of its credit-triangle arithmetic.

**What they miss, and four things to get right.** Rates use the **static-pool** or cohort method, and the treatment of issuers whose ratings are **withdrawn** changes the answer materially, since withdrawal is not random with respect to credit quality. Rates are **issuer-weighted by default**, and volume-weighted rates answer a different question. **Recovery** usually means the trading price about thirty days after default — a market expectation of ultimate recovery rather than the recovery itself. And the **rating scale is not a constant**: a Baa assigned in 1935 came out of a process that no longer exists.

**Free or licensed.** Split: the **annual studies and their summary tables are free**, typically behind registration; the **issuer-level histories are licensed** (Moody's Default and Recovery Database, S\&P's CreditPro and RatingsXpress).

**Where this book uses it.** §§10.3 and 10.6; the calibration behind Chapter 26's credit stress example.

### Call Reports (FFIEC)

**What it measures.** Every insured US commercial bank and savings institution files a Consolidated Report of Condition and Income quarterly: full balance sheet and income statement plus schedules covering loans by type, securities by category and by held-to-maturity or available-for-sale designation, deposits including insured and uninsured splits, commitments, derivatives, and regulatory capital. Machine-readable data run from the 1970s, with bulk downloads from 2001. Bank holding companies file the parallel FR Y-9C above a size threshold raised repeatedly (to $3 billion in 2018). It is a complete, free, quarterly, institution-level census of every bank balance sheet in the country; very little else in the financial system is measured this well.

**What it misses, and the four things that break panels.** **The reporting entity is the bank, not the group** — a large organization comprises a holding company, insured banks, a broker-dealer, and foreign affiliates, and dealer inventories often sit outside the Call Report, which is why Chapter 19 uses FR 2004 and Y-9C for dealers. **RCFD against RCON**: nearly every item exists in a consolidated-including-foreign-offices version and a domestic-only version, small banks file only one, and mixing them builds in a discontinuity correlated with size. **Mergers destroy panels**: a merged bank's history simply ends, so a consistent panel needs National Information Center structure histories and merger adjustments, keyed on `RSSD` ID rather than name or charter number. **The form changes**: items are added, split, and retired as regulation changes, so check an item's history before running a long regression.

**Free or licensed.** **Free**: FFIEC Central Data Repository bulk downloads; the FDIC's BankFind and Statistics on Depository Institutions with a public API; the National Information Center for structure; the Chicago Fed for some historical files.

**Where this book uses it.** Chapter 19's bank balance-sheet material, including the 2023 held-to-maturity unrealized-loss episode visible in Schedule RC-B and the AOCI opt-out election; Chapter 25's evidence on how much corporate borrowing is bank-funded; Appendix C's worked bank balance sheet.

### Supervisory stress test disclosures

**Free.** The Federal Reserve publishes the annual scenarios each winter — baseline and severely adverse paths for a couple of dozen macro-financial variables — and the results each June, with projected losses by portfolio and post-stress capital ratios for each firm. What is **not** public is the underlying FR Y-14 supervisory collection, which students routinely assume is. Used in Chapter 19 and in Chapter 26's treatment of regulatory risk measurement as a portfolio constraint.

### Leveraged loan compilers

Chapter 24's covenant-lite share and the syndicated-loan pricing series in Chapters 10 and 25 come from commercial compilers — LSEG LPC, PitchBook LCD and their predecessors — which are **licensed**, expensive, and mutually inconsistent, since each defines "institutional leveraged loan" and "covenant-lite" its own way. Any figure from them must name the compiler and the vintage, which is what the open marker on Chapter 24's Table 24.2 records. **Free substitutes**, all aggregate: the interagency Shared National Credit review, the Senior Loan Officer Opinion Survey, and the syndicated-loan sections of the Fed's *Financial Stability Report*.

***

## B.6 Issuance, firms, and filings

### EDGAR

**What it measures.** Every document filed with the SEC since the mid-1990s, free, immediately, in bulk. For this book: **10-K** and **10-Q** for financial statements; **8-K** for material events, with Item 2.02 carrying earnings announcements and their timestamps; **DEF 14A** for compensation, boards, and shareholder proposals; **13F-HR** for institutional holdings; **SC 13D** and **13G** for beneficial ownership above five percent, with the 13D deadline shortened to five business days in 2024; **Form D** for private placements, the only systematic window onto exempt-offering volume; **N-PORT**, **N-CEN**, and **N-PX** for registered funds; and **S-1** and **424B** for offerings.

**What it misses.** Private companies that never register; the pre-1993 record; and structure — a filing is a document, and turning documents into a panel is the work. The **Financial Statement Data Sets**, quarterly ZIPs of every numeric XBRL fact from every filing, solve most of that for accounting data from roughly 2009 and are the free Compustat substitute named in Table B.2. Identifier linkage remains yours: EDGAR keys on `CIK`, CRSP on `PERMNO`, Compustat on `GVKEY`.

**Free or licensed.** **Free**, no registration. Programmatic access via `data.sec.gov` (submissions and XBRL company-facts APIs), the daily and quarterly index files, and full-text search over filings from 2001, subject to the ten-requests-per-second fair-access policy and a descriptive user-agent header.

**Where this book uses it.** The baseline exercises of Chapters 12, 16, 17, 20, 24, and 25; Box 20.1; the 13F entry above.

### Jay Ritter's IPO data

**Free**, at Ritter's University of Florida page: first-day returns by year from 1980, long-run performance by cohort, founding dates, and IPO and listing counts — the standard reference for Chapter 12. The screens matter and are documented there: the tables exclude offerings below $5, unit offers, closed-end funds, REITs, ADRs, and — critically for any statement about the 2020s — special purpose acquisition companies. A count of "IPOs" including SPACs and one excluding them differ by a factor that swamps every economic effect under discussion, which is why Chapter 12's Table 12.1, regenerated in the pass §B.7 describes, now prints the full screen list in its source line rather than the phrase "Ritter IPO data."

### SIFMA issuance statistics

**Free** summary spreadsheets of issuance and outstandings by market — Treasury, agency MBS, corporate, municipal, asset-backed. Convenient, and defined differently from the Z.1: coverage of private placements and non-agency structures diverges, so a SIFMA outstanding and a Z.1 level for nominally the same instrument will not match, and mixing them inside one table is an error. Used in Chapters 12, 13, and 25.

### Treasury Monthly Statement of the Public Debt

**Free**, with the modern access path at `fiscaldata.treasury.gov`. The MSPD reports every outstanding Treasury security by type and maturity date, monthly — the source for the maturity distribution Chapter 9 uses in its gap-filling discussion, and the denominator for the Treasury row of Chapter 2's Table 2.5 when a reader wants par outstanding rather than the Z.1's market-value measure. The two differ by hundreds of billions in any period when rates have moved.

### World Bank GFDD and BIS statistics

Both **free**, both used in Chapter 27's cross-country comparisons. The World Bank's Global Financial Development Database assembles indicators of financial depth, access, efficiency, and stability for most countries annually from 1960, including the listed-domestic-companies counts Chapter 12's Table 12.2 cross-checks against; its indicators come from national sources of varying quality, so cross-country levels are orders of magnitude. The BIS credit statistics — total credit to the non-financial sector by borrowing sector and lender, spliced quarterly over long spans, plus the debt securities and locational banking statistics — are the better source wherever the two overlap, because the BIS does the definitional harmonization the GFDD does not.

***

## B.7 The `TODO(data)` resolution workflow

Thirteen `TODO(data)` markers were open in the manuscript as drafted. They mark tables whose magnitudes are right and whose figures are approximate — a deliberate choice, since a table stated to the nearest trillion teaches the map while one stated to three decimals invites a student to believe a precision the next revision will erase. Each is discharged the same way, and the discharge is a pipeline, not a proofreading pass.

**Eight are now discharged and five remain open.** The eight closed on 25 August 2026: Chapter 9's Table 9.3, Chapter 12's Tables 12.1 and 12.2, Chapter 13's Table 13.3, and all four of Chapter 14's. Of the five still open, **four are Chapter 2's**, held back deliberately rather than blocked — the chapter's dollar levels turn out to be a stale vintage, and they belong in one reconciliation pass over Table 2.5, Table 2.6, §2.5 and Box 2.2 rather than in four marker-by-marker patches. **The fifth is Chapter 24's covenant-lite share**, and it is blocked on a licensed source: the number requires one named PitchBook LCD or LSTA vintage, and the free substitutes named in §B.5 were checked against this need and cannot supply it — the Shared National Credit review, the Senior Loan Officer Opinion Survey and the *Financial Stability Report* discuss covenant-lite lending in prose and charts, but none publishes the share as a series a script can read. That is a clean illustration of what the licensed tier actually buys: not better data about the same object, but the only machine-readable statement of it.

The steps, as the discharged eight were actually run:

1. **Each tagged table gets a table-generating script**, at `code/tables/tab_NN_MM.py` — one file per table, named for the table it produces, so `tab_09_03.py` builds Table 9.3 — drawing on the Z.1, DFA, and SCF adapters in `fe_figures/sources/`. A table and a figure sharing a source share the adapter, which is why a chapter's figure and its table cannot silently disagree.
2. **The script writes the table as a markdown fragment** to `data/tables/tab_09_03.md`, source line included, and any figure it also produces to `figures/`. The fragment is the artifact the chapter carries; nothing is transcribed by hand.
3. **The chapter includes the fragment and records the vintage** in its *Source:* line: "Financial Accounts of the United States (Z.1), table L.210, holder rows through their FRED mirrors; data through 2026:Q1, retrieved 25 August 2026" replaces "magnitudes approximate." The cross-checks the marker called for are stated in the same line rather than kept in a working file, so a reader can see what was checked and against what.
4. **The marker is replaced by a provenance comment** naming the script, the series, the vintage and the cross-checks: `<!-- data: generated by code/tables/tab_09_03.py, Z.1 L.210 via the FRED mirror, vintage 2026:Q1 retrieved 2026-08-25; cross-checked against TIC official/non-official, H.4.1 SOMA, and the Z.1 hedge fund memo -->`. The comment travels with the table, so the next editor who wonders where a cell came from does not have to find this appendix first.
5. **`GENERATION_LOG.md` carries the vintage**, one row per successful run, so a re-run at the next Z.1 release is one command and the diff is reviewable line by line.
6. **Chapter 16's Table 16.4 overlap problem resolves inside the workflow, not in the prose.** Its rows sum to roughly $132 trillion against a stated total near $120 trillion because the categories overlap. The script computes the overlap explicitly and emits either a non-overlapping restatement or an explicit overlap row.
7. **Chapter 16's defined-benefit scope question resolves the same way.** Form 5500 is private-sector only (§B.3), so a DB/DC series built from it omits several trillion dollars of state and local plans; the script pulls those from Z.1 L.120 and the table states its scope in the *Source:* line.

The rule the workflow enforces is one sentence long: **no number appears in this book that cannot name the file and the release date it came from.** That is what "verify against the current release" means concretely, and it is why this appendix ends in a list rather than a peroration.

**Table B.7: The `TODO(data)` markers and what closes them**

| Chapter table                           | Primary source                                         | Cross-check                                                                    | Appendix section | Status                 |
| --------------------------------------- | ------------------------------------------------------ | ------------------------------------------------------------------------------ | ---------------- | ---------------------- |
| 2.3 Sector balance sheets               | Z.1 B.101, B.103, L.106, L.109, L.111                  | Sector totals against L.108                                                    | B.1              | Open — vintage pass    |
| 2.4 Fund, insurance, pension complex    | Z.1 L.116-L.124                                        | ICI Fact Book; Form 5500                                                       | B.1, B.3         | Open — vintage pass    |
| 2.5 Master holdings table               | Z.1 L.204-L.205, L.210-L.213, L.223-L.224, L.117-L.120 | Holdings sum to outstanding; discrepancy line stated                           | B.1              | Open — vintage pass    |
| 2.6 Three reallocations                 | Z.1 L.223, L.117-L.120, L.108, L.111                   | Form 5500 and ICI for the DB/DC split                                          | B.1, B.3         | Open — vintage pass    |
| 9.3 Treasury holders over time          | Z.1 L.210                                              | TIC official against private; H.4.1 SOMA; OFR and CFTC for the residual row    | B.1, B.3, B.4    | Discharged 2026-08-25  |
| 12.1 IPO first-day returns              | Ritter IPO tables                                      | Screens (offer price, exchange, units, ADRs, SPACs) printed in the source line | B.6              | Discharged 2026-08-25  |
| 12.2 Listed domestic companies          | Doidge-Karolyi-Stulz counts (CRSP-based)               | World Bank listed-companies series; the two differ by several hundred firms    | B.2, B.6         | Discharged 2026-08-25  |
| 13.3 Agency and MBS holders             | Z.1 L.211                                              | H.4.1 SOMA MBS; Call Report securities schedules (HTM against AFS); TIC        | B.1, B.4, B.5    | Discharged 2026-08-25  |
| 14.1 Equity participation by income     | SCF 2022 summary tables                                | Update to the next wave when released                                          | B.1              | Discharged 2026-08-25  |
| 14.2 Asset composition by wealth group  | SCF 2022 plus DFA                                      | Column shares sum to 100 after rounding                                        | B.1              | Discharged 2026-08-25  |
| 14.3 Household debt and leverage        | Z.1 B.101                                              | NY Fed Household Debt and Credit Report; debt-service ratio via FRED           | B.1              | Discharged 2026-08-25  |
| 14.4 Asset class shares by wealth group | DFA, latest quarter                                    | Note that the DFA distributes Z.1 aggregates using SCF shares                  | B.1              | Discharged 2026-08-25  |
| 24.2 Covenant-lite share                | Leveraged loan compiler, named vintage                 | Definitions vary by compiler; state which                                      | B.5              | Open — licensed source |

*Source: Author's compilation from the drafted manuscript and REVIEW\_FLAGS.md. "Vintage pass" means held deliberately for the author's single reconciliation pass over Chapter 2, not blocked. The discharged rows were run as `code/tables/tab_NN_MM.py` and their fragments sit in `data/tables/`; the cross-check column records what the source lines now state.*

***

## Where this appendix is used

Chapter 2 defers to this appendix for "data sources, free and licensed, including the Z.1 file structure," and §B.1 discharges that deferral: the table families, the sector-and-instrument numbering, the series-code grammar, the Data Download Program, the renumbering warning, the discrepancy line, and the household-as-residual construction that Chapter 2 states twice and Box 2.4 dramatizes. Section B.0's roster carries the free-or-licensed flag for every source in the book, which is the other half of what Chapter 2 promises.

Beyond that: Chapter 5 for consumption vintages and timing; Chapters 6 and 7 and Appendix A for the French library and the Goyal-Welch predictors; Chapters 9, 10, and 13 for rates, spreads, and holder decompositions; Chapters 14 through 20 for the holdings sources that make the investor ecology measurable at all; Chapters 22 through 25 for filings and fundamentals; and Chapter 26 for the supervisory data that make risk constraints visible from outside a firm.
