> For the complete documentation index, see [llms.txt](https://laurence-wilse-samson.gitbook.io/textbooks/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://laurence-wilse-samson.gitbook.io/textbooks/financial-economics-claims-prices-holders/appendices/appendix_d_math_review.md).

# Appendix D: Mathematical Review

This appendix carries the mathematics the chapters use and decline to derive. It is not a course in optimization, linear algebra, probability, or stochastic calculus; it is the specific set of results the body leans on, worked in the body's own notation, so that the algebra behind a stated result sits in one place rather than in four textbooks with four conventions.

The prerequisite is the one Chapter 1 states: calculus through partial derivatives, and one statistics course. Nothing here requires measure theory. Two results genuinely do — the fundamental theorem of asset pricing in infinite state spaces, and the construction of the Itô integral — and in both cases §§D.5 and D.7 state the finite-dimensional version completely, say where the general version stops being elementary, and name the book that proves it. Each section names the chapter that consumes it; the sections are independent except that D.7 uses D.2 and D.5 uses D.4.

***

## D.1 Notation and Conventions

*Consumed by: everywhere.*

`NOTATION.md` is the book's registry of symbols and it governs this appendix as it governs every chapter. Six conventions do most of the work, and a reader who arrives here first should have them.

**Gross against net.** Capital $$R$$ is a gross return (1.07), lower-case $$r$$ is net ($$0.07$$), and $$R = 1 + r$$. The pricing equation is always written gross, $$1 = E\[mR]$$; present-value formulas, mean-variance algebra, and empirical regressions are written net, so a risk premium is $$\mu\_i - r\_f$$. Where a chapter announces it — §5.2 and §7.4 do — $$r$$ additionally denotes the log return $$\ln R$$, which agrees with the net return to first order. Section D.2 is about how far "to first order" goes.

$$m$$**, never** $$M$$**.** The stochastic discount factor is lower-case throughout; $$M$$ is reserved for the market-portfolio subscript, as in $$R\_M$$.

**The star.** On probabilities and measures, $$^{\ast}$$ marks the risk-neutral object ($$\pi^{\ast}$$, $$E^{\ast}$$); on choice variables it marks an optimum ($$w^{\ast}$$, $$n^{\ast}$$, $$\Psi^{\ast}$$); in headings, ★ marks a PhD-track section. Never star a probability to mean an optimum.

**Preference parameters.** $$\beta$$ is always a regression coefficient. Time preference is $$\delta$$, relative risk aversion is $$\gamma$$, absolute risk aversion is $$\tau$$, the elasticity of intertemporal substitution is $$\psi$$.

**Indices, dates, and arrays.** Asset $$i$$ or $$j$$; state $$s$$; time $$t$$; factor $$k$$. Dates are 0 and 1 in two-date settings, $$t$$ and $$t+1$$ in multi-period ones. Vectors are columns, $$'$$ denotes transpose, and $$\mathbf{1}$$ is a conformable vector of ones.

**Functions take arguments.** A symbol followed by parentheses is a function, and that is what separates several look-alikes: $$N(\cdot)$$ is the standard normal cumulative distribution function while $$N$$ is the number of risky assets; $$\phi(\cdot)$$ is the standard normal density while every bare $$\phi$$ in the registry is a fraction; $$u(\cdot)$$ is a period utility function while $$u$$ is a binomial up-multiplier.

Because the chapters never wrote a Lagrangian or a stochastic differential equation, this appendix introduces a handful of symbols the registry does not yet carry. They are listed here so they can be checked rather than discovered in an equation.

**Table D.1: Symbols this appendix adds**

| Symbol                                                     | Meaning                                                                                       | Section  | Separated from                                                                                                                                                                                                   |
| ---------------------------------------------------------- | --------------------------------------------------------------------------------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| $$\mathcal{L}$$                                            | Lagrangian                                                                                    | D.3      | The loss $$\mathcal{L}$$ of Ch 10 §10.8 and Ch 26 §26.3; no loss distribution appears here                                                                                                                       |
| $$\lambda\_1$$, $$\lambda\_2$$, $$\lambda\_j$$             | Lagrange multipliers                                                                          | D.3      | The price of risk, which always carries a model subscript ($$\lambda\_m$$, $$\lambda\_M$$, $$\lambda\_k$$), and Ch 10 §10.3's default intensity, which appears in no optimization                                |
| $$A\_\Sigma$$, $$B\_\Sigma$$, $$C\_\Sigma$$, $$D\_\Sigma$$ | The four scalars of the mean-variance frontier                                                | D.3      | The unsubscripted $$A$$ (assets), $$B$$ (bond position), $$C$$ (call price), $$D$$ (CARA demand); the $$\Sigma$$ subscript marks them as functions of the covariance matrix                                      |
| $$X$$                                                      | The $$\lvert S\rvert \times N$$ matrix of traded payoffs; its column span is the payoff space | D.4, D.5 | Ch 5 §5.5's habit level and Ch 9 §9.2's state vector, both dated $$X\_t$$. Ch 3 §3.5 already writes $$\mathrm{proj}(m \mid X)$$ in this sense                                                                    |
| $$\theta$$                                                 | Portfolio of traded assets, in units of each asset                                            | D.3, D.5 | Ch 4 §4.3's scalar risky share; the two never co-occur                                                                                                                                                           |
| $$\varphi$$                                                | Normal vector of the separating hyperplane                                                    | D.5      | Every $$\phi$$ in the registry: those are fractions, or the density $$\phi(\cdot)$$, which takes an argument                                                                                                     |
| $$\mathrm{CE}$$                                            | Certainty equivalent                                                                          | D.6      | Roman, a derived quantity, in the manner of $$\mathrm{SR}$$ and $$\mathrm{PME}$$                                                                                                                                 |
| $$z\_t$$, $$dz$$                                           | Standard Brownian motion and its increment                                                    | D.7      | **Deliberate reuse.** Ch 7 §7.2's supply noise and Ch 11 §11.3's order flow are unsubscripted scalars in models with no continuous time; here $$z$$ always carries a time subscript or appears as a differential |
| $$\Pi\_t$$                                                 | Value of the hedged portfolio in the Black-Scholes derivation                                 | D.7      | The physical probability $$\pi\_s$$, which is lower case and carries a state subscript                                                                                                                           |

*Source: Author's construction, checked against `NOTATION.md`.*

***

## D.2 Logs, Lognormals, and the Jensen Correction

*Consumed by: Ch 5 §§5.2-5.4; Ch 7 §7.4's log-linear identity; Ch 8's Black-Scholes.*

Finance runs on logarithms for two reasons that have nothing to do with each other. Log returns add across time, which makes multi-period statements linear. And log returns are approximately normal when net returns are not, which makes distributional statements tractable. Both conveniences are bought at the price of a correction term, and the correction term is where the mistakes live.

### D.2.1 Log against net

Define the **log return** (or continuously compounded return) as $$\ln R = \ln(1+r)$$. The Taylor series around $$r = 0$$ gives

$$
\ln(1+r) = r - \tfrac{1}{2}r^2 + \tfrac{1}{3}r^3 - \cdots
$$

so the two agree to first order and the log return is always the smaller of the two for $$r > 0$$. The wedge is second order in $$r$$, which is why nobody worries about it at monthly frequency and everybody should worry about it at decadal frequency.

**Table D.2: How far "agrees to first order" goes**

| Net return $$r$$ | Log return $$\ln(1+r)$$ | Difference   |
| ---------------- | ----------------------- | ------------ |
| 0.01             | 0.00995                 | $$-0.00005$$ |
| 0.05             | 0.04879                 | $$-0.00121$$ |
| 0.10             | 0.09531                 | $$-0.00469$$ |
| 0.25             | 0.22314                 | $$-0.02686$$ |
| 0.50             | 0.40546                 | $$-0.09454$$ |
| $$-0.50$$        | $$-0.69315$$            | $$-0.19315$$ |

*Source: Author's calculation.*

The last row is the one that matters. The approximation is not symmetric: a 50 percent loss is a $$-69$$ percent log return, because a claim that has halved must double to recover. This is not a curiosity — it is the reason a levered fund's arithmetic average return can be positive while its terminal wealth is a fraction of what it started with, and it is why Chapter 18's performance measurement insists on the internal rate of return rather than an average.

The additivity property is exact and worth stating: over $$T$$ periods, $$\ln(R\_1 R\_2 \cdots R\_T) = \sum\_t \ln R\_t$$. Net returns do not add; log returns do.

### D.2.2 The lognormal moment formulas

A random variable is **lognormal** if its logarithm is normal. Let $$\ln y \sim N(\mu\_y, \sigma\_y^2)$$. Then, completing the square inside the defining integral,

$$
E\[y] = \exp\left(\mu\_y + \tfrac{1}{2}\sigma\_y^2\right), \qquad \mathrm{Var}(y) = \exp\left(2\mu\_y + \sigma\_y^2\right)\left(e^{\sigma\_y^2} - 1\right)
$$

The first of these, taken in logs, is the workhorse:

$$
\ln E\[y] = E\[\ln y] + \tfrac{1}{2}\mathrm{Var}(\ln y)
$$

The $$\tfrac{1}{2}\sigma^2$$ term is the **Jensen correction**. It is not a modeling choice; it is the exact size of the gap that Jensen's inequality (§D.8) says must be there whenever a convex function meets an expectation. The expected value of a lognormal claim exceeds the exponential of its expected log by a factor that grows with its variance, and a great deal of finance consists of remembering which of the two objects a formula is talking about.

Dividing the second formula by the square of the first gives the coefficient of variation, which is the form §5.4 uses:

$$
\frac{\sigma(y)}{E\[y]} = \sqrt{e^{\sigma\_y^2} - 1} \approx \sigma\_y \quad \text{for small } \sigma\_y
$$

![Figure D.2: Logs and lognormals](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-614f1eaeee4d992c18f7671d686ddb45505d3a92%2Ffig_D_02_lognormal.png?alt=media)

**Figure D.2: Logs and lognormals.** Panel (a) is one random variable seen twice: the log return, normal and symmetric about zero, and the gross return it exponentiates to, lognormal and skewed right with a floor at zero. The median of the gross return is one — the exponential of the log's mean — and its mean is strictly above that, at the exponential of the mean plus half the variance. The gap between the two vertical rules is the entire content of Jensen's inequality for this distribution. Panel (b) plots that gap against volatility, and it is exactly half sigma squared. Three familiar facts are the same fact drawn here. An arithmetic average return exceeds a geometric one, by this amount. A volatile fund compounds to less than a steady one with the same average return, by this amount per period. And Chapter 5's log-linear derivation carries a half-sigma-squared term that is not an approximation error but the exact Jensen wedge — which is why dropping it changes the answer rather than rounding it. *Source: Author's construction from Section D.2.2.*

### D.2.3 Chapter 5's two-equation system, in full

§5.2 states that "the derivation is three lines" and then writes the three lines. Here they are with the steps between them.

Assume $$\ln m$$ and $$r = \ln R$$ are jointly normal and identically distributed over time. Under power utility $$m = \delta e^{-\gamma g}$$, where $$g$$ is log consumption growth with mean $$\mu\_g$$ and standard deviation $$\sigma\_g$$, so $$\ln m = \ln\delta - \gamma g$$ is normal whenever $$g$$ is.

Start from the pricing equation $$1 = E\[mR] = E\[e^{\ln m + r}]$$. The exponent is normal, so apply the moment formula and take logs of both sides:

$$
0 = E\[\ln m + r] + \tfrac{1}{2}\mathrm{Var}(\ln m + r)
$$

Expand the variance of a sum:

$$
0 = E\[\ln m] + E\[r] + \tfrac{1}{2}\mathrm{Var}(\ln m) + \tfrac{1}{2}\mathrm{Var}(r) + \mathrm{Cov}(\ln m, r)
$$

which is Chapter 5's first display. Now apply it twice.

**The riskless asset.** Here $$r = r\_f$$ is a constant, so $$\mathrm{Var}(r) = 0$$ and $$\mathrm{Cov}(\ln m, r) = 0$$. Substituting $$E\[\ln m] = \ln\delta - \gamma\mu\_g$$ and $$\mathrm{Var}(\ln m) = \gamma^2\sigma\_g^2$$,

$$
0 = \ln\delta - \gamma\mu\_g + r\_f + \tfrac{1}{2}\gamma^2\sigma\_g^2 \quad\Longrightarrow\quad r\_f = -\ln\delta + \gamma\mu\_g - \tfrac{1}{2}\gamma^2\sigma\_g^2
$$

**A risky asset.** Subtract the riskless equation from the general one. The $$E\[\ln m]$$ and $$\tfrac{1}{2}\mathrm{Var}(\ln m)$$ terms are common to both and cancel, leaving

$$
0 = \big(E\[r] - r\_f\big) + \tfrac{1}{2}\mathrm{Var}(r) + \mathrm{Cov}(\ln m, r)
$$

Since $$\ln m = \ln\delta - \gamma g$$ and $$\ln\delta$$ is a constant, $$\mathrm{Cov}(\ln m, r) = -\gamma\mathrm{Cov}(g, r)$$, and therefore

$$
E\[r] - r\_f + \tfrac{1}{2}\mathrm{Var}(r) = \gamma\mathrm{Cov}(g, r)
$$

The left-hand side is the log premium plus its own Jensen correction, which together approximate the arithmetic premium $$E\[R] - R\_f$$. That is the premium equation.

**The Hansen-Jagannathan ratio.** §5.4 reports that $$\sigma(m)/E\[m] = \sqrt{e^{\gamma^2\sigma\_g^2} - 1}$$ *exactly*, not approximately. Apply §D.2.2's coefficient of variation to $$m = \delta e^{-\gamma g}$$: here $$\ln m$$ is normal with variance $$\gamma^2\sigma\_g^2$$, so

$$
\frac{\sigma(m)}{E\[m]} = \sqrt{e^{\gamma^2\sigma\_g^2} - 1} \approx \gamma\sigma\_g
$$

The subjective discount factor $$\delta$$ appears in neither expression, because it scales $$m$$ and the ratio is scale-free. That is worth noticing: the Hansen-Jagannathan bound is a statement about the *volatility* of marginal utility growth, and impatience does not supply any.

***

## D.3 Optimization and Lagrangians

*Consumed by: Ch 4 §§4.2-4.3 (the frontier algebra the body states without deriving); Ch 3 §3.7; Ch 16 §16.6.*

### D.3.1 Equality constraints and the multiplier as a shadow price

Consider maximizing $$f(x)$$ over $$x \in \mathbb{R}^N$$ subject to $$h(x) = b$$. Form the **Lagrangian**

$$
\mathcal{L}(x, \lambda) = f(x) - \lambda\big(h(x) - b\big)
$$

and set its partial derivatives to zero. The first-order conditions are $$\nabla f(x) = \lambda \nabla h(x)$$ together with $$h(x) = b$$. Geometrically: at an optimum the objective's gradient is parallel to the constraint's, because any direction that raises $$f$$ while staying on the constraint surface would be an improvement, and there is none.

The economics is in the multiplier. Let $$f^{\ast}(b)$$ denote the maximized value as a function of the constraint level. The **envelope theorem** says

$$
\frac{df^{\ast}(b)}{db} = \lambda
$$

**The multiplier is the shadow price of the constraint** — the rate at which the objective improves if the constraint is relaxed by one unit. This is the interpretation §3.7 and §4.9 both invoke when they say that a binding constraint's shadow price enters the pricing relation, and it is worth seeing exactly how it gets there.

Take a household choosing holdings $$\theta\_j$$ of asset $$j$$ to maximize $$u(c\_0) + \delta E\[u(c\_1)]$$, with $$c\_0 = e\_0 - \sum\_j p\_j\theta\_j$$ and $$c\_1 = e\_1 + \sum\_j \theta\_j x\_j$$ for endowments $$e\_0, e\_1$$. With no constraints, differentiating with respect to $$\theta\_j$$ gives $$-p\_j u'(c\_0) + \delta E\[u'(c\_1)x\_j] = 0$$, which is $$p\_j = E\[mx\_j]$$ with $$m = \delta u'(c\_1)/u'(c\_0)$$ — §3.5's derivation.

Now impose a short-sale or borrowing limit, $$\theta\_j \ge -\bar\theta\_j$$. The first-order condition acquires a multiplier $$\lambda\_j \ge 0$$ on that inequality:

$$
-p\_j u'(c\_0) + \delta E\[u'(c\_1)x\_j] + \lambda\_j = 0 \quad\Longrightarrow\quad p\_j = E\[m x\_j] + \frac{\lambda\_j}{u'(c\_0)}
$$

When the constraint is slack, $$\lambda\_j = 0$$ and the Euler equation holds. When it binds, $$\lambda\_j > 0$$ and the observed price *exceeds* what this investor's marginal utility would justify. That is Palm in March 2000, and it is the CDS-bond basis in 2008, and it is the covered-interest-parity deviation in every year since: a price set by the marginal buyer because the marginal seller is locked out. The Euler equation has not failed. It has acquired a term, and the term is a fact about somebody's balance sheet rather than about their preferences.

### D.3.2 Inequality constraints and corners

The general statement is the **Kuhn-Tucker** (or Karush-Kuhn-Tucker) conditions. To maximize $$f(x)$$ subject to $$h\_j(x) \le b\_j$$ for each constraint $$j$$, the necessary conditions at an optimum — under a regularity condition on the constraint gradients — are

$$
\nabla f(x) = \sum\_j \lambda\_j \nabla h\_j(x), \qquad \lambda\_j \ge 0, \qquad \lambda\_j\big(h\_j(x) - b\_j\big) = 0
$$

The third line is **complementary slackness**: either the constraint binds or its multiplier is zero, never both non-zero. It is the formal content of the sentence "the investor sits at a corner." A mean-variance investor forbidden to short holds $$w\_i = 0$$ in an asset whose unconstrained optimal weight was negative, and the multiplier on that constraint measures how much the prohibition costs — which is also, by §D.3.1, how much the asset would have to fall in price before the investor would want to hold a positive amount.

Two cautions. The conditions are necessary, and sufficient only when $$f$$ is concave and the feasible set convex — which covers every optimization in this book. And they are silent about *which* constraints bind, so finding the active set is a computational problem, which is why constrained portfolio optimization is a quadratic program rather than a formula. Simon and Blume (1994) give the regularity conditions; Boyd and Vandenberghe (2004) give the machinery.

### D.3.3 The mean-variance frontier, derived

§4.2 states that the frontier is a hyperbola and does not derive it. Here is the derivation.

Let $$\mu$$ be the $$N$$-vector of expected net returns and $$\Sigma$$ the covariance matrix, assumed positive definite (see §D.4). The frontier solves

$$
\min\_w \tfrac{1}{2}w'\Sigma w \quad \text{subject to} \quad w'\mu = \mu\_p, \quad w'\mathbf{1} = 1
$$

The factor $$\tfrac{1}{2}$$ is cosmetic and clears a two from the derivative. The Lagrangian is

$$
\mathcal{L} = \tfrac{1}{2}w'\Sigma w - \lambda\_1\big(w'\mu - \mu\_p\big) - \lambda\_2\big(w'\mathbf{1} - 1\big)
$$

Differentiating with respect to the vector $$w$$ and setting to zero,

$$
\Sigma w = \lambda\_1\mu + \lambda\_2\mathbf{1} \quad\Longrightarrow\quad w = \Sigma^{-1}\big(\lambda\_1\mu + \lambda\_2\mathbf{1}\big)
$$

Every frontier portfolio is therefore a combination of two fixed vectors, $$\Sigma^{-1}\mu$$ and $$\Sigma^{-1}\mathbf{1}$$, with the weights on them determined by the target return. **That is two-fund separation**, and it has just fallen out of the first-order condition without further argument: the set of frontier portfolios is a two-dimensional family, so any two of its members span it.

To pin down $$\lambda\_1$$ and $$\lambda\_2$$, define the four scalars

$$
A\_\Sigma = \mathbf{1}'\Sigma^{-1}\mathbf{1}, \qquad B\_\Sigma = \mathbf{1}'\Sigma^{-1}\mu, \qquad C\_\Sigma = \mu'\Sigma^{-1}\mu, \qquad D\_\Sigma = A\_\Sigma C\_\Sigma - B\_\Sigma^2
$$

with $$A\_\Sigma > 0$$ and $$C\_\Sigma > 0$$ because $$\Sigma^{-1}$$ is positive definite, and $$D\_\Sigma > 0$$ by the Cauchy-Schwarz inequality applied in the inner product $$\langle a, b\rangle = a'\Sigma^{-1}b$$, with equality only if $$\mu$$ is proportional to $$\mathbf{1}$$ — that is, only if every asset has the same expected return and there is no frontier to speak of. Substituting $$w$$ into the two constraints gives

$$
\mu\_p = \lambda\_1 C\_\Sigma + \lambda\_2 B\_\Sigma, \qquad 1 = \lambda\_1 B\_\Sigma + \lambda\_2 A\_\Sigma
$$

a linear system whose solution is

$$
\lambda\_1 = \frac{A\_\Sigma\mu\_p - B\_\Sigma}{D\_\Sigma}, \qquad \lambda\_2 = \frac{C\_\Sigma - B\_\Sigma\mu\_p}{D\_\Sigma}
$$

Now compute the minimized variance. Because $$\Sigma w = \lambda\_1\mu + \lambda\_2\mathbf{1}$$, we have $$w'\Sigma w = \lambda\_1 w'\mu + \lambda\_2 w'\mathbf{1} = \lambda\_1\mu\_p + \lambda\_2$$, and substituting,

$$
\sigma\_p^2 = \frac{A\_\Sigma\mu\_p^2 - 2B\_\Sigma\mu\_p + C\_\Sigma}{D\_\Sigma}
$$

This is the **frontier hyperbola**. It is a quadratic in $$\mu\_p$$ with positive leading coefficient, so in $$(\sigma\_p^2, \mu\_p)$$ space it is a parabola opening rightward, and in $$(\sigma\_p, \mu\_p)$$ space — the picture Chapter 4 draws — a hyperbola. Its leftmost point, the **global minimum-variance portfolio**, is found by setting the derivative to zero:

$$
\mu\_p = \frac{B\_\Sigma}{A\_\Sigma}, \qquad \sigma\_p^2 = \frac{1}{A\_\Sigma}, \qquad w = \frac{\Sigma^{-1}\mathbf{1}}{A\_\Sigma}
$$

and the last of these is the only frontier portfolio whose weights do not depend on $$\mu$$ at all — the reason minimum-variance investing survives the estimation problems of §D.4.

One more line, because it is the cleanest demonstration of §D.3.1 in the book. Differentiate the frontier with respect to the target return:

$$
\frac{d}{d\mu\_p}\left(\tfrac{1}{2}\sigma\_p^2\right) = \frac{A\_\Sigma\mu\_p - B\_\Sigma}{D\_\Sigma} = \lambda\_1
$$

The multiplier on the expected-return constraint *is* the marginal variance cost of an extra unit of expected return. The shadow price is not a metaphor here; it is the slope of the frontier.

### D.3.4 The tangency portfolio

Add a riskless asset at net return $$r\_f$$. §4.3 states that the tangency weights are $$w^T \propto \Sigma^{-1}(\mu - r\_f\mathbf{1})$$. Derive it by maximizing the Sharpe ratio directly:

$$
\max\_w \mathrm{SR}(w) = \frac{w'(\mu - r\_f\mathbf{1})}{\left(w'\Sigma w\right)^{1/2}}
$$

Note first that $$\mathrm{SR}(cw) = \mathrm{SR}(w)$$ for any $$c > 0$$: the Sharpe ratio is homogeneous of degree zero, so the normalization $$w'\mathbf{1} = 1$$ can be imposed after the fact rather than as a constraint. Differentiating,

$$
\frac{\mu - r\_f\mathbf{1}}{\left(w'\Sigma w\right)^{1/2}} - \frac{\big(w'(\mu - r\_f\mathbf{1})\big)\Sigma w}{\left(w'\Sigma w\right)^{3/2}} = 0
$$

Multiply through by $$(w'\Sigma w)^{1/2}$$ and rearrange:

$$
\mu - r\_f\mathbf{1} = \frac{w'(\mu - r\_f\mathbf{1})}{w'\Sigma w}\Sigma w
$$

The scalar in front is positive at any portfolio with a positive premium, so

$$
w^T \propto \Sigma^{-1}\big(\mu - r\_f\mathbf{1}\big)
$$

normalized to sum to one. Reward divided by risk, in the matrix sense, exactly as Chapter 4 reads it.

The same vector arrives by a different road. An investor maximizing $$w'(\mu - r\_f\mathbf{1}) - \tfrac{\gamma}{2}w'\Sigma w$$ — expected excess return penalized by variance, with $$\gamma$$ the coefficient of risk aversion — has first-order condition $$\mu - r\_f\mathbf{1} = \gamma\Sigma w$$, so $$w = \gamma^{-1}\Sigma^{-1}(\mu - r\_f\mathbf{1})$$. Same direction, and now with the scale pinned: risk aversion determines *how much* of the tangency portfolio to hold and nothing about its composition. That is Chapter 4's two-fund separation in its sharp form, and §D.6.2 will show that the CARA-normal demand function of §7.2 is this same expression with one asset and absolute rather than relative risk aversion.

§16.6's divestment arithmetic is the comparative static of this expression under a constraint. If holders of a fraction $$\phi$$ of wealth are forbidden asset $$i$$, the remaining holders must absorb weight $$w\_i/(1-\phi)$$ rather than $$w\_i$$, and the first-order condition $$\mu - r\_f\mathbf{1} = \gamma\Sigma w$$ prices the extra exposure at $$\Delta E\[r\_i] \approx \gamma\big\[\phi/(1-\phi)\big]w\_i\sigma\_i^2$$, with $$\sigma\_i^2$$ the undiversifiable part. The constraint is on holders; the effect is in the price; the multiplier is the mechanism.

***

## D.4 Linear Algebra for Portfolio Problems

*Consumed by: Ch 3 §3.5; Ch 4; Ch 6 §6.1.*

### D.4.1 Quadratic forms and positive definiteness

A symmetric $$N \times N$$ matrix $$\Sigma$$ is **positive semi-definite** if $$w'\Sigma w \ge 0$$ for every $$w$$, and **positive definite** if the inequality is strict for every $$w \ne 0$$. Any covariance matrix is automatically positive semi-definite, because $$w'\Sigma w = \mathrm{Var}(w'r) \ge 0$$ — a variance cannot be negative, and that is the whole proof.

Positive definiteness is the stronger and more consequential property. It fails exactly when some portfolio $$w \ne 0$$ has $$\mathrm{Var}(w'r) = 0$$ — a combination of the assets that is riskless. Equivalently, $$\Sigma$$ is **singular**, $$\Sigma^{-1}$$ does not exist, and every formula in §D.3 is undefined. Two ways this happens. Either an asset in the list is a portfolio of the others (an exchange-traded fund alongside its own basket, two share classes of one company, a futures contract and its cash-and-carry equivalent), in which case deleting the redundant asset loses nothing, since no payoff has left the opportunity set. Or there are too few observations: a sample covariance matrix built from $$T$$ periods on $$N$$ assets has rank at most $$\min(N, T-1)$$, so with 500 stocks and 60 months of data it is singular by construction and the "optimal" portfolio is a numerical artifact.

The dangerous case is neither of these but **near**-singularity. When the smallest eigenvalue of $$\Sigma$$ is small but positive, $$\Sigma^{-1}$$ has a very large one, and the frontier weights amplify estimation error in exactly the directions where the data say least. This is Michaud's (1989) "error maximization" — mean-variance optimization is an efficient estimator of the largest expected-return errors, not of the frontier. The standard repairs are shrinkage of the covariance estimate (Ledoit and Wolf 2004), shrinkage of the expected returns (Black and Litterman 1992), or no-short-sale constraints, which Jagannathan and Ma (2003) show are *equivalent* to a particular shrinkage of $$\Sigma$$. The constraint that looks like a limitation is doing statistical work.

### D.4.2 Projection and the discount factor in the payoff space

The idea that makes §3.5's $$m^{\ast}$$ intelligible is that random variables form an inner-product space. Define the inner product of two payoffs as

$$
\langle x, y \rangle = E\[xy]
$$

This satisfies everything an inner product must, and it induces the norm $$\lVert x\rVert = \sqrt{E\[x^2]}$$ and the notion of orthogonality: $$x \perp y$$ means $$E\[xy] = 0$$. Two payoffs are orthogonal when they are uncorrelated *and* at least one has mean zero — a distinction worth keeping, because "orthogonal" and "uncorrelated" are not synonyms here.

Let $$X$$ be the matrix of traded payoffs and let its column span — the set of payoffs attainable as portfolios of traded assets — be the **payoff space**. In a finite-dimensional setting this is a linear subspace, and the **projection theorem** applies: for any random variable $$m$$, there is a unique element $$m^{\ast}$$ of the payoff space such that $$m - m^{\ast}$$ is orthogonal to every element of it. That element is the projection, written $$m^{\ast} = \mathrm{proj}(m \mid X)$$, and it is characterized by the **normal equations** $$E\[(m - m^{\ast})x] = 0$$ for every traded $$x$$.

Two consequences, both of which the body uses.

$$m^{\ast}$$ **prices everything** $$m$$ **prices.** If $$p(x) = E\[mx]$$ for every traded $$x$$, then $$E\[m^{\ast}x] = E\[mx] = p(x)$$ as well, by the normal equations. In the finite case there is a formula: with $$x$$ the $$N$$-vector of traded payoffs and $$p$$ the corresponding price vector,

$$
m^{\ast} = p'E\[xx']^{-1}x
$$

which one verifies by computing $$E\[m^{\ast}x] = E\[xx']E\[xx']^{-1}p = p$$. The discount factor that lives in the payoff space is itself a portfolio, and this is the sense in which Chapter 3 says it is "the unique discount factor that can be written as a portfolio of traded assets."

$$m^{\ast}$$ **is the least volatile discount factor.** Suppose the riskless payoff is traded, so the constant $$1$$ lies in the payoff space. Then $$E\[m - m^{\ast}] = 0$$, hence $$E\[m] = E\[m^{\ast}]$$, and because $$m - m^{\ast}$$ is orthogonal to $$m^{\ast}$$,

$$
\mathrm{Var}(m) = \mathrm{Var}(m^{\ast}) + \mathrm{Var}(m - m^{\ast}) \ge \mathrm{Var}(m^{\ast})
$$

So among all discount factors consistent with observed prices, $$m^{\ast}$$ has the smallest variance and therefore the smallest $$\sigma(m)/E\[m]$$. That is the Hansen-Jagannathan bound of §5.4 read geometrically: the bound is the length of a projection, and the reason it holds with equality in Chapter 3's complete two-state economy is that there the payoff space is everything, so $$m = m^{\ast}$$ and there is no residual to discard.

One honesty note that §D.5 will need: $$m^{\ast}$$ is not guaranteed to be positive. Projection is a linear operation and it does not respect the positive orthant. The discount factor whose existence the fundamental theorem guarantees is strictly positive; the one that lives in the payoff space is the one empirical work can construct. They coincide only when markets are complete.

***

## D.5 State Prices, Separating Hyperplanes, and the Fundamental Theorem

*Consumed by: Ch 3 §§3.4-3.5 (the body defers exactly this, at §3.5's "Complete and incomplete markets"); Ch 8's risk-neutral valuation.*

§3.4 extracted state prices from two equations in two unknowns and observed that the system was square because the market was complete. §3.5 then asserted the general result and sent the reader here: **no arbitrage holds if and only if there exists a strictly positive** $$m$$**; markets are complete if and only if it is unique.** This section proves it in a finite state space, which is where all of the intuition and none of the technical difficulty lives.

### D.5.1 The payoff space, the cone, and arbitrage as geometry

Fix two dates and a finite state space $$S$$ with $$\lvert S\rvert$$ states, all of which have strictly positive physical probability. (States with zero probability can be deleted without changing any price, and doing so is what makes $$m\_s = q\_s/\pi\_s$$ well defined.) There are $$N$$ traded assets. Write $$X$$ for the $$\lvert S\rvert \times N$$ payoff matrix, whose $$(s,j)$$ entry is asset $$j$$'s payoff in state $$s$$, and $$p$$ for the $$N$$-vector of prices.

A **portfolio** is a vector $$\theta \in \mathbb{R}^N$$, with negative entries permitted — short sales are allowed, and this matters. It costs $$p'\theta$$ today and pays $$X\theta \in \mathbb{R}^{\lvert S\rvert}$$ tomorrow.

An **arbitrage** is a portfolio that is never worse and sometimes better than nothing. Formally, $$\theta$$ is an arbitrage if the vector

$$
\big(-p'\theta, X\theta\big) \in \mathbb{R}^{1+\lvert S\rvert}
$$

has every component non-negative and at least one component strictly positive. The first component is money received today; the remaining components are payoffs tomorrow. This single definition covers both cases §3.2 distinguishes: a free lunch today with no liability tomorrow (first component positive, rest zero) and a costless position with a chance of a gain (first component zero, some later component positive).

Now the geometry. The set

$$
\mathcal{K} = \left\lbrace\big(-p'\theta, X\theta\big) : \theta \in \mathbb{R}^N\right\rbrace
$$

is the image of $$\mathbb{R}^N$$ under a linear map, hence a **linear subspace** of $$\mathbb{R}^{1+\lvert S\rvert}$$ — it contains the origin, and it is closed under addition and under multiplication by any scalar, positive or negative. The last clause is where short selling enters: reversing a portfolio reverses its point in $$\mathcal{K}$$.

The non-negative orthant $$\mathbb{R}^{1+\lvert S\rvert}\_+$$ is a **convex cone**: closed under addition and under multiplication by non-negative scalars. And the definition above says exactly this:

$$
\textbf{No arbitrage} \iff \mathcal{K} \cap \mathbb{R}^{1+\lvert S\rvert}\_+ = \lbrace 0\rbrace
$$

A subspace and a cone that meet only at the origin. That is the entire content of no arbitrage as a geometric object, and everything that follows is a theorem about subspaces and cones that happens to be about finance.

![Figure D.1: Why a positive m exists](https://846781005-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F3EupdX99vVBoNySDtmxb%2Fuploads%2Fgit-blob-10107346d16a959d42abd8467833e888610a45c2%2Ffig_D_01_separating_hyperplane.png?alt=media)

**Figure D.1: Why a positive m exists.** The geometry of no arbitrage in the smallest case that shows it: two states, so that the payoff space is a plane. The marketed subspace is a line through the origin — a line, and not a ray, because short selling means every portfolio's reflection is also available. The non-negative orthant is the shaded quadrant, and a point in it other than the origin is an arbitrage: a payoff that is never negative and somewhere positive, at a price of zero. No arbitrage is therefore the statement that the line meets the quadrant only at the origin, and that is all it is. Everything else follows from a theorem with no finance in it: a subspace and a convex cone meeting only at the origin can be separated by a hyperplane, and the normal to that hyperplane can be chosen strictly positive in every coordinate. That normal is the vector of state prices, and a strictly positive state-price vector is a strictly positive stochastic discount factor. Chapter 3 §3.4's fundamental theorem is this picture in an arbitrary number of states. *Source: Author's construction from Section D.5.*

### D.5.2 The separating hyperplane theorem

> **Separating hyperplane theorem (finite-dimensional).** Let $$\mathcal{C}$$ and $$\mathcal{D}$$ be disjoint non-empty convex subsets of $$\mathbb{R}^n$$, with $$\mathcal{C}$$ compact and $$\mathcal{D}$$ closed. Then there exist a vector $$\varphi \ne 0$$ and a scalar $$a$$ such that $$\varphi'v < a < \varphi'c$$ for every $$v \in \mathcal{D}$$ and every $$c \in \mathcal{C}$$.

The proof in finite dimensions is elementary and worth a sketch, because the intuition transfers. Consider the set of differences $$\mathcal{C} - \mathcal{D}$$, which is convex and does not contain the origin, and which is closed because $$\mathcal{C}$$ is compact. A closed convex set that omits a point contains a unique point nearest to it — uniqueness because the midpoint of two equally near points would be nearer, by the parallelogram law. Let $$\varphi$$ be that nearest point. Then the hyperplane through $$\varphi/2$$ perpendicular to $$\varphi$$ separates: every point of the set lies on the far side, since a point on the near side would admit a shorter connection. Rockafellar (1970) gives the general statement, including the infinite-dimensional versions that Chapter 3's Duffie reference needs.

The version we want specializes this to a subspace and a cone.

> **Corollary.** Let $$\mathcal{K} \subset \mathbb{R}^n$$ be a linear subspace with $$\mathcal{K} \cap \mathbb{R}^n\_+ = \lbrace 0\rbrace$$. Then there exists $$\varphi \in \mathbb{R}^n$$ with **every component strictly positive** such that $$\varphi'v = 0$$ for all $$v \in \mathcal{K}$$.

*Proof.* Let $$\mathcal{C}$$ be the unit simplex $$\lbrace c \in \mathbb{R}^n\_+ : \sum\_i c\_i = 1\rbrace$$, which is compact and convex, and note that $$\mathcal{K} \cap \mathcal{C} = \emptyset$$ by hypothesis, since every non-zero point of the orthant is a positive multiple of a point in $$\mathcal{C}$$. Apply the theorem with $$\mathcal{D} = \mathcal{K}$$: there are $$\varphi$$ and $$a$$ with $$\varphi'v < a < \varphi'c$$ for all $$v \in \mathcal{K}$$, $$c \in \mathcal{C}$$. Because $$\mathcal{K}$$ is a subspace, it contains $$tv$$ for every real $$t$$; if $$\varphi'v$$ were non-zero for some $$v \in \mathcal{K}$$, then $$\varphi'(tv) = t\varphi'v$$ would exceed $$a$$ for $$t$$ of the right sign and magnitude. So $$\varphi'v = 0$$ on $$\mathcal{K}$$, and therefore $$a > 0$$. Finally, each standard basis vector $$e\_i$$ lies in $$\mathcal{C}$$, so $$\varphi\_i = \varphi'e\_i > a > 0$$. $$\blacksquare$$

Notice what the proof used. Convexity, compactness of the simplex, and the fact that a subspace is closed under multiplication by negative scalars. Nothing about preferences, equilibrium, or behavior. The theorem is about shapes.

### D.5.3 The fundamental theorem of asset pricing: existence

> **Theorem (finite-state fundamental theorem, existence half).** There is no arbitrage if and only if there exists a strictly positive vector of state prices $$q \in \mathbb{R}^{\lvert S\rvert}$$, $$q\_s > 0$$ for every $$s$$, such that $$p\_j = \sum\_s q\_s x\_{j,s}$$ for every traded asset $$j$$.

*Proof of the hard direction.* Suppose no arbitrage. By §D.5.1,

$$
\mathcal{K} \cap \mathbb{R}^{1+\lvert S\rvert}\_+ = \lbrace 0\rbrace
$$

and $$\mathcal{K}$$ is a subspace, so the corollary supplies a strictly positive $$\varphi = (\varphi\_0, \varphi\_1, \ldots, \varphi\_{\lvert S\rvert})$$ with $$\varphi'v = 0$$ for every $$v \in \mathcal{K}$$. Writing that out for $$v = (-p'\theta, X\theta)$$,

$$
-\varphi\_0p'\theta + \sum\_s \varphi\_s\thinspace(X\theta)\_s = 0 \qquad \text{for every } \theta \in \mathbb{R}^N
$$

Since this holds for every $$\theta$$, it holds for each basis vector, giving $$\varphi\_0 p\_j = \sum\_s \varphi\_s x\_{j,s}$$ for each asset $$j$$. Divide by $$\varphi\_0 > 0$$ and define

$$
q\_s \equiv \frac{\varphi\_s}{\varphi\_0} > 0
$$

to obtain $$p\_j = \sum\_s q\_s x\_{j,s}$$. $$\blacksquare$$

*Proof of the easy direction.* Suppose such a $$q$$ exists and let $$\theta$$ satisfy $$X\theta \ge 0$$ with some component strictly positive. Then $$p'\theta = \sum\_s q\_s (X\theta)\_s > 0$$, because every $$q\_s$$ is strictly positive: the portfolio costs money, so it is not an arbitrage. The case $$X\theta = 0$$ with $$p'\theta < 0$$ is excluded the same way. $$\blacksquare$$

Read the hard direction again, because the punchline is easy to walk past. **The normal vector of the separating hyperplane, normalized by its date-0 component, is the state-price vector.** Arrow-Debreu prices are not an additional modeling assumption laid on top of no arbitrage; they are the geometric dual of it. The strict positivity of every $$q\_s$$ — the fact that a dollar in any state has a positive price — is inherited directly from the strict positivity the corollary delivers, which in turn came from the simplex being compact.

The translation to the stochastic discount factor is one line. Since every state has $$\pi\_s > 0$$, define $$m\_s = q\_s/\pi\_s$$, and

$$
p\_j = \sum\_s q\_s x\_{j,s} = \sum\_s \pi\_s \frac{q\_s}{\pi\_s} x\_{j,s} = E\[mx\_j]
$$

with $$m$$ strictly positive because $$q$$ is. That is §3.5's fundamental pricing equation, now with an existence theorem behind it.

The general version — infinitely many states, infinitely many dates, continuous trading — is substantially harder, and the difficulty is real rather than technical fussiness. In infinite dimensions a subspace need not be closed, "no arbitrage" has to be strengthened to a no-free-lunch condition, and the separating functional need not be representable by a vector. Harrison and Kreps (1979) and Harrison and Pliska (1981) established the martingale form; Delbaen and Schachermayer (1994) gave the definitive general statement. Ross (1978) is the finite-state original and is the closest published relative of the argument above. Duffie (2001) is the standard graduate treatment and is what Chapter 3's Readings point to.

### D.5.4 Completeness and uniqueness

Existence gives at least one strictly positive $$q$$. Whether there is exactly one is a question about the rank of a matrix.

Markets are **complete** when every payoff vector in $$\mathbb{R}^{\lvert S\rvert}$$ is attainable — for every target $$x$$ there is a portfolio $$\theta$$ with $$X\theta = x$$. That is precisely the statement that $$X$$ has rank $$\lvert S\rvert$$, which requires $$N \ge \lvert S\rvert$$ and no fewer than $$\lvert S\rvert$$ linearly independent columns. §3.4's economy had two assets and two states with independent payoffs, so it was complete; §3.4.7's three-state variant with the same two assets was not.

> **Theorem (uniqueness).** Given no arbitrage, the state-price vector is unique if and only if markets are complete.

*Proof.* The state prices solve the linear system $$X'q = p$$. If $$X$$ has rank $$\lvert S\rvert$$, the null space of $$X'$$ is trivial and the solution is unique. If $$X$$ has rank $$k < \lvert S\rvert$$, take the strictly positive solution $$q^0$$ that §D.5.3 supplies; the full solution set is $$q^0 + \ker(X')$$, an affine set of dimension $$\lvert S\rvert - k \ge 1$$. Its intersection with the strictly positive orthant is relatively open in that affine set and contains $$q^0$$, hence contains a neighborhood of $$q^0$$ within the affine set and therefore infinitely many points. $$\blacksquare$$

The corollary is worth stating because students expect it to be false: an arbitrage-free market has either exactly one state-price vector or infinitely many. There is no economy with two.

In discount-factor language this is the family §3.5 describes. If $$m$$ prices all traded assets, so does $$m + \varepsilon$$ for any $$\varepsilon$$ orthogonal to the payoff space, since $$E\[\varepsilon x] = 0$$ for every traded $$x$$ by construction. Under incompleteness the orthogonal complement is non-trivial, and the family is infinite. Two members deserve names.

The **strictly positive** members are the ones the fundamental theorem produces. At least one exists; under incompleteness there are many; and positivity is what rules out arbitrage, so it is the property that has economic content.

The **projection** $$m^{\ast} = \mathrm{proj}(m \mid X)$$ of §D.4.2 is the unique member lying in the payoff space, and it is the one empirical work uses, for two reasons. It is constructible from data — it is a portfolio, with weights $$p'E\[xx']^{-1}$$ — and by §D.4.2 it is the minimum-variance member, so it is the one that attains the Hansen-Jagannathan bound rather than merely satisfying it. Its cost is that it need not be positive, and when it is not, the honest reading is that the model has been asked for something the traded assets cannot deliver. Cochrane (2005, Chs 4 and 18) works through the geometry of the positive-$$m$$ frontier that this raises; Hansen and Jagannathan (1991) is where the constraint was first imposed.

§3.4.7's practical statement now has a proof behind it. Where a claim can be replicated, $$m$$'s value on that claim is pinned down and the identity of the holder is irrelevant. Where it cannot, the discount factor has a free direction, arbitrage returns a range rather than a number, and something outside arbitrage — who wants the claim, how badly, and what constrains them — settles where in the range the price lands. That "something" is Part IV of this book.

### D.5.5 Change of measure, formally

The risk-neutral probabilities of §3.4.5 are now easy to place. Assume a riskless asset trades, with gross return $$R\_f$$, so that $$\sum\_s q\_s = 1/R\_f$$. Define

$$
\pi^{\ast}\_s = q\_s R\_f = \pi\_sm\_s R\_f
$$

Three properties make this a probability measure and not merely a normalization. Each $$\pi^{\ast}\_s > 0$$, by §D.5.3. They sum to one, since $$\sum\_s q\_s R\_f = 1$$. And $$\pi^{\ast}\_s > 0$$ exactly when $$\pi\_s > 0$$, which is the definition of two measures being **equivalent** — they agree about which states are possible and disagree only about how likely. Equivalence is what no arbitrage delivers and all that it delivers.

The ratio

$$
\frac{\pi^{\ast}\_s}{\pi\_s} = m\_s R\_f
$$

is the **Radon-Nikodym derivative** of $$\pi^{\ast}$$ with respect to $$\pi$$, usually written $$d\pi^{\ast}/d\pi$$. In a finite state space it is nothing more exotic than the vector of ratios, one per state, and the general definition in measure theory is designed to say the same thing when there are too many states to list. Its expectation under $$\pi$$ is $$E\[mR\_f] = R\_f E\[m] = 1$$, as any density must have.

Expectations translate by reweighting: for any random variable $$y$$,

$$
E^{\ast}\[y] = \sum\_s \pi^{\ast}\_s y\_s = \sum\_s \pi\_sm\_s R\_fy\_s = R\_fE\[my]
$$

Setting $$y = x$$ and dividing gives the pricing formula in its risk-neutral form:

$$
p = E\[mx] = \frac{E^{\ast}\[x]}{R\_f}
$$

**Price equals the expected payoff discounted at the riskless rate, with the expectation taken under the wrong probabilities.** The risk adjustment has not vanished; it has been moved out of the discount rate and into the measure, where it is carried by $$m R\_f$$ — that is, by marginal utility, normalized. States are made to look more likely in exact proportion to how much a dollar is worth in them.

Iterating over dates gives the **martingale** property: the deflated price process $$p\_t/R\_f^{t}$$ satisfies

$$
E^{\ast}\_t\[p \_{t+1}/R\_f^{t+1}] = p\_t/R\_f^{t}
$$

which is why $$\pi^{\ast}$$ is also called an equivalent martingale measure. This is the statement §8.3.2 makes for the binomial tree and §D.7.4 uses in continuous time, and it is the compact form of the whole of Part II: under one particular reweighting of the probabilities, every price in the economy is a martingale.

***

## D.6 CARA-Normal Expected Utility

*Consumed by: Ch 7 §7.2; Ch 11's Kyle model, which is the same machinery with a different information structure.*

§7.2 states the CARA-normal demand function "in words rather than deriving it," and its starred subsection states the expected-utility ratio and sketches why the level terms cancel. This section carries both derivations. Why the model is written this way is worth naming first: constant absolute risk aversion plus normal payoffs is the only combination making demand linear in the conditional mean and independent of wealth, and linearity is what permits an equilibrium price that is itself a normal signal. The tractability is bought with an assumption nobody believes — that a billionaire and a graduate student buy the same number of shares — and the model earns its place by isolating the information mechanism, not by being realistic about preferences.

### D.6.1 The certainty equivalent

Let utility be $$U(W) = -e^{-\tau W}$$, with $$\tau > 0$$ the coefficient of **absolute** risk aversion — the curvature of utility in dollars rather than in proportions, which is what distinguishes $$\tau$$ from the book's $$\gamma$$. A trader with information set $$\Omega$$ buys $$D$$ units of a claim at price $$p$$, financing the purchase at the riskless gross return $$R\_f$$, so terminal wealth is

$$
W = R\_f W\_0 + D\thinspace(x - R\_f p)
$$

Write $$\hat\mu = E\[x \mid \Omega]$$ and $$\hat\sigma^2 = \mathrm{Var}\[x \mid \Omega]$$. Conditional on $$\Omega$$, $$W$$ is normal with mean $$R\_fW\_0 + D(\hat\mu - R\_fp)$$ and variance $$D^2\hat\sigma^2$$. Apply §D.2.2's moment formula to the exponential:

$$
E\big\[-e^{-\tau W} \thinspace\big|\thinspace \Omega\big] = -\exp\left(-\tau\big\[R\_fW\_0 + D(\hat\mu - R\_fp)\big] + \tfrac{1}{2}\tau^2 D^2\hat\sigma^2\right)
$$

Because $$-e^{-\tau\cdot}$$ is a strictly increasing transformation, maximizing this is the same as maximizing the exponent's negative, and the **certainty equivalent** — the certain wealth that would deliver the same utility — reads off directly:

$$
\mathrm{CE} = R\_fW\_0 + D(\hat\mu - R\_fp) - \frac{\tau}{2}D^2\hat\sigma^2
$$

Mean minus $$\tau/2$$ times variance. Mean-variance preferences here are not an approximation; under CARA and normality they are exact, which is the second reason the pairing is standard.

### D.6.2 The demand function

Differentiate the certainty equivalent with respect to $$D$$ and set to zero:

$$
\hat\mu - R\_fp - \tau D\hat\sigma^2 = 0 \quad\Longrightarrow\quad D = \frac{E\[x\mid\Omega] - R\_fp}{\tau\mathrm{Var}\[x\mid\Omega]}
$$

which is §7.2's demand function. Note $$W\_0$$ has disappeared: initial wealth entered the certainty equivalent additively and therefore not at all in the derivative. That is the wealth-independence the model needs in order to aggregate across traders without tracking who is rich.

Substituting the optimum back gives the value of being informed, which §D.6.4 needs:

$$
\mathrm{CE}^{\ast} = R\_fW\_0 + \frac{\big(\hat\mu - R\_fp\big)^2}{2\tau\hat\sigma^2}, \qquad E\[U\mid\Omega] = -\exp\left(-\tau R\_fW\_0 - \frac{(\hat\mu - R\_fp)^2}{2\hat\sigma^2}\right)
$$

The gain from information is quadratic in the perceived mispricing and inversely proportional to the conditional variance. Information helps by shrinking the denominator, which lets the trader take a larger position — not by making the asset cheap.

### D.6.3 Normal-normal updating

Let $$y \sim N(\bar y, \sigma\_y^2)$$ and let a signal $$w = y + u$$ be observed, with $$u \sim N(0, \sigma\_u^2)$$ independent of $$y$$. Then $$y$$ and $$w$$ are jointly normal, and the conditional distribution of $$y$$ given $$w$$ is normal with

$$
E\[y\mid w] = \frac{\sigma\_u^{-2}}{\sigma\_y^{-2} + \sigma\_u^{-2}}w + \frac{\sigma\_y^{-2}}{\sigma\_y^{-2} + \sigma\_u^{-2}}\bar y, \qquad \mathrm{Var}\[y\mid w] = \frac{1}{\sigma\_y^{-2} + \sigma\_u^{-2}}
$$

**Precisions add and the posterior mean is a precision-weighted average of the prior and the signal.** Defining the signal-to-noise ratio $$\Psi = \sigma\_y^2/\sigma\_u^2$$ puts the variance in the form Chapter 7 uses:

$$
\mathrm{Var}\[y \mid w] = \frac{\sigma\_y^2}{1 + \Psi}
$$

At $$\Psi = 0$$ the signal is pure noise and the posterior is the prior; as $$\Psi \to \infty$$ the posterior variance goes to zero. **Price informativeness**, defined as the fraction of the prior variance the signal removes, is $$\Psi/(1+\Psi)$$, which is Chapter 7's measure.

### D.6.4 The unconditional expectation of the exponential-quadratic

This is the step §7.2's starred subsection sketches. It needs one lemma.

> **Lemma.** If $$v \sim N(\mu\_v, \sigma\_v^2)$$ and $$a > -1/\sigma\_v^2$$, then
>
> $$
> E\left\[e^{-\frac{a}{2}v^2}\right] = \left(1 + a\sigma\_v^2\right)^{-1/2}\exp\left(-\frac{a\mu\_v^2}{2\left(1 + a\sigma\_v^2\right)}\right)
> $$

*Proof.* Write the expectation as an integral against the normal density and collect the exponent, which is $$-\tfrac{1}{2}\big\[(a + \sigma\_v^{-2})v^2 - 2\mu\_v\sigma\_v^{-2}v + \mu\_v^2\sigma\_v^{-2}\big]$$. Complete the square in $$v$$ with leading coefficient $$a + \sigma\_v^{-2}$$; the Gaussian integral contributes $$\big\[\sigma\_v^2(a + \sigma\_v^{-2})\big]^{-1/2} = (1 + a\sigma\_v^2)^{-1/2}$$, and the residual constant is $$-\tfrac{1}{2}\mu\_v^2\sigma\_v^{-2}\big\[1 - (a\sigma\_v^2 + 1)^{-1}\big] = -a\mu\_v^2/\big\[2(1 + a\sigma\_v^2)\big]$$. $$\blacksquare$$

Apply it with $$v = \hat\mu - R\_fp$$, the conditional expected excess payoff as seen ex ante, and $$a = 1/\hat\sigma^2$$. From §D.6.2,

$$
E\thinspace U = -e^{-\tau R\_f W\_0}E\left\[\exp\left(-\frac{v^2}{2\hat\sigma^2}\right)\right] = -e^{-\tau R\_fW\_0}\left(\frac{\hat\sigma^2}{\hat\sigma^2 + \sigma\_v^2}\right)^{1/2}\exp\left(-\frac{\mu\_v^2}{2\left(\hat\sigma^2 + \sigma\_v^2\right)}\right)
$$

Now the cancellation that Chapter 7 asserts. Whatever the trader's information set, the excess payoff decomposes as

$$
x - R\_f p = \underbrace{\big(\hat\mu - R\_fp\big)}\_{v} + \underbrace{\big(x - \hat\mu\big)} \_{\text{forecast error}}
$$

and the forecast error is, by construction, mean zero, uncorrelated with anything in $$\Omega$$, and of variance $$\hat\sigma^2$$. Therefore, **for every type of trader**,

$$
\mu\_v = E\big\[x - R\_fp\big], \qquad \hat\sigma^2 + \sigma\_v^2 = \mathrm{Var}\big(x - R\_fp\big)
$$

Both are properties of the traded claim and the equilibrium price, not of the information set. The exponential factor and the denominator inside the square root are therefore identical across types, and in the ratio they cancel — leaving only the numerators, which are the conditional variances:

$$
\frac{E\thinspace U\_{\text{informed}}}{E\thinspace U\_{\text{uninformed}}} = \left(\frac{\mathrm{Var}\[x\mid y]}{\mathrm{Var}\[x\mid p]}\right)^{1/2}
$$

Finally, the cost of information. Charging $$\xi$$ against terminal wealth replaces $$W$$ by $$W - \xi$$ and multiplies $$-e^{-\tau W}$$ by exactly $$e^{\tau\xi}$$ — the property that makes a fixed charge tractable under CARA and intractable under almost anything else. So

$$
\frac{E\thinspace U\_{\text{informed}}}{E\thinspace U\_{\text{uninformed}}} = e^{\tau\xi}\left(\frac{\mathrm{Var}\[x\mid y]}{\mathrm{Var}\[x\mid p]}\right)^{1/2}
$$

Both expected utilities are negative, so the informed are better off when the ratio is *below* one, and free entry is the ratio equal to one. Squaring gives Chapter 7's **free-entry condition**,

$$
\frac{\mathrm{Var}\[x\mid p]}{\mathrm{Var}\[x\mid y]} = e^{2\tau\xi}
$$

with $$\mathrm{Var}\[x\mid y] = \sigma\_\varepsilon^2$$. The level of the expected excess payoff — the risk premium, whatever it is — has dropped out, which is why the equilibrium informational gap depends on the cost of information and on nothing else.

### D.6.5 Solving the linear equilibrium

The remaining gap is why the price carries exactly the information in $$w = y - (\tau\sigma\_\varepsilon^2/n)z$$, where $$n$$ is the informed fraction and $$z$$ the random per-capita supply. Conjecture that the price is a strictly monotone function of some scalar, so that conditioning on $$p$$ is conditioning on that scalar; then verify.

Market clearing sets aggregate demand equal to supply:

$$
n\frac{y - R\_fp}{\tau\sigma\_\varepsilon^2} + (1-n)\frac{E\[x\mid p] - R\_fp}{\tau\mathrm{Var}\[x\mid p]} = z
$$

Multiply through by $$\tau\sigma\_\varepsilon^2/n$$ and move terms:

$$
y - \frac{\tau\sigma\_\varepsilon^2}{n}z = R\_fp - \frac{(1-n)\sigma\_\varepsilon^2}{n\mathrm{Var}\[x\mid p]}\Big(E\[x\mid p] - R\_fp\Big)
$$

The right-hand side is a function of $$p$$ alone. So $$w$$, the left-hand side, is a function of $$p$$, and the conjecture is confirmed: observing $$p$$ is informationally equivalent to observing $$w$$, whatever the coefficients of the price function turn out to be.

That licenses conditioning on $$w$$ directly. It is normal, with mean $$\bar y - (\tau\sigma\_\varepsilon^2/n)\bar z$$ and noise variance $$(\tau\sigma\_\varepsilon^2/n)^2\sigma\_z^2$$, so §D.6.3 gives

$$
\Psi(n) = \frac{\sigma\_y^2}{\big(\tau\sigma\_\varepsilon^2/n\big)^2\sigma\_z^2} = \frac{n^2\sigma\_y^2}{\tau^2\sigma\_z^2\sigma\_\varepsilon^4}, \qquad \mathrm{Var}\[x\mid p] = \sigma\_\varepsilon^2 + \frac{\sigma\_y^2}{1+\Psi}
$$

Substituting into the free-entry condition, $$\sigma\_\varepsilon^2 + \sigma\_y^2/(1+\Psi) = e^{2\tau\xi}\sigma\_\varepsilon^2$$, and solving,

$$
\Psi^{\ast} = \frac{\sigma\_y^2}{\left(e^{2\tau\xi} - 1\right)\sigma\_\varepsilon^2} - 1, \qquad n^{\ast} = \frac{\tau\sigma\_z\sigma\_\varepsilon^2}{\sigma\_y}\sqrt{\Psi^{\ast}}
$$

truncated to $$\[0,1]$$, which are the expressions Chapter 7 tabulates.

The price itself follows. Substitute $$y = w + (\tau\sigma\_\varepsilon^2/n)z$$ into the market-clearing condition; the supply shock cancels, and writing $$\hat\mu\_U = E\[x\mid w]$$ and $$V\_U = \mathrm{Var}\[x\mid w]$$,

$$
R\_fp = \frac{\dfrac{n}{\tau\sigma\_\varepsilon^2}w + \dfrac{1-n}{\tau V\_U}\hat\mu\_U}{\dfrac{n}{\tau\sigma\_\varepsilon^2} + \dfrac{1-n}{\tau V\_U}}
$$

The equilibrium price is a precision-weighted average of the informed traders' statistic and the uninformed traders' posterior mean, with the weights being each group's mass divided by its conditional variance. It is affine in $$w$$ with a positive coefficient, confirming the monotonicity the conjecture required, and the risk premium enters through the mean of $$w$$, which carries $$\bar z$$. In the limiting case $$n = 1$$ the expression collapses to $$R\_fp = w = y - \tau\sigma\_\varepsilon^2 z$$, informed demand equals $$z$$, and the expected excess payoff is $$\tau\sigma\_\varepsilon^2\bar z$$ — risk aversion times conditional variance times average supply, which is the compensation for bearing the claim and is exactly what a one-asset CAPM would charge.

***

## D.7 Stochastic Calculus, at the Minimum Needed

*Consumed by: Ch 8 (body and the derivation §8.4 gates here); Ch 13's option-adjusted spread discussion.*

### D.7.1 Brownian motion

A **standard Brownian motion** (or Wiener process) $$z\_t$$ is defined by three properties:

1. $$z\_0 = 0$$, and the path $$t \mapsto z\_t$$ is continuous.
2. Increments over disjoint intervals are independent.
3. The increment over an interval of length $$\Delta t$$ is normal with mean zero and variance $$\Delta t$$: $$z\_{t + \Delta t} - z\_t \sim N(0, \Delta t)$$.

Three properties, and every consequence in this section is one of them applied carefully. The one that does the real work is the third, because it says the *standard deviation* of an increment is $$\sqrt{\Delta t}$$ rather than $$\Delta t$$. Over a short interval, the random part of the motion is large relative to anything proportional to $$\Delta t$$ — which is why a Brownian path is continuous everywhere and differentiable nowhere, and why the calculus of Brownian motion is not the calculus of smooth functions.

The precise form of that statement is **quadratic variation**. Partition $$\[0, T]$$ into $$n$$ equal pieces and sum the squared increments. Each squared increment has expectation $$T/n$$, so the sum has expectation $$T$$ regardless of $$n$$; and its variance goes to zero as $$n$$ grows. In the limit the sum converges to $$T$$ exactly rather than to a random quantity. Written as a rule of thumb,

$$
(dz)^2 = dt
$$

For a smooth function the analogous sum vanishes, which is why second-order terms are discarded in ordinary calculus. Here they are not. Everything strange about Itô's lemma is this one line. Karatzas and Shreve (1991) construct the process and prove convergence properly; the construction requires measure theory and is the boundary this appendix declines to cross.

### D.7.2 Geometric Brownian motion

§8.4 says that as the binomial step shrinks, "the multiplicative random walk converges to geometric Brownian motion." The limit object is the solution to

$$
dS = \mu S\thinspace dt + \sigma S\thinspace dz
$$

read as: over an instant, the stock's proportional change $$dS/S$$ has mean $$\mu\thinspace dt$$ and standard deviation $$\sigma\sqrt{dt}$$. Proportional, not absolute — which is the modeling content, and which guarantees $$S$$ stays positive, since a proportional shock cannot take a positive number below zero.

Its solution, derived in §D.7.3, is

$$
S\_T = S\_0\exp\left\[\left(\mu - \tfrac{1}{2}\sigma^2\right)T + \sigma\big(z\_T - z\_0\big)\right]
$$

so that $$\ln S\_T$$ is normal with mean $$\ln S\_0 + (\mu - \tfrac{1}{2}\sigma^2)T$$ and variance $$\sigma^2 T$$: $$S\_T$$ **is lognormal**, and §D.2's formulas apply to it verbatim. In particular $$E\[S\_T] = S\_0 e^{\mu T}$$, with the $$-\tfrac{1}{2}\sigma^2$$ in the exponent and the $$+\tfrac{1}{2}\sigma^2$$ from the lognormal moment formula cancelling exactly. The drift of the *log* is smaller than the drift of the *level* by $$\tfrac{1}{2}\sigma^2$$, which is the Jensen correction of §D.2.2 wearing a different hat, and which is the same $$\tfrac{1}{2}\sigma^2$$ that appears in $$d\_1$$ and $$d\_2$$.

### D.7.3 Itô's lemma

> **Itô's lemma.** Let $$S$$ follow $$dS = a(S,t)\thinspace dt + b(S,t)\thinspace dz$$ and let $$f(S,t)$$ be twice continuously differentiable in $$S$$ and once in $$t$$. Then
>
> $$
> df = \left(\frac{\partial f}{\partial t} + a\frac{\partial f}{\partial S} + \frac{1}{2}b^2\frac{\partial^2 f}{\partial S^2}\right)dt + b\frac{\partial f}{\partial S}\thinspace dz
> $$

Here is the paragraph the chapters keep deferring. Expand $$f$$ to second order:

$$
df = \frac{\partial f}{\partial t}dt + \frac{\partial f}{\partial S}dS + \frac{1}{2}\frac{\partial^2 f}{\partial S^2}(dS)^2 + \cdots
$$

In ordinary calculus the third term is discarded because $$(dS)^2$$ is of order $$(dt)^2$$. Here it is not. Substituting $$dS = a\thinspace dt + b\thinspace dz$$ and squaring,

$$
(dS)^2 = a^2(dt)^2 + 2ab\thinspace dt\thinspace dz + b^2(dz)^2
$$

The first term is order $$(dt)^2$$ and vanishes. The second is order $$dt \cdot \sqrt{dt} = (dt)^{3/2}$$ and vanishes. The third is $$b^2(dz)^2 = b^2\thinspace dt$$ by §D.7.1, and it does **not** vanish — it is first order in $$dt$$, the same order as the terms being kept. Collecting the $$dt$$ terms gives the lemma.

So Itô's lemma is the chain rule plus one extra term, and the extra term is the second derivative times half the instantaneous variance. Its economic content is that **a convex function of a volatile variable drifts upward relative to the same function of a smooth variable**, at a rate proportional to the curvature. Convexity pays; that is gamma, and it is why Chapter 8's Greeks come in a pair.

**The worked application.** Take $$f = \ln S$$ with $$S$$ geometric Brownian, so $$a = \mu S$$, $$b = \sigma S$$, and

$$
\frac{\partial f}{\partial t} = 0, \qquad \frac{\partial f}{\partial S} = \frac{1}{S}, \qquad \frac{\partial^2 f}{\partial S^2} = -\frac{1}{S^2}
$$

Substituting into the lemma,

$$
d(\ln S) = \left(0 + \mu S \cdot \frac{1}{S} + \frac{1}{2}\sigma^2S^2\cdot\left(-\frac{1}{S^2}\right)\right)dt + \sigma S\cdot\frac{1}{S}\thinspace dz = \left(\mu - \tfrac{1}{2}\sigma^2\right)dt + \sigma\thinspace dz
$$

The log of the stock has *constant* drift and *constant* diffusion, so it is ordinary Brownian motion with drift and integrates immediately to §D.7.2's solution. The $$-\tfrac{1}{2}\sigma^2$$ was manufactured by the second-order term; naive application of the chain rule would have produced $$\mu\thinspace dt + \sigma\thinspace dz$$ and a stock with the wrong expected value. Every $$\tfrac{1}{2}\sigma^2$$ in Chapter 8 traces to this line.

### D.7.4 The Black-Scholes derivation

§8.4 gates this here. Two routes, and both are worth having: the hedging route explains *why* the price is what it is, and the risk-neutral route computes it.

**Route one: the hedged portfolio and the PDE.**

Let $$C(S,t)$$ be the value of a European call on a non-dividend-paying stock following $$dS = \mu S\thinspace dt + \sigma S\thinspace dz$$, with continuously compounded riskless rate $$r$$. By Itô's lemma,

$$
dC = \left(\frac{\partial C}{\partial t} + \mu S\frac{\partial C}{\partial S} + \frac{1}{2}\sigma^2S^2\frac{\partial^2C}{\partial S^2}\right)dt + \sigma S\frac{\partial C}{\partial S}\thinspace dz
$$

The option and the stock are driven by the *same* $$dz$$. That is the whole idea: two claims on one source of uncertainty can be combined to cancel it. Form the portfolio

$$
\Pi\_t = -C + \frac{\partial C}{\partial S}S
$$

— short one call, long $$\partial C/\partial S$$ shares — and hold the share count fixed over the instant $$dt$$. Then

$$
d\Pi = -dC + \frac{\partial C}{\partial S}\thinspace dS = -\left(\frac{\partial C}{\partial t} + \frac{1}{2}\sigma^2S^2\frac{\partial^2C}{\partial S^2}\right)dt
$$

Both the $$\mu$$ terms and the $$dz$$ terms have cancelled. The first cancellation is §8.3.3's "the true drift drops out," now visible as algebra rather than as an assertion; the second is the hedge. What remains is a portfolio whose change over the instant is known with certainty, and by no arbitrage a certain return must be the riskless one: $$d\Pi = r\Pi\thinspace dt$$. Substituting and cancelling $$dt$$,

$$
-\frac{\partial C}{\partial t} - \frac{1}{2}\sigma^2S^2\frac{\partial^2C}{\partial S^2} = r\left(\frac{\partial C}{\partial S}S - C\right)
$$

which rearranges to the **Black-Scholes partial differential equation**:

$$
\frac{\partial C}{\partial t} + \frac{1}{2}\sigma^2S^2\frac{\partial^2 C}{\partial S^2} + rS\frac{\partial C}{\partial S} - rC = 0
$$

with **boundary conditions** $$C(S,T) = \max(S - K, 0)$$ at expiry, $$C(0,t) = 0$$ for all $$t$$ (a worthless stock cannot make the option valuable, since zero is absorbing), and $$C(S,t)/S \to 1$$ as $$S \to \infty$$ (deep in the money the call is the stock, less the discounted strike).

Two honest notes. The step "hold the share count fixed over $$dt$$" is the informal version of a **self-financing** condition, and making it precise requires the stochastic integral rather than a heuristic differential; Harrison and Pliska (1981) and Duffie (2001) do it properly, and nothing in the answer changes. And the PDE has already delivered the substantive result: **the option's price does not depend on** $$\mu$$**.** Two investors who disagree completely about the stock's expected return must agree on the option's price, because $$\mu$$ appears nowhere in the equation they must both solve.

The PDE is a **backward parabolic equation**, and three substitutions convert it into the heat equation of physics: replace $$S$$ by log-moneyness $$\ln(S/K)$$, replace calendar time by scaled time to expiry $$\tfrac{1}{2}\sigma^2(T-t)$$, and rescale $$C$$ by an exponential factor in those two variables chosen to kill the first- and zeroth-order terms. The heat equation with a known initial condition has a known solution — a convolution of that condition with the Gaussian kernel — and undoing the substitutions produces the formula. Wilmott, Howison and Dewynne (1995, Ch 5) carry that algebra, which is mechanical and long; the route below reaches the same place in a page.

**Route two: risk-neutral valuation.**

Section D.5.5 established that pricing is expectation under $$\pi^{\ast}$$, discounted at the riskless rate. The continuous-time version says the same thing: there is a measure $$\pi^{\ast}$$, equivalent to the physical one, under which the discounted stock price is a martingale — equivalently, under which the stock's drift is $$r$$ rather than $$\mu$$. The theorem that constructs it is **Girsanov's**, and it is the one genuinely measure-theoretic ingredient in this section; Björk (2009, Ch 11) and Shreve (2004, Ch 5) state and prove it. Its content, in words, is that changing the drift of a Brownian motion is a change of measure and not a change of the possible paths, which is exactly the equivalence of §D.5.5.

Under $$\pi^{\ast}$$, then, §D.7.2 gives $$\ln S\_T$$ normal with mean $$\ln S\_0 + (r - \tfrac{1}{2}\sigma^2)T$$ and variance $$\sigma^2T$$, so writing $$\varepsilon$$ for a standard normal draw under $$\pi^{\ast}$$,

$$
S\_T = S\_0\exp\left\[\left(r - \tfrac{1}{2}\sigma^2\right)T + \sigma\sqrt{T}\varepsilon\right]
$$

The call is worth its discounted risk-neutral expected payoff:

$$
C\_0 = e^{-rT}E^{\ast}\big\[\max(S\_T - K, 0)\big]
$$

The option finishes in the money when $$S\_T > K$$, which happens when

$$
\varepsilon > \frac{\ln(K/S\_0) - \left(r - \tfrac{1}{2}\sigma^2\right)T}{\sigma\sqrt{T}} = -d\_2, \qquad d\_2 \equiv \frac{\ln(S\_0/K) + \left(r - \tfrac{1}{2}\sigma^2\right)T}{\sigma\sqrt{T}}
$$

Split the integral over that region into two pieces.

*The strike piece.* Using the symmetry $$1 - N(-d\_2) = N(d\_2)$$,

$$
e^{-rT}K\int\_{-d\_2}^{\infty}\phi(\varepsilon)\thinspace d\varepsilon = Ke^{-rT}N(d\_2)
$$

*The stock piece.* Pull out the constants and complete the square in the exponent, using $$\sigma\sqrt{T}\varepsilon - \tfrac{1}{2}\varepsilon^2 = -\tfrac{1}{2}\left(\varepsilon - \sigma\sqrt T\right)^2 + \tfrac{1}{2}\sigma^2T$$:

$$
e^{-rT}S\_0e^{\left(r - \frac{1}{2}\sigma^2\right)T}\int\_{-d\_2}^{\infty}e^{\sigma\sqrt T\varepsilon}\phi(\varepsilon)\thinspace d\varepsilon = S\_0\int\_{-d\_2}^{\infty}\phi\left(\varepsilon - \sigma\sqrt{T}\right)d\varepsilon = S\_0N\left(d\_2 + \sigma\sqrt{T}\right)
$$

where the $$e^{-rT}$$, the $$e^{rT}$$, the $$e^{-\sigma^2T/2}$$ and the $$e^{+\sigma^2T/2}$$ have all cancelled, and the last step substituted the shifted variable and used the symmetry again. Defining $$d\_1 = d\_2 + \sigma\sqrt T$$ and subtracting,

$$
\boxed{C\_0 = S\_0N(d\_1) - Ke^{-rT}N(d\_2)}
$$

$$
d\_1 = \frac{\ln(S\_0/K) + \left(r + \tfrac{1}{2}\sigma^2\right)T}{\sigma\sqrt{T}}, \qquad d\_2 = d\_1 - \sigma\sqrt{T}
$$

which is §8.4's formula. The put follows from parity.

Two closing observations. The completion of the square is where $$d\_1$$ acquires $$+\tfrac{1}{2}\sigma^2$$ against $$d\_2$$'s $$-\tfrac{1}{2}\sigma^2$$, which is the whole difference between them and is the reason Chapter 8 can read $$N(d\_2)$$ as a probability under $$\pi^{\ast}$$ and $$N(d\_1)$$ as the corresponding probability under a measure in which the stock rather than the money-market account is the unit of account. And differentiating the formula with respect to $$S\_0$$ gives $$\partial C/\partial S\_0 = N(d\_1)$$ exactly, because the two density terms $$S\_0\phi(d\_1)\partial d\_1/\partial S\_0$$ and $$Ke^{-rT}\phi(d\_2)\partial d\_2/\partial S\_0$$ cancel identically — a small miracle that is really the statement that the hedge ratio is the hedge ratio. Delta is $$N(d\_1)$$, and route one's portfolio was $$\partial C/\partial S$$ shares all along.

### D.7.5 What the derivation assumes

§8.4 walks through the assumptions and says what each buys. Here is the same list mapped onto the line of algebra it is protecting, so the reader can see which step fails when the assumption does.

**Table D.3: Each Black-Scholes assumption and the step it protects**

| Assumption (Ch 8 §8.4)                                | The step it protects                                                                                          | What breaks without it                                                                                                                                                                                        |
| ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Constant, known $$\sigma$$                            | The coefficient $$\tfrac{1}{2}\sigma^2S^2$$ in the PDE, and $$\sigma^2T$$ in the lognormal variance           | $$\sigma$$ becomes a second state variable; the PDE gains a dimension and has no closed-form solution. This is §8.5's smile                                                                                   |
| Lognormal returns with continuous paths               | $$(dz)^2 = dt$$, and the cancellation of the $$dz$$ terms in $$d\Pi$$                                         | A jump is not a small move, so a hedge chosen to cancel small moves does not cancel it. The option stops being replicable, and a formula that prices manufacturing cost prices nothing                        |
| Continuous, costless trading                          | Holding $$\partial C/\partial S$$ fixed over $$dt$$ and rebalancing thereafter — the self-financing condition | The hedge becomes discrete and approximate; the "price" becomes a band whose width grows with the rebalancing interval and the cost                                                                           |
| Frictionless borrowing and shorting at one rate $$r$$ | The step $$d\Pi = r\Pi\thinspace dt$$, and the discounting $$e^{-rT}$$                                        | Two rates give two prices and a no-arbitrage interval. A hard-to-borrow name carries the borrow cost as a negative dividend                                                                                   |
| No dividends                                          | That the drift of $$dS$$ is the stock's total return                                                          | Removable, not an assumption. With payout yield $$y$$, $$dS = (\mu - y)S\thinspace dt + \sigma S\thinspace dz$$ and $$S\_0$$ is replaced by $$S\_0e^{-yT}$$, which prices index, currency and futures options |
| Price-taking hedger                                   | That $$\mu$$, $$\sigma$$ and the law of motion of $$S$$ do not respond to the hedger's trades                 | The hedging demand feeds back into $$S$$; delta hedging becomes destabilizing rather than neutral. October 1987, and §8.6 in miniature every month                                                            |

*Source: Author's construction, mapping Ch 8 §8.4's assumption list onto §D.7.4's derivation.*

The pattern is worth naming. Every assumption in the list is protecting either the *exactness of the hedge* or the *existence of a single discount rate*. Black-Scholes is not a statistical model of stock returns that happens to price options; it is an engineering claim about a manufacturing process, and the assumptions are the tolerances the process requires. When they fail, the formula does not become a slightly worse forecast — it stops being a replication argument, which is a different kind of failure and the one §8.5 is about.

***

## D.8 Probability Facts Used Without Proof

*Consumed by: Chs 3, 5, 7, 8.*

**Jensen's inequality.** For a convex function $$f$$ and a random variable $$X$$ with finite mean, $$E\[f(X)] \ge f(E\[X])$$, with equality only if $$f$$ is linear on the support of $$X$$ or $$X$$ is degenerate. Reversed for concave $$f$$. This is the source of the $$\tfrac{1}{2}\sigma^2$$ in §D.2, of the gap between arithmetic and geometric average returns, and of the fact that $$E\[1/m] \ne 1/E\[m]$$ — which is why the riskless rate is $$1/E\[m]$$ and not $$E\[1/m]$$, a distinction that has caused more sign errors than any other in this literature.

**The law of iterated expectations.** For any information set $$\Omega$$, $$E\big\[E\[X \mid \Omega]\big] = E\[X]$$, and more generally $$E\big\[E\[X\mid\Omega\_2] \thinspace\big|\thinspace \Omega\_1\big] = E\[X\mid\Omega\_1]$$ whenever $$\Omega\_1 \subset \Omega\_2$$ — the tower property. Its most consequential appearance in this book is §7.4's excess-volatility argument. If $$P\_t = E\[P^{\ast}\_t \mid \Omega\_t]$$, where $$P^{\ast}\_t$$ is the perfect-foresight price, then the variance decomposition

$$
\mathrm{Var}\big(P^{\ast}\big) = \mathrm{Var}\big(E\[P^{\ast}\mid\Omega]\big) + E\big\[\mathrm{Var}(P^{\ast}\mid\Omega)\big] = \mathrm{Var}(P) + E\big\[\mathrm{Var}(P^{\ast}\mid\Omega)\big]
$$

gives $$\sigma(P) \le \sigma(P^{\ast})$$ immediately, since the second term is non-negative. **A forecast is less variable than the thing it forecasts.** Shiller's bound is that inequality met with data, and the objections to it are all objections to the premise $$P\_t = E\[P^{\ast}\_t\mid\Omega\_t]$$ — a constant discount rate — rather than to the algebra.

**Covariance decompositions.** $$E\[XY] = E\[X]E\[Y] + \mathrm{Cov}(X,Y)$$, used to turn $$1 = E\[mR]$$ into the beta representation in one line. $$\mathrm{Var}(X + Y) = \mathrm{Var}(X) + \mathrm{Var}(Y) + 2\mathrm{Cov}(X,Y)$$, used in §D.2.3. And the law of total covariance, which is the two-variable version of the decomposition above.

**Cauchy-Schwarz.** $$\lvert\mathrm{Cov}(X,Y)\rvert \le \sigma(X)\sigma(Y)$$, equivalently $$\lvert\rho\_{X,Y}\rvert \le 1$$. This one inequality delivers the Hansen-Jagannathan bound of §5.4 and the maximum-Sharpe-ratio statement at the end of §3.5, and it delivers $$D\_\Sigma > 0$$ in §D.3.3. It is doing more work in this book than any other result in it.

**Martingales.** A process $$\lbrace y\_t\rbrace$$ is a **martingale** with respect to an information sequence $$\lbrace \Omega\_t\rbrace$$ if $$E\[y\_{t+1} \mid \Omega\_t] = y\_t$$: the best forecast of tomorrow's value is today's. The definition says nothing about the distribution of the increment, only about its mean, which is why "prices are a martingale" is compatible with volatility clustering, fat tails, and everything else the data show.

Two martingale statements in this book must not be confused. Under the *physical* measure, prices are approximately a martingale after adjustment for a drift — that is Chapter 7's weak-form efficiency, an empirical claim that can be false. Under the *risk-neutral* measure, deflated prices are a martingale exactly — that is §D.5.5's theorem, a consequence of no arbitrage that cannot be false unless arbitrage exists. The first is a claim about information; the second is a claim about geometry. Confusing them is the most common error in reading this literature, and the whole of Part II is in the difference.

***

## Where the Proofs Live

Author-date, and each entry says which step it finishes.

* Björk, T. (2009). *Arbitrage Theory in Continuous Time*, 3rd ed. Oxford University Press. *Girsanov, and the bridge from §D.5 to §D.7.*
* Black, F. and M. Scholes (1973). "The Pricing of Options and Corporate Liabilities." *Journal of Political Economy* 81(3): 637-654. *The original of §D.7.4's route one.*
* Boyd, S. and L. Vandenberghe (2004). *Convex Optimization.* Cambridge University Press. *Kuhn-Tucker, duality, and the quadratic programs constrained portfolio choice actually requires. Free from the authors.*
* Cochrane, J. H. (2005). *Asset Pricing*, revised ed. Princeton University Press. *Chapters 4-6 are the discount-factor geometry of §§D.4-D.5, including the frontier of positive discount factors this appendix only names.*
* Delbaen, F. and W. Schachermayer (1994). "A General Version of the Fundamental Theorem of Asset Pricing." *Mathematische Annalen* 300: 463-520. *The definitive general form of §D.5.3, and the reason its finite-state proof cannot be waved at infinite dimensions.*
* Duffie, D. (2001). *Dynamic Asset Pricing Theory*, 3rd ed. Princeton University Press. *The graduate treatment of §§D.5 and D.7, and the reference Ch 3's Readings send the reader to.*
* Hansen, L. P. and R. Jagannathan (1991). "Implications of Security Market Data for Models of Dynamic Economies." *Journal of Political Economy* 99(2): 225-262. *The bound of §D.4.2, including its positivity-constrained version.*
* Harrison, J. M. and D. M. Kreps (1979). "Martingales and Arbitrage in Multiperiod Securities Markets." *Journal of Economic Theory* 20(3): 381-408. *No arbitrage as the existence of an equivalent martingale measure.*
* Harrison, J. M. and S. R. Pliska (1981). "Martingales and Stochastic Integrals in the Theory of Continuous Trading." *Stochastic Processes and their Applications* 11(3): 215-260. *Self-financing strategies made precise — the step §D.7.4 flags.*
* Karatzas, I. and S. E. Shreve (1991). *Brownian Motion and Stochastic Calculus*, 2nd ed. Springer. *The construction of §D.7.1 and the proof of Itô's lemma.*
* Merton, R. C. (1972). "An Analytic Derivation of the Efficient Portfolio Frontier." *Journal of Financial and Quantitative Analysis* 7(4): 1851-1872. *Section D.3.3, in the original, with the four scalars.*
* Michaud, R. O. (1989). "The Markowitz Optimization Enigma: Is 'Optimized' Optimal?" *Financial Analysts Journal* 45(1): 31-42; with Black, F. and R. Litterman (1992), *Financial Analysts Journal* 48(5): 28-43; Jagannathan, R. and T. Ma (2003), *Journal of Finance* 58(4): 1651-1683; and Ledoit, O. and M. Wolf (2004), *Journal of Multivariate Analysis* 88(2): 365-411. *The four repairs named in §D.4.1, in the order the text names them.*
* Rockafellar, R. T. (1970). *Convex Analysis.* Princeton University Press. *The separating hyperplane theorem of §D.5.2 in full generality.*
* Ross, S. A. (1978). "A Simple Approach to the Valuation of Risky Streams." *Journal of Business* 51(3): 453-475. *The finite-state fundamental theorem, and the closest published relative of §D.5.3.*
* Shreve, S. E. (2004). *Stochastic Calculus for Finance II: Continuous-Time Models.* Springer. *Written for exactly this appendix's reader, one level up. The obvious next book after §D.7.*
* Simon, C. P. and L. Blume (1994). *Mathematics for Economists.* Norton. *Lagrangians, Kuhn-Tucker with the regularity conditions, quadratic forms, and the envelope theorem, at the level §§D.3-D.4 assume.*
* Wilmott, P., S. Howison and J. Dewynne (1995). *The Mathematics of Financial Derivatives: A Student Introduction.* Cambridge University Press. *The heat-equation reduction of the Black-Scholes PDE, worked in full.*
