Chapter 2: Stochastic Convergence

Stochastic convergence is the language for separating a stable first-order limit from terms that vanish asymptotically. This chapter develops the parts of van der Vaart’s Chapter 2 that are used repeatedly on the route to semiparametric efficiency.

TipWhy this chapter is here
  • Prerequisites: elementary probability, expectations, and ordinary convergence of deterministic sequences.
  • Purpose: move limits through transformations, combine weak and probability limits, and control stochastic remainders.
  • Chapter 25 payoff: the same tools justify the asymptotically linear representations and negligible remainders used in Lemma 25.23. Before then, they support contiguity, LAN expansions, and function-valued convergence in Chapters 18 and 19.
NoteCompanion readings

For a systematic review of convergence and stochastic order, see §§2.1–2.7 and §3.4 of Jiang (2022). For a compact treatment centered on weak convergence, see §1.1 of DasGupta (2008).

NoteReading guide
  • Core — §2.1 through Lemma 2.8: Portmanteau, continuous mapping, the Euclidean Prohorov–Helly compactness step, relations among convergence modes, and Slutsky’s lemma.
  • Core — §2.2 and Lemma 2.12: stochastic \(o_P/O_P\) calculus and the conversion of deterministic remainder bounds into stochastic ones.
  • Compressed — Example 2.9: an illustration of weak convergence combined with a random normalizer in a \(t\) statistic.
  • Omitted/deferred — remaining Chapter 2 material: Example 2.10, Lemma 2.11, characteristic functions, almost-sure representations, moment convergence, convergence-determining classes, the law of the iterated logarithm, Lindeberg–Feller theory, and total variation. Chapter 18 supplies the metric-space generalization of tightness and Prohorov’s theorem needed later.

2.1 Basic Theory

For random vectors in \(\mathbb R^k\):

  • Convergence in distribution, also called weak convergence or convergence in law, is written \(X_n\rightsquigarrow X\). It means

    \[ P(X_n\leq x)\to P(X\leq x) \]

    at every continuity point of the distribution function of \(X\).

  • Convergence in probability is written \(X_n\overset P\to X\). It means

    \[ P\{d(X_n,X)>\epsilon\}\to0 \]

    for every \(\epsilon>0\).

  • Almost-sure convergence is written \(X_n\overset{\mathrm{a.s.}}\longrightarrow X\). It means \(d(X_n,X)\to0\) with probability one.

Convergence in probability and almost-sure convergence require \(X_n\) and \(X\) to live on the same probability space. Weak convergence depends only on their laws. For a careful measure-theoretic treatment of these convergence modes and the role of uniform integrability, see Pollard (2002).

ImportantLemma

Lemma 2.2: Portmanteau

For random vectors \(X_n\) and \(X\), the following statements are equivalent:

  1. \(P(X_n\leq x)\to P(X\leq x)\) at every continuity point of \(x\mapsto P(X\leq x)\);
  2. \(Ef(X_n)\to Ef(X)\) for every bounded continuous function \(f\);
  3. \(Ef(X_n)\to Ef(X)\) for every bounded Lipschitz function \(f\);
  4. \(\liminf_n Ef(X_n)\geq Ef(X)\) for every nonnegative continuous function \(f\);
  5. \(\liminf_nP(X_n\in G)\geq P(X\in G)\) for every open set \(G\);
  6. \(\limsup_nP(X_n\in F)\leq P(X\in F)\) for every closed set \(F\); and
  7. \(P(X_n\in B)\to P(X\in B)\) for every Borel set \(B\) with \(P(X\in\partial B)=0\).

Proof roadmap.

Approximate a bounded continuous function by step functions on continuity rectangles to obtain (1) \(\Rightarrow\) (2). The remaining implications pass between test functions and sets: Lipschitz approximations approach indicators of open sets, complements exchange open and closed sets, and the interior and closure squeeze continuity sets. Truncation supplies the equivalence of (2) and (4).

Complete proof

(1) \(\Rightarrow\) (2). First suppose the distribution function of \(X\) is continuous. Choose a compact rectangle \(I\) with \(P(X\notin I)<\epsilon\). Partition \(I\) into finitely many rectangles \(I_j\) so finely that a bounded continuous \(f\), rescaled if necessary to have \(|f|\leq1\), varies by at most \(\epsilon\) on each cell. With \(x_j\in I_j\), set

\[ f_\epsilon=\sum_j f(x_j)\mathbf 1_{I_j}. \]

Condition (1) gives \(P(X_n\in I_j)\to P(X\in I_j)\), so \(Ef_\epsilon(X_n)\to Ef_\epsilon(X)\). Tightness of the single vector \(X\), together with convergence on \(I\), gives \(P(X_n\notin I)<2\epsilon\) eventually. The two step-function approximation errors are therefore uniformly small, and \(Ef(X_n)\to Ef(X)\).

For a general distribution, use rectangles whose faces have zero \(X\)-probability. For each coordinate, only countably many levels can have positive probability, so such rectangle endpoints form a dense set. The same step-function argument then applies.

(2) \(\Rightarrow\) (3). Every bounded Lipschitz function is bounded and continuous.

(3) \(\Rightarrow\) (5). For open \(G\), define

\[ f_m(x)=m\,d(x,G^c)\wedge1. \]

Then \(f_m\) is bounded and Lipschitz, \(0\leq f_m\leq\mathbf 1_G\), and \(f_m\uparrow\mathbf 1_G\). Hence, for every fixed \(m\),

\[ \liminf_nP(X_n\in G) \geq\lim_nEf_m(X_n) =Ef_m(X). \]

Letting \(m\to\infty\) and using monotone convergence gives (5).

(5) \(\Leftrightarrow\) (6). Take complements, since a set is open exactly when its complement is closed.

(5) and (6) \(\Rightarrow\) (7). For any Borel set \(B\),

\[ P(X\in B^\circ) \leq\liminf_nP(X_n\in B) \leq\limsup_nP(X_n\in B) \leq P(X\in\overline B). \]

If \(P(X\in\partial B)=0\), the outside probabilities agree, proving (7).

(7) \(\Rightarrow\) (1). If \(x\) is a continuity point of the distribution function, the lower orthant \((-\infty,x]\) has boundary of \(X\)-probability zero. Apply (7).

It remains to connect (2) and (4). (2) \(\Rightarrow\) (4): for nonnegative continuous \(f\), each \(f\wedge M\) is bounded and continuous, so

\[ \liminf_nEf(X_n)\geq E\{f(X)\wedge M\}. \]

Let \(M\to\infty\). (4) \(\Rightarrow\) (2): if \(|f|\leq M\), apply (4) to the two nonnegative continuous functions \(M+f\) and \(M-f\). The first bounds the limiting expectation from below and the second bounds it from above. Thus \(Ef(X_n)\to Ef(X)\).

The source leaves this last equivalence as an exercise; the truncation argument above completes it here.

Portmanteau allows a proof to switch to whichever characterization best matches the problem. The open-set form is central in weak-convergence arguments; the test-function form drives the change-of-measure result in Theorem 6.6.

NoteTheorem

Theorem 2.3: continuous mapping

Let \(g:\mathbb R^k\to\mathbb R^m\) be continuous at every point of a set \(C\) such that \(P(X\in C)=1\). Then:

  1. if \(X_n\rightsquigarrow X\), then \(g(X_n)\rightsquigarrow g(X)\);
  2. if \(X_n\overset P\to X\), then \(g(X_n)\overset P\to g(X)\); and
  3. if \(X_n\overset{\mathrm{a.s.}}\longrightarrow X\), then \(g(X_n)\overset{\mathrm{a.s.}}\longrightarrow g(X)\).

Proof roadmap.

For weak convergence, pull a closed set back through \(g\) and apply Portmanteau. For convergence in probability, separate failure of continuity near the random limit from failure of \(X_n\) to approach \(X\). The almost-sure statement then follows pointwise.

Complete proof

For a closed set \(F\), continuity of \(g\) on \(C\) implies

\[ g^{-1}(F) \subseteq\overline{g^{-1}(F)} \subseteq g^{-1}(F)\cup C^c. \]

Indeed, if \(x_j\to x\in C\) and \(g(x_j)\in F\), then \(g(x_j)\to g(x)\) and closedness gives \(g(x)\in F\). If \(X_n\rightsquigarrow X\), Portmanteau therefore yields

\[ \begin{aligned} \limsup_nP\{g(X_n)\in F\} &\leq \limsup_nP\{X_n\in\overline{g^{-1}(F)}\}\\ &\leq P\{X\in\overline{g^{-1}(F)}\}\\ &\leq P\{g(X)\in F\}+P(X\notin C)\\ &=P\{g(X)\in F\}. \end{aligned} \]

Another application of Portmanteau proves the weak-convergence claim.

For convergence in probability, fix \(\epsilon>0\). For \(\delta>0\), let \(B_\delta\) contain those \(x\) for which some \(y\) satisfies \(d(x,y)<\delta\) but \(d\{g(x),g(y)\}>\epsilon\). Then

\[ P\bigl[d\{g(X_n),g(X)\}>\epsilon\bigr] \leq P(X\in B_\delta)+P\{d(X_n,X)\geq\delta\}. \]

For fixed \(\delta\), the second term tends to zero. As \(\delta\downarrow0\), the sets \(B_\delta\cap C\) decrease to the empty set by continuity, while \(P(X\in C)=1\); hence the first term tends to zero. This proves convergence in probability.

Finally, on the probability-one event where \(X_n\to X\) and \(X\in C\), ordinary continuity gives \(g(X_n)\to g(X)\). This proves the almost-sure claim.

The theorem passes a known limit through a statistic. Its discontinuity caveat is substantive: if \(X_n=1/n\), \(X=0\), and \(g(x)=\mathbf 1\{x>0\}\), then \(X_n\to0\) but \(g(X_n)=1\not\to g(0)=0\).

Tightness and subsequences

A sequence of random vectors is uniformly tight if, for every \(\varepsilon>0\), there is an \(M<\infty\) such that

\[ \sup_nP(\|X_n\|>M)<\varepsilon. \tag{2.4a} \]

Thus all laws place arbitrarily high probability in one common compact ball. In Euclidean spaces this is equivalent to boundedness in probability. The source spells the next result Prohorov’s theorem; Prokhorov is another common transliteration.

NoteTheorem

Theorem 2.4: Prohorov’s theorem

Let \(X_n\) be random vectors in \(\mathbb R^k\).

  1. If \(X_n\rightsquigarrow X\), then \((X_n)\) is uniformly tight.
  2. If \((X_n)\) is uniformly tight, then some subsequence \(X_{n_j}\) converges weakly to a random vector \(X\).

Proof roadmap. Portmanteau puts a weakly convergent sequence inside one large ball. Conversely, Helly’s lemma extracts a subsequential limit of the distribution functions; uniform tightness prevents that limit from losing probability mass at infinity.

Complete proof

Part 1. Fix \(\varepsilon>0\). Choose \(M\) so large that

\[ P(\|X\|\geq M)<\varepsilon. \]

The set \(\{x:\|x\|\geq M\}\) is closed, so Portmanteau gives

\[ \limsup_{n\to\infty}P(\|X_n\|\geq M) \leq P(\|X\|\geq M)<\varepsilon. \]

Hence the same tail probability is below \(2\varepsilon\) for all sufficiently large \(n\). Each of the finitely many remaining random vectors is tight; enlarging \(M\) once more makes their tail probabilities smaller than \(2\varepsilon\) as well. Since \(\varepsilon\) was arbitrary, the sequence is uniformly tight.

Part 2. Let \(F_n\) be the distribution function of \(X_n\). Lemma 2.5 supplies a subsequence \(F_{n_j}\) converging at continuity points to a possibly defective distribution function \(F\). It remains to show that \(F\) has total mass one.

Fix \(\varepsilon>0\). Uniform tightness gives \(M\) such that

\[ P(\|X_n\|\leq M)>1-\varepsilon \quad\text{for every }n. \]

Choose a continuity point \(x_M\) of \(F\) whose every coordinate exceeds \(M\); continuity points of a multivariate distribution function are dense from above. Then

\[ F(x_M) =\lim_{j\to\infty}F_{n_j}(x_M) \geq1-\varepsilon. \]

Letting \(M\to\infty\) and then \(\varepsilon\downarrow0\) shows that \(F(x)\to1\) as every coordinate of \(x\) tends to \(+\infty\). The same tightness bound applied to the event that any coordinate is below \(-M\) shows that \(F(x)\to0\) when any coordinate tends to \(-\infty\). Thus \(F\) is a proper distribution function. It is therefore the distribution function of some random vector \(X\), and convergence at all continuity points gives \(X_{n_j}\rightsquigarrow X\).

ImportantLemma

Lemma 2.5: Helly’s lemma

Every sequence \(F_n\) of distribution functions on \(\mathbb R^k\) has a subsequence \(F_{n_j}\) such that

\[ F_{n_j}(x)\longrightarrow F(x) \]

at every continuity point \(x\) of a possibly defective distribution function \(F\).

Proof roadmap. Diagonalize over the rational grid, extend the rational limits by right-continuity, prove convergence by rational squeezing, and pass nonnegative cell probabilities to the limit.

Complete proof

Enumerate \(\mathbb Q^k\) as \(q_1,q_2,\ldots\). Since \(0\leq F_n(q_1)\leq1\), choose a subsequence on which \(F_n(q_1)\) converges. From it choose a further subsequence on which \(F_n(q_2)\) converges, and continue. The diagonal subsequence, still denoted \(F_{n_j}\), then satisfies

\[ F_{n_j}(q)\longrightarrow G(q) \qquad\text{for every }q\in\mathbb Q^k. \]

Each \(F_n\) is coordinatewise nondecreasing, so \(G(q)\leq G(q')\) whenever \(q\leq q'\) coordinatewise. Define

\[ F(x)=\inf_{\substack{q\in\mathbb Q^k\\q>x}}G(q), \tag{2.5a} \]

where \(q>x\) means strict inequality in every coordinate. This definition makes \(F\) coordinatewise nondecreasing. It is right-continuous: given \(\varepsilon>0\), choose rational \(q>x\) with \(G(q)<F(x)+\varepsilon\). Whenever \(x\leq y<q\) coordinatewise,

\[ F(x)\leq F(y)\leq G(q)<F(x)+\varepsilon. \]

Now let \(x\) be a continuity point of \(F\). Rational points \(q<x<q'\) can be chosen sufficiently close to \(x\) that

\[ G(q')-G(q)<\varepsilon. \]

Monotonicity gives

\[ F_{n_j}(q)\leq F_{n_j}(x)\leq F_{n_j}(q'). \]

After taking lower and upper limits in \(j\),

\[ G(q)\leq\liminf_jF_{n_j}(x) \leq\limsup_jF_{n_j}(x)\leq G(q'). \]

The definition of \(F\) and the continuity of \(F\) at \(x\) place \(F(x)\) between the same two bounds. Since their gap is smaller than \(\varepsilon\), and \(\varepsilon\) is arbitrary, \(F_{n_j}(x)\to F(x)\).

For \(k=1\), monotonicity and right-continuity are all the finite-point requirements for a possibly defective distribution function. In higher dimensions one must also verify that every half-open rectangle has nonnegative \(F\)-mass. If all \(2^k\) corners of a rectangle are continuity points, its alternating corner sum is the limit of the corresponding nonnegative corner sums for \(F_{n_j}\) and is therefore nonnegative. For an arbitrary rectangle, move each face downward through continuity points and use right-continuity. Hence \(F\) is \(k\)-increasing and so is a possibly defective distribution function.

Example 2.6 gives a quick sufficient condition for (2.4a): if \(\sup_nE\|X_n\|^p<\infty\) for some \(p>0\), then Markov’s inequality yields

\[ \sup_nP(\|X_n\|>M) \leq M^{-p}\sup_nE\|X_n\|^p \longrightarrow0. \]

The compactness step reappears in Chapter 18, where closed balls need not be compact and tightness must be formulated with general compact sets.

NoteTheorem

Theorem 2.7: relations among modes of convergence

Let \(X_n\), \(X\), and \(Y_n\) be random vectors. Then:

  1. \(X_n\overset{\mathrm{a.s.}}\longrightarrow X\) implies \(X_n\overset P\to X\);
  2. \(X_n\overset P\to X\) implies \(X_n\rightsquigarrow X\);
  3. \(X_n\overset P\to c\) for a constant \(c\) if and only if \(X_n\rightsquigarrow c\);
  4. if \(X_n\rightsquigarrow X\) and \(d(X_n,Y_n)\overset P\to0\), then \(Y_n\rightsquigarrow X\);
  5. if \(X_n\rightsquigarrow X\) and \(Y_n\overset P\to c\) for a constant \(c\), then \((X_n,Y_n)\rightsquigarrow(X,c)\); and
  6. if \(X_n\overset P\to X\) and \(Y_n\overset P\to Y\), then \((X_n,Y_n)\overset P\to(X,Y)\).

Proof roadmap.

Almost-sure convergence controls the tail supremum and hence gives convergence in probability. A bounded-Lipschitz comparison shows that a probability-small perturbation does not change a weak limit. The other statements follow by specializing that comparison or by applying it jointly.

Complete proof

For (1), fix \(\epsilon>0\) and define

\[ A_n=\left\{\sup_{m\geq n}d(X_m,X)>\epsilon\right\}. \]

On the event of almost-sure convergence, \(A_n\) decreases to the empty set. Thus

\[ P\{d(X_n,X)>\epsilon\}\leq P(A_n)\to0. \]

For (4), let \(f\) take values in \([0,1]\) and have Lipschitz constant at most one. For every \(\epsilon>0\),

\[ |Ef(X_n)-Ef(Y_n)| \leq \epsilon+2P\{d(X_n,Y_n)>\epsilon\}. \]

The right side has limiting upper bound \(\epsilon\), and \(\epsilon\) is arbitrary. Hence \(Ef(X_n)\) and \(Ef(Y_n)\) have the same limit. Portmanteau and \(X_n\rightsquigarrow X\) give \(Y_n\rightsquigarrow X\).

For (2), apply (4) with the constant sequence equal to \(X\) and the other sequence equal to \(X_n\). For (3), the forward direction is (2). Conversely, if \(X_n\rightsquigarrow c\), Portmanteau applied to the closed complement of an open ball about \(c\) gives \(X_n\overset P\to c\).

For (5), \(d\{(X_n,Y_n),(X_n,c)\}=d(Y_n,c)\overset P\to0\). Moreover, for every bounded continuous \(f\), the map \(x\mapsto f(x,c)\) is bounded and continuous, so \((X_n,c)\rightsquigarrow(X,c)\). Apply (4).

For (6), use a product-space metric bounded by the sum of component distances:

\[ P\{d((X_n,Y_n),(X,Y))>\epsilon\} \leq P\{d(X_n,X)>\epsilon/2\}+P\{d(Y_n,Y)>\epsilon/2\}\to0. \]

Part 5 is the joint-convergence step behind Slutsky’s lemma and many later plug-in arguments.

ImportantLemma

Lemma 2.8: Slutsky

If \(X_n\rightsquigarrow X\) and \(Y_n\overset P\to c\) for a constant \(c\), then:

  1. \(X_n+Y_n\rightsquigarrow X+c\);
  2. \(Y_nX_n\rightsquigarrow cX\); and
  3. \(Y_n^{-1}X_n\rightsquigarrow c^{-1}X\), provided \(c\neq0\).

The statements also cover conformable matrices; in the third statement, \(c\neq0\) then means that \(c\) is invertible.

Proof roadmap.

First form the joint limit \((X_n,Y_n)\rightsquigarrow(X,c)\) using Theorem 2.7. Then apply the continuous mapping theorem to addition, multiplication, or inversion followed by multiplication.

Complete derivation in these notes

By Theorem 2.7,

\[ (X_n,Y_n)\rightsquigarrow(X,c). \]

The maps \((x,y)\mapsto x+y\) and \((x,y)\mapsto yx\) are continuous, so Theorem 2.3 gives the first two conclusions. If \(c\) is nonzero, or invertible in the matrix case, the map \((x,y)\mapsto y^{-1}x\) is continuous in a neighborhood of every point \((x,c)\). The continuous mapping theorem gives the third conclusion.

Van der Vaart states Slutsky’s lemma as an immediate consequence of the two preceding results; the derivation is made explicit here.

The recurring asymptotic pattern is

\[ T_n=A_n+R_n,\qquad A_n\rightsquigarrow A,\qquad R_n=o_P(1). \]

Slutsky removes \(R_n\). The LAN remainder in Theorem 7.2, the empirical-process remainder in Theorem 19.26, and the efficient-estimator expansion in Lemma 25.23 all have this structure.

TipExample

Example 2.9: the \(t\) statistic

Let \(Y_1,Y_2,\ldots\) be i.i.d. with \(EY_1=0\) and \(EY_1^2<\infty\), and let

\[ S_n^2=\frac{1}{n-1}\sum_{i=1}^n(Y_i-\overline Y_n)^2. \]

The weak law and continuous mapping give \(S_n^2\overset P\to\operatorname{var}(Y_1)\) and hence \(S_n\overset P\to\operatorname{sd}(Y_1)\). The central limit theorem gives \(\sqrt n\,\overline Y_n\rightsquigarrow N\{0,\operatorname{var}(Y_1)\}\). Slutsky then yields

\[ \frac{\sqrt n\,\overline Y_n}{S_n}\rightsquigarrow N(0,1). \]

The denominator is a random normalizer. Its probability limit is enough to replace it asymptotically by the fixed standard deviation.

2.2 Stochastic \(o\) and \(O\) Symbols

The source defines stochastic order relative to a possibly random sequence \(R_n\) by factorization:

\[ \begin{aligned} X_n=o_P(R_n) &\quad\text{means}\quad X_n=Y_nR_n\ \text{for some }Y_n\overset P\to0,\\ X_n=O_P(R_n) &\quad\text{means}\quad X_n=Y_nR_n\ \text{for some }Y_n=O_P(1). \end{aligned} \]

Here \(O_P(1)\) means bounded in probability. This factorization remains meaningful when \(R_n\) is random or can equal zero. When \(R_n\) is strictly positive and division is valid, it reduces to the familiar shorthand \(X_n/R_n=o_P(1)\) or \(O_P(1)\).

For deterministic \(a_n>0\), a root-\(n\) consistent estimator satisfies

\[ \widehat\theta_n-\theta=O_P(n^{-1/2}). \]

Frequently used rules include

\[ \begin{gathered} o_P(1)+o_P(1)=o_P(1),\qquad o_P(1)+O_P(1)=O_P(1),\\ O_P(1)o_P(1)=o_P(1),\qquad \{1+o_P(1)\}^{-1}=O_P(1),\\ o_P(R_n)=R_no_P(1),\qquad O_P(R_n)=R_nO_P(1),\qquad o_P\{O_P(1)\}=o_P(1). \end{gathered} \]

These displays are implications read from left to right, not algebraic identities. Each occurrence of \(o_P(1)\) or \(O_P(1)\) may denote a different sequence. In particular, \(O_P(1)o_P(1)=o_P(1)\) says that the product of a bounded-in-probability sequence and a sequence converging to zero in probability also converges to zero in probability.

The product rule is especially important in semiparametric problems. Two nuisance errors may each be larger than \(n^{-1/2}\) while their product is \(o_P(n^{-1/2})\), so the product disappears from the target estimator’s first-order expansion; see the no-bias condition in §25.8.

ImportantLemma

Lemma 2.12: deterministic remainders at random arguments

Let \(R\) be defined on a subset of \(\mathbb R^k\), with \(R(0)=0\), and let \(X_n\) take values in that domain with \(X_n\overset P\to0\). For every \(p>0\):

  1. if \(R(h)=o(\|h\|^p)\) as \(h\to0\), then \(R(X_n)=o_P(\|X_n\|^p)\); and
  2. if \(R(h)=O(\|h\|^p)\) as \(h\to0\), then \(R(X_n)=O_P(\|X_n\|^p)\).

Proof roadmap.

Factor the remainder as \(R(h)=g(h)\|h\|^p\). A deterministic little-\(o\) assumption makes \(g\) continuous at zero; a deterministic big-\(O\) assumption makes \(g\) locally bounded. Continuous mapping or tightness then gives the corresponding stochastic order.

Complete proof

Define

\[ g(h)= \begin{cases} R(h)/\|h\|^p,&h\neq0,\\ 0,&h=0. \end{cases} \]

Then \(R(X_n)=g(X_n)\|X_n\|^p\).

Under the first assumption, \(g(h)\to0=g(0)\) as \(h\to0\). The continuous mapping theorem gives \(g(X_n)\overset P\to0\), proving

\[ R(X_n)=o_P(\|X_n\|^p). \]

Under the second assumption, there are finite \(M\) and \(\delta>0\) such that \(|g(h)|\leq M\) whenever \(\|h\|\leq\delta\). Therefore

\[ P\{|g(X_n)|>M\}\leq P(\|X_n\|>\delta)\to0. \]

Thus \(g(X_n)=O_P(1)\) and

\[ R(X_n)=O_P(\|X_n\|^p). \]

Lemma 2.12 is the bridge from ordinary differentiability to stochastic expansions. It supplies the remainder step in the delta method in Chapter 3, and the same mechanism reappears in likelihood and estimating-equation expansions throughout Chapters 7, 19, and 25.

Weak convergence controls bounded continuous test functions, but it does not by itself justify taking limits of expectations or risks. A sequence of integrable random variables is uniformly integrable if

\[ \lim_{M\to\infty} \sup_nE\bigl[|X_n|\mathbf 1\{|X_n|>M\}\bigr]=0. \]

If \(X_n\rightsquigarrow X\) and \((X_n)\) is uniformly integrable, then \(X\) is integrable and \(EX_n\to EX\). A convenient sufficient condition is

\[ \sup_nE|X_n|^{1+\delta}<\infty \]

for some \(\delta>0\). This extra tail control is what turns a distributional efficiency statement into convergence of mean loss when the loss is unbounded. It is optional for the first-order Chapter 25 route, whose principal conclusions are weak limits and local risk lower bounds stated with their own regularity conditions.

TipChapter takeaway

Weak convergence, continuous mapping, Slutsky’s lemma, and stochastic-order calculus separate stable first-order limits from negligible remainders. These tools underwrite the local expansions used throughout Chapters 6–8, 18–19, and 25.