How statement and proof provenance work
The first chip identifies the source of the statement or construction; the second identifies the source of its local proof or verification.
- Literature-sourced: the exact statement appears in a cited source; only wording and notation differ.
- AI-adapted: a semantically identical restatement of literature-sourced material, modulo indexing, notation, and boundary cases adopted by the library.
- AI-generated: a genuinely novel statement formulated by AI, with no source for the claim itself.
These labels describe origin, not correctness: citations and verification chips remain separate evidence.
Central Limit Theorems
1 · Prerequisites
- Absolute and Conditional Convergence; Rearrangement; Products
- Algebraic Closure, Embeddings, and Separability
- Algebraic Extensions, Extension Degree, and Finite Fields
- Areas of Elementary Plane Figures
- Binary Operations, Monoids, Groups and Subgroups
- Characteristic Functions Inversion and Continuity
- Compactness
- Compactness in Metric Spaces
- Complete Metrizability, Čech-Completeness, and Baire Category
- Completeness, Completion, and Uniform Continuity
- Complex Lp Spaces and Test-Function Conventions
- Composition Series, the Jordan–Hölder Theorem and Solvable Groups
- Congruences, the Integers Modulo n and the Chinese Remainder Theorem
- Construction of the Natural Numbers
- Construction of the Real Numbers via Cauchy Sequences
- Construction of the Real Numbers via Dedekind Cuts
- Continuity, IVT, EVT, and Uniform Continuity
- Cosets, Index and Lagrange's Theorem
- Countability and Uncountability
- Countability Axioms and Cardinal Functions
- Cyclic Groups and Direct Products
- Darboux, L'Hôpital, and Taylor's Theorem
- Density Separability and Convolution in Lᵖ
- Determinants of Matrices over a Commutative Ring
- Diagonalisation and the Minimal Polynomial
- Divisibility, Euclidean Domains, Principal Ideal Domains and Unique Factorisation
- Divisibility, Greatest Common Divisors and Bézout's Identity
- Dual Spaces, Bilinear and Quadratic Forms, and Sylvester's Law of Inertia
- Eigenvalues, Eigenvectors and the Characteristic Polynomial
- Filters and Ultrafilters
- Finite Counting, Factorials and Binomial Coefficients
- Finite Probability and the Probabilistic Method
- Finite Probability Spaces and Random Variables
- Foundations of the Real Numbers for Analysis
- Fourier Transform Convolution and Approximate Identities
- Fubini and Change of Variables
- Fundamental Trigonometric Identities
- Gaussian Elimination, Elementary Matrices and Reduced Row Echelon Form
- Group Actions, Orbits, Stabilisers and Cayley's Theorem
- Group Homomorphisms and the Isomorphism Theorems
- Hausdorff via the Diagonal
- Ideals, Quotient Rings and the Isomorphism Theorems for Rings
- Improper and Parameter-Dependent Multiple Integrals
- Improper Integrals
- Independence Borel Cantelli and Zero One Laws
- Infinite Product Measures and Kolmogorov Extension
- Inner Product Spaces, Gram-Schmidt, Projections and Adjoints
- Lebesgue Measure on Euclidean Space
- Lebesgue-Stieltjes Measures and Distribution Functions
- Limits of Real Functions
- limsup, liminf, and Subsequential Limits
- Linear Independence, Bases and Dimension
- Linear Transformations, Rank-Nullity and Quotient Spaces
- Matrices, the Matrix of a Linear Map, and Change of Basis
- Measurable Functions and Simple Approximation
- Measures and Their Basic Properties
- Metric Spaces
- Mixed Partials, Taylor Formulae, and Extrema
- Modes of Convergence Egorov and Lusin
- Modes of Convergence for Random Variables
- Monotone Functions, Discontinuities, and Continuity Sets
- Monotone Sequences, Bolzano-Weierstrass, and Cauchy Completeness
- Normal Subgroups and Quotient Groups
- Order, Zorn's Lemma, and the Axiom of Choice
- Outer Measure and the Caratheodory Extension Theorem
- Partitions of Unity and Paracompactness
- Polynomial Rings, the Division Algorithm and Roots
- Power Series and Real-Analytic Functions
- Primes, Euclid's Lemma and the Fundamental Theorem of Arithmetic
- Probability Spaces Random Variables and Expectation
- Product Measures and the Fubini Tonelli Theorems
- Properties of the Integral and the Working FTC
- Radon Measures and the Riesz Markov Kakutani Theorem
- Relations, Functions, and Quotients
- Rings, Subrings, Integral Domains and Fields
- Rⁿ as a Normed Space; Vector-Valued Functions
- Roots, Rational Powers, and Classical Inequalities
- Separation Axioms: the Hierarchy
- Sequences and Limits
- Sequences and Series of Functions; Uniform Convergence
- Series: Convergence and the Nonnegative Tests
- Sigma Algebras and Borel Sets
- Signed and Complex Measures Hahn and Jordan
- Simple Field Extensions and the Construction of the Complex Numbers
- Sine, Cosine, and the Definition of Pi
- Splitting Fields
- Subspaces, Products, and Quotients
- Suprema and Infima
- Sylow's Theorems, p-Groups and Nilpotent Groups
- Symmetric Groups, Cycle Decomposition and the Sign Homomorphism
- The Cantor Set, Baire Category, and Measure Zero in ℝ
- The Complex Exponential and Euler's Formula
- The Derivative and the Mean Value Theorems
- The Determinant of a Linear Operator, Cofactors and Cramer's Rule
- The Exponential Function
- The Fundamental Theorem of Algebra
- The Fundamental Theorem of Finite Abelian Groups
- The Galois Correspondence
- The Inverse and Implicit Function Theorems
- The Lebesgue and Riemann Integrals Compared
- The Lebesgue Integral and the Convergence Theorems
- The Logarithm and General Powers
- The Lᵖ Spaces Holder Minkowski and Riesz Fischer
- The Maximal Function and Lebesgue Differentiation
- The Radon Nikodym Theorem and Lebesgue Decomposition
- The Riemann Integral in Rᵐ and Jordan Content
- The Riemann Integral: Definition and Integrability
- The Spectral Theorem, Positive Operators and Singular Value Decomposition
- The Topology of Euclidean Space
- The Total Derivative in ℝᵐ → ℝⁿ
- The ZFC Axioms and the Basic Set Constructions
- Topological Spaces and Continuity
- Topology of ℝ
- Triangularisation, Generalised Eigenspaces and Jordan Canonical Form
- Urysohn's Lemma and the Tietze Extension Theorem
- Vector Spaces, Linear Subspaces, Span and Direct Sums
- Weak Convergence Tightness and Representation
- Weak Laws and Series of Independent Random Variables
2 · Summary
Central limit theorems describe the laws of normalized sums. The scalar argument begins with an explicit normal transform and a second-order expansion requiring only a finite second moment. A finite-product estimate controls the accumulated error when the number of factors grows.
A local AC-to-countable-choice and dependent-choice lemma supplies the choices used in probability-product constructions. The iid theorem leads to de Moivre–Laplace. Triangular arrays then allow the distributions to change within each row. Lindeberg's tail condition implies that no individual variance dominates and gives a Gaussian limit. Conversely, when maximal summand variance tends to zero and total variance is one, a normal limit forces Lindeberg: the proof uses a nonnegative cosine deficit and a fixed frequency chosen for each tail threshold. Lyapunov's higher-moment condition is a convenient sufficient condition.
Multivariate normal laws are constructed by a positive semidefinite matrix square root. Their projection characterization includes singular covariance matrices, so the multivariate iid theorem separates positive-variance projections from those that vanish almost surely. Cramer–Wold then combines these scalar conclusions.
The results state their choice assumptions. AC covers the published Gaussian, independent-copy, uniqueness and weak-convergence machinery where used; the scalar remainder and finite-product arguments are choice-free. The final remark distinguishes convergence of laws from convergence in probability or almost surely on a common space. No quantitative approximation rate is asserted.
3 · Logical flowchart
4 · Definitions, theorems and proofs
Characteristic function of a normal law
Statement
Assume AC. If has law with , then for every real . Moreover and , including .
Facts & Assumptions
Under AC the normal laws are the affine images of the standard density law. Standard normal and normal laws.
The positive Borel density g has total integral one. The standard normal density has total mass one.
The real exponential is smooth with derivative itself. The exponential function is smooth and .
The chain rule differentiates the quadratic composition. The chain rule, in one line from Carathéodory: if is differentiable at and is differentiable at , then is differentiable at with .
The product and linearity rules apply to real components. Sums, scalar multiples, products and quotients: , , , and when .
Continuous real integrands are Riemann integrable on compact intervals. A continuous function on is Riemann integrable, by Heine-Cantor and Riemann's criterion.
Integrable derivatives integrate to their endpoint differences. The second fundamental theorem: if is differentiable on with and is integrable, then .
Integration by parts holds for continuously differentiable real functions on compact intervals. If are differentiable on with integrable, then .
Under countable choice the compact Riemann and Lebesgue integrals agree. A bounded Riemann integrable function on a closed bounded interval is Lebesgue measurable and has the same integral.
Increasing nonnegative truncations converge in integral. Monotone convergence for the integral.
Integration against a density agrees with integration of the product for nonnegative measurable integrands. Integrating against a density agrees with integrating the product.
Finite first absolute moment permits differentiation of the characteristic function. Moments give derivatives of the characteristic function.
An integrable absolute majorant permits complex integral limits. Dominated convergence.
Euler form has unit modulus on imaginary arguments. , , and .
The sine and cosine derivatives justify real-component integration by parts. The derivatives of sine and cosine are cosine and minus sine.
A real function with zero derivative on an interval is constant. A function continuous on an interval whose derivative vanishes at every interior point of is constant on ; consequently two such functions with the same derivative differ by a constant.
Affine changes give the stated characteristic-function transformation. Characteristic functions under affine maps and independent sums.
The exponential addition law holds, and the complex exponential extends the real exponential. , and the complex exponential extends the real exponential.
Real integrals are defined through positive and negative parts, and complex integrals through real and imaginary parts. Integrable real and complex functions, and their integrals.
The real exponential is defined by the everywhere convergent series . The real exponential function and the number by a power series.
Proof
Given: Assume AC. If has law with , then for every real . Moreover and , including .
Write and let be the coordinate under its probability law. By [F1]–[F2], this is a probability law with density . For any measurable complex with , apply [F11] to the positive and negative parts of and . The definitions in [F19] then give ; in particular, every density-integral identity used below is covered. The derivative rules give . All compact-interval functions below are continuously differentiable, so [F6]–[F9] apply also to each real and imaginary component. AC is used through the normal-law construction and the countable-choice Riemann/Lebesgue bridge; no sequence of arbitrary witnesses is selected.
For , the series in [F20] has nonnegative terms at , so . By [F18], ; hence and both and tend to zero. For each positive integer , FTC on each half interval gives . Thus MCT proves . Also ; DCT with majorant proves . Integration by parts with gives . MCT and [F2] now give . All truncations here use the explicit integers R.
By the finite first moment and [F12], . On , componentwise integration by parts, using [F14]–[F15], gives . The boundary has modulus at most . The two integrands are dominated respectively by and , already integrable. DCT therefore gives for every t, with no improper differentiation left unjustified.
By the real-component product and chain rules, has derivative zero. Apply [F16] to its real and imaginary parts on every real interval. Since , for all t, hence . For in law, [F17] yields , and [F18] combines the exponents. Its moments follow by expanding the finite integrable expressions: and . For the variable equals m almost surely and both formulas give the Dirac law directly.
Second-order characteristic-function expansion
Statement
If and , then as . No third moment or choice axiom is assumed. The same scalar estimates give for every real u. More precisely, for ,
Facts & Assumptions
Real Taylor remainders are bounded by the uniform next derivative bound. A uniform derivative bound gives a uniform Taylor remainder bound.
Sine and cosine have derivatives of all orders bounded by one. The derivatives of sine and cosine are cosine and minus sine.
Euler form has real cosine and imaginary sine components. , , and .
DCT applies to the prescribed nonnegative majorant sequence below. Dominated convergence.
Integrable real and complex linear combinations commute with integration. The Lebesgue integral is linear on .
The characteristic function is the expectation of the unit exponential. Characteristic function of a real random variable.
Proof
Given: If and , then as . No third moment or choice axiom is assumed. The same scalar estimates give for every real u. More precisely, for ,
Taylor at zero through degree two for cosine and sine, with third derivatives bounded by one, gives and . Euler form and the triangle inequality give . Taylor through degree one gives and , hence . Adding bounds for all real u. At u=0 all remainders vanish.
For , the first step gives . This is a measurable nonnegative sequence tending pointwise to zero, bounded by the integrable . DCT therefore makes its expectations tend to zero. The bound is uniform over all such t, so as a genuine two-sided real limit, without choosing a sequence of frequencies or requiring . Also gives integrability of the linear term.
By linearity, . Insert the stated moments and the preceding remainder limit. If , the same majorants have zero integral, so the remainder vanishes and the formula still holds. At t=0 the defining expectation is one. Every limiting integrand was explicitly specified; no choice principle enters.
Products of near-one characteristic factors
Statement
Let each row be a finite family of complex numbers; empty rows are allowed with maximum zero. Suppose , , and . Then
Facts & Assumptions
The exponential is its power series. The complex exponential by its power series.
The series converges absolutely at every complex argument. The complex exponential series converges absolutely for every complex argument.
Finite products of exponentials equal the exponential of their sum. , and the complex exponential extends the real exponential.
Proof
Given: Let each row be a finite family of complex numbers; empty rows are allowed with maximum zero. Suppose , , and . Then
Put . For , absolute convergence gives . Also and . The real series is nonnegative term by term, which proves the middle inequality without a logarithm.
For arbitrary finite lists a,b of length r, subtracting successive mixed products gives . The cancellation follows by expanding each difference; for r=0 both products are one and the sum zero. Apply it with and . For all sufficiently large n every modulus of w is at most 1/2. Each pair of partial products is bounded by , so the difference of the row products has modulus at most .
The bound tends to zero by hypothesis. Repeated exponential addition identifies the comparison product with , including empty and one-factor rows. If M=0 then all w vanish and equality is exact in every row. All products and telescoping sums are finite; no branch of logarithm or choice principle is used.
Lindeberg-Levy iid central limit theorem
Statement
Assume AC. Let be iid real random variables with mean and variance . Then
Facts & Assumptions
Centered variance-one variables have characteristic function 1-t^2/2+o(t^2). Second-order characteristic-function expansion.
Near-one rows with bounded absolute sum and vanishing square sum admit exponential comparison. Products of near-one characteristic factors.
Affine transformations and finite independent sums have the stated transform identities. Characteristic functions under affine maps and independent sums.
Under AC the standard normal has transform exp(-t^2/2). Characteristic function of a normal law.
Under AC convergence to the transform of a specified law gives weak convergence. Characteristic function criterion for weak convergence.
Linearity permits centering and variance calculations. The Lebesgue integral is linear on .
Proof
Given: Assume AC. Let be iid real random variables with mean and variance . Then
Set . The affine functions are Borel; independence is preserved because an event concerning Y_k is the corresponding inverse-image event concerning X_k. Their laws agree, and linearity gives , . The normalized sum in the statement is . The positive finite variance makes every division well defined.
Fix a nonzero real t and write . By [F1], . Therefore , is bounded, and . The bounds extend to the finitely many early n since their values are finite. Apply [F2] to the row with n copies of . Its product differs from by a quantity tending to zero. Continuity of the exponential yields the limit . At t=0 each factor is exactly one.
By [F3], row independence identifies that product with the characteristic function of the normalized sum. By [F4] its limit is the standard-normal transform, continuous at zero and equal to one there. Thus [F5] proves the claimed convergence of laws. AC is inherited only through the normal-law and Levy-criterion suppliers; the iid sequence is given and no new copies are constructed. This theorem excludes sigma=0 because its displayed normalization divides by sigma.
AC supplies countable selections and prescribed serial paths
Statement
Assume AC. Then countable choice holds. Moreover, if is serial on a nonempty set X and , there is a sequence with and for every n. Thus the countable and dependent choices required by the probability-product suppliers are available under AC.
Facts & Assumptions
AC supplies a choice function on a set of nonempty sets. The Axiom of Choice.
Countable choice is selection from an omega-indexed nonempty family. The Axiom of Countable Choice ().
DC requires a serial path starting at a prescribed point. The axiom of dependent choice: a relation in which every element is related to something admits an -indexed chain.
A given self-map and initial point have a uniquely specified natural-number iterate sequence. The recursion theorem.
Proof
Given: Assume AC. Then countable choice holds. Moreover, if is serial on a nonempty set X and , there is a sequence with and for every n. Thus the countable and dependent choices required by the probability-product suppliers are available under AC.
For an omega-indexed family of nonempty sets, its image is a set of nonempty sets. By [F1] choose with for every . Then is a function and , which is [F2]. Repeated sets use the same selected value and cause no ambiguity. For an empty index family the empty function already suffices; singleton fibers force their unique value.
For the stated serial R, every fiber is nonempty. AC applied to the set of all these fibers gives c with . Define the self-map on X. No recursively changing choice is being assumed: s is now one fixed function chosen from a fixed set of nonempty fibers. It satisfies for every x.
Apply [F4] to X, the prescribed a and s to obtain and for every natural n. The defining property of s gives , precisely [F3]. If X has one element and R is serial, this is its constant sequence. Empty X is excluded by the prescribed a. The uses of AC are exactly the two fixed-family selections in steps 1.1 and 1.2; recursion itself requires no further choice. This proves the two implications from AC, not either converse.
De Moivre-Laplace central limit theorem
Statement
Assume AC and fix . If has law for every , then No relationship between the probability spaces of the is required.
Facts & Assumptions
Under dependent and countable choice a prescribed Bernoulli law has independent copies. Countably many independent copies of a prescribed law exist.
AC implies dependent choice and countable choice. AC supplies countable selections and prescribed serial paths.
Bernoulli(p) has mean p and variance p(1-p). A Bernoulli variable has mean and variance ; a binomial variable has mean and variance .
The iid finite-positive-variance CLT applies under AC. Lindeberg-Levy iid central limit theorem.
A binomial law is the law of a finite independent Bernoulli sum. Bernoulli random variables and binomial random variables as sums of independent Bernoulli trials.
Proof
Given: Assume AC and fix . If has law for every , then No relationship between the probability spaces of the is required.
By [F2], AC supplies the countable and dependent choice required in [F1]. Apply that result to the two-point probability with masses 1-p and p to obtain iid Bernoulli variables . The law of is Bin(n,p) by [F5]. In particular it agrees with the specified law of B_n for each n; equality persists under the displayed affine standardization.
By [F3], and . The variables are bounded, hence their second moments are finite. [F4] gives . Equality of laws in step 1.1 transfers the conclusion to B_n. The excluded p=0,1 and n=0 would make the denominator zero; no claim using that denominator is made there. Neither a finite-n error estimate nor continuity correction follows from this limit theorem.
Row-wise independent centered triangular array
Definition
For each integer , let be a finite integer and let be integrable real random variables on a probability space . The family is a centered triangular array when for every admissible pair . It is row-wise independent when, for each fixed n and all Borel sets , Taking unused equal to the real line gives the same factorization for each subfamily; by Independent random elements are characterized by finite rectangle probabilities, this is exactly independence in Independent random elements. Random variables and integrability have the meanings in Random elements and real random variables and Expectation of a nonnegative or integrable random variable. There is no independence requirement between different rows. The spaces may differ with n; assertions about row sums compare their laws. A single-entry row is independent automatically, and deterministic zero entries are allowed. Empty rows are excluded by . This definition makes no existence or choice assumption.
Total row variance and the Lindeberg condition
Definition
For a centered triangular array as in Row-wise independent centered triangular array, suppose every entry has finite second moment. Write Variance is defined in Moments, variance, and covariance on a probability space. Every term is finite and nonnegative, so the finite sum and its nonnegative square root s_n exist. Whenever , define The Lindeberg condition is for every fixed , with for all n under consideration. The event is measurable because absolute value is continuous and the entries are measurable; products are measurable by Arithmetic and lattice operations preserve measurability whenever they are defined. Its nonnegative integrand is bounded by , so its expectation exists and is finite by Monotonicity and nonnegative homogeneity of the nonnegative integral and Expectation of a nonnegative or integrable random variable. In particular .
For , scalar homogeneity gives . The equality gives Thus the normalized and unnormalized conditions are exactly equivalent, not merely asymptotic. The cutoff uses strict inequality; equality at the threshold is excluded. A zero-variance row is allowed in the initial variance definition, but its normalization and Lindeberg expression above are undefined. No independence or choice is needed to define these quantities.
The Lindeberg condition implies Feller negligibility
Statement
For a centered triangular array with , the Lindeberg condition implies Feller negligibility: . Independence is not needed.
Facts & Assumptions
In normalized rows the Lindeberg quantity is the sum of the truncated second moments. Total row variance and the Lindeberg condition.
An integrable function splits into its two complementary restrictions. The Lebesgue integral is linear on .
Nonnegative integrals preserve pointwise inequalities. Monotonicity and nonnegative homogeneity of the nonnegative integral.
Proof
Given: For a centered triangular array with , the Lindeberg condition implies Feller negligibility: . Independence is not needed.
Fix . For each entry split its square on and its complement. The first integral is at most . The second is one nonnegative term of . Centering identifies variance with second moment, hence . The maximum exists because the row is finite and nonempty.
For any , take . Lindeberg gives an index after which . The preceding bound is then less than eta, proving convergence to zero. Values at the cutoff stay in the small part, zero entries satisfy the bound, and a single variance-one summand in every row would contradict this conclusion and hence could not satisfy Lindeberg. The proof uses only given finite rows and explicit inequalities, not independence or choice.
Lindeberg-Feller central limit theorem: sufficiency
Statement
Assume AC. If a centered row-wise independent triangular array has finite second moments, and the Lindeberg condition, then
Facts & Assumptions
Normalization makes total variance one and preserves the Lindeberg quantity. Total row variance and the Lindeberg condition.
The scalar remainder obeys min(|u|^3/3,4u^2), and the centered exponential increment is bounded by u^2. Second-order characteristic-function expansion.
Lindeberg implies maximal entry variance tends to zero. The Lindeberg condition implies Feller negligibility.
The near-one product estimate controls the full growing row. Products of near-one characteristic factors.
Finite row independence identifies the product transform. Characteristic functions under affine maps and independent sums.
The standard-normal transform is exp(-t^2/2). Characteristic function of a normal law.
Under AC pointwise convergence to a specified characteristic function implies weak convergence. Characteristic function criterion for weak convergence.
Finite sums and centered integrable terms may be integrated linearly. The Lebesgue integral is linear on .
The modulus of a complex integral is bounded by the integral of the modulus. The modulus of an integral is bounded by the integral of the modulus.
Proof
Given: Assume AC. If a centered row-wise independent triangular array has finite second moments, and the Lindeberg condition, then
Put and . By [F1], the normalized row has total variance one, remains centered and independent, and its tail sum tends to zero. Fix real t and write . The centering and scalar bound in [F2] give , using [F8]–[F9]. Consequently , , and by [F3].
With r as in [F2], linearity gives . On the cubic bound gives ; on the complement the quadratic bound gives . Therefore . First take limsup in n, then let the arbitrary positive epsilon decrease to zero. This proves , without interchanging an unbounded number of unquantified little-o terms.
All hypotheses of [F4] were verified in step 1.1, so . Step 2.1 makes its limit . By [F5] this product is the row-sum transform, and by [F6] its limit is the transform of N(0,1), continuous at zero. [F7] proves the result. At t=0 all factors and the limit are exactly one. AC is inherited in [F6]–[F7]; rows on different probability spaces cause no difficulty because only their laws are compared.
Feller converse to Lindeberg-Feller
Statement
Assume AC. Let a centered row-wise independent triangular array satisfy and . If , then it satisfies the Lindeberg condition.
Facts & Assumptions
The scalar centered exponential increment has modulus at most u^2; its proof also gives 1-cos(u)<=u^2/2. Second-order characteristic-function expansion.
Near-one products differ from the exponential of their summed increments by o(1). Products of near-one characteristic factors.
The independent row sum has product characteristic function. Characteristic functions under affine maps and independent sums.
The assumed weak convergence gives pointwise characteristic-function convergence. Levy continuity theorem forward direction.
Under AC N(0,1) has transform exp(-t^2/2). Characteristic function of a normal law.
Euler form gives unit modulus and -1<=cos<=1. , , and .
Exponential addition and real extension identify the modulus as exp(real part). , and the complex exponential extends the real exponential.
The real nonnegative exponential series gives exp(d)>=1+d for d>=0. The complex exponential by its power series.
Centering, real parts and finite sums commute with integration. The Lebesgue integral is linear on .
Complex expectation is bounded by the expectation of its absolute value. The modulus of an integral is bounded by the integral of the modulus.
For total row variance one, Lindeberg is exactly convergence of the summed tail second moments. Total row variance and the Lindeberg condition.
Proof
Given: Assume AC. Let a centered row-wise independent triangular array satisfy and . If , then it satisfies the Lindeberg condition.
Fix real t and put and . Centering, [F1], [F9] and [F10] give . Thus , , and . By [F2]–[F5] and the assumed normal limit, . This inference uses the forward continuity theorem only, not Lindeberg sufficiency.
Define the nonnegative function . Nonnegativity follows from the cosine Taylor bound in [F1]; also because . Hence is finite and nonnegative. Total variance one and [F9] give . Exponential addition and Euler modulus show . Step 1.1 therefore implies . Since the nonnegative series gives , we obtain . No complex logarithm or subsequence of measures is needed.
Now fix any and take the single frequency . On we have , and yields . On the complementary set d_t is nonnegative. Integrating and summing gives . This is precisely [F11], for every positive epsilon. AC is inherited from the target normal-law construction; the proof uses no Helly selection, uniqueness inversion or backward application of sufficiency. Zero entries and t=0 in the earlier steps are harmless, but the final chosen t is nonzero.
Lyapunov central limit theorem
Statement
Assume AC. Let a centered row-wise independent triangular array have finite second moments and . If for some its moments are finite and then .
Facts & Assumptions
The unnormalized tail expression defines Lindeberg. Total row variance and the Lindeberg condition.
Under AC Lindeberg implies a standard-normal limit for the normalized row sums. Lindeberg-Feller central limit theorem: sufficiency.
Pointwise bounds pass to nonnegative expectations. Monotonicity and nonnegative homogeneity of the nonnegative integral.
Positive real powers obey product and exponent laws. The exponent, product, quotient, and iterated-power laws for positive real bases and real exponents.
For positive a, a^r=exp(r log a); zero to a positive power is zero. Real powers for positive bases, with the zero-base positive-exponent convention.
Proof
Given: Assume AC. Let a centered row-wise independent triangular array have finite second moments and . If for some its moments are finite and then .
Fix . On , monotonicity of the positive real power gives . Multiplying by gives there. Off that event the truncated square is zero and the right side is nonnegative, including X=0. [F3] therefore bounds . The exponent simplification uses [F4]. For positive delta, monotonicity of log and exp in the defining formula [F5] gives the asserted power monotonicity.
The right side tends to zero for this fixed positive epsilon, so Lindeberg holds for every epsilon. All hypotheses of [F2] are now satisfied: finite second moments, centered independent rows and positive total standard deviations were given. Apply it to obtain the stated limit. The positive delta is fixed across all rows; delta=0 is excluded because the displayed normalized second-moment sum would be one. AC is inherited from [F2].
Multivariate normal law, including singular covariance
Definition
Assume AC and let be finite. For and a real symmetric positive semidefinite matrix , a Borel probability law is denoted when a vector X with that law satisfies for every . Such a law exists, has mean m and covariance Sigma, and can be realized as with independent standard-normal coordinates. Singular Sigma is allowed. Uniqueness will be proved in the following characteristic-function lemma.
Facts & Assumptions
Normal laws have their specified means, variances and characteristic functions. Characteristic function of a normal law.
A finite-dimensional positive semidefinite symmetric operator has a positive semidefinite square root. A non-negative operator has a unique non-negative square root.
Independent standard-normal coordinates can be realized under DC and countable choice. Countably many independent copies of a prescribed law exist.
AC supplies DC and countable choice. AC supplies countable selections and prescribed serial paths.
Finite independent linear combinations have product transforms. Characteristic functions under affine maps and independent sums.
Under AC equality of scalar characteristic functions determines scalar laws. Uniqueness of a law from its characteristic function.
Integrable products in distinct independent coordinates factor. Expectations factor over finite products of independent random variables.
Finite linear combinations commute with expectations. The Lebesgue integral is linear on .
A law is the probability pushforward of a measurable random element. Law or distribution of a random element.
Proof
Given: Assume AC and let be finite. For and a real symmetric positive semidefinite matrix , a Borel probability law is denoted when a vector X with that law satisfies for every . Such a law exists, has mean m and covariance Sigma, and can be realized as with independent standard-normal coordinates. Singular Sigma is allowed. Uniqueness will be proved in the following characteristic-function lemma.
For any square-integrable vector X, covariance entries exist because . They are symmetric. Finite linearity gives , so covariance matrices are positive semidefinite. Conversely let a symmetric positive semidefinite Sigma be given and take its nonnegative symmetric square root A by [F2], so .
By [F3]–[F4], AC realizes d independent standard normals , each of mean zero and second moment one by [F1]. Define . Coordinate linear combinations are measurable, so X is a Borel random vector and its pushforward is a probability by [F9]. For any u and real t, [F5] and [F1] give . This equals the characteristic function of by [F1]; [F6] gives equality of scalar laws. Thus the required projection condition holds, including u=0 and all null directions.
Finite linearity gives . By [F7], for i different from j, and [F1] gives . Hence the covariance of AZ is , with every product integrable by the bound in step 1.1. If Sigma=0 then A=0 and the law is the point mass at m; no inverse or density is required even when only some directions are null. For d=1 the construction agrees with the scalar affine normal. If dimension zero is admitted, use the unique law on the singleton empty tuple instead. AC is spent in the standard-normal construction, independent-copy realization and scalar uniqueness supplier.
Characteristic function of a multivariate normal law
Statement
Assume AC. If , then This transform uniquely determines the law, including singular Sigma.
Facts & Assumptions
Every linear projection has the specified scalar normal law, and the vector law exists. Multivariate normal law, including singular covariance.
A scalar normal has transform exp(ims-sigma^2s^2/2). Characteristic function of a normal law.
Under AC scalar laws with equal characteristic functions agree. Uniqueness of a law from its characteristic function.
Under AC all linear projection laws determine the Borel vector law. Cramer wold device.
Proof
Given: Assume AC. If , then This transform uniquely determines the law, including singular Sigma.
For fixed t, [F1] makes scalar normal of mean and variance . Evaluate its characteristic function from [F2] at scalar frequency one. This gives the displayed formula. If t=0 both sides are one; if the scalar law is the point mass at and the same formula applies.
Let Y have another Borel probability law with the same displayed vector transform. For every u and scalar s, . Scalar uniqueness [F3] identifies each pair of projection laws. Then [F4] identifies the vector laws. Thus the construction in [F1] is independent of any realization choices. No determinant or inverse of Sigma is used. AC is inherited from [F1]–[F4]; in dimension zero the single possible law makes the assertion immediate.
Multivariate iid central limit theorem
Statement
Assume AC and let be a finite integer. Let be iid -valued random vectors with , mean m and covariance Sigma. Then The covariance may be singular.
Facts & Assumptions
The scalar iid CLT applies in each positive-variance projection. Lindeberg-Levy iid central limit theorem.
The Gaussian target with a positive semidefinite covariance exists, including singular covariance. Multivariate normal law, including singular covariance.
Convergence of every projection to those of a specified Borel probability implies vector weak convergence. Cramer wold device.
A nonnegative measurable function of integral zero vanishes almost everywhere. A nonnegative measurable function has integral exactly when it vanishes almost everywhere.
Finite means and covariance expansions obey linearity. The Lebesgue integral is linear on .
The Euclidean Cauchy–Schwarz inequality bounds a projection by the vector norm. Cauchy-Schwarz with its equality case, the triangle inequality for , the parallelogram law and polarisation.
Continuous scalar scaling preserves convergence in distribution. Continuous mapping theorem.
Proof
Given: Assume AC and let be a finite integer. Let be iid -valued random vectors with , mean m and covariance Sigma. Then The covariance may be singular.
For each fixed set . These are iid: inverse images under the continuous projection preserve the finite independence identities and the common law. By [F6], ; the latter is integrable since . Each centered coordinate product is integrable by , and symmetry of these products gives . Finite linearity gives and . Because this holds for every , the symmetric matrix is positive semidefinite. Consequently [F2] supplies the target law .
If v>0, [F1] gives . Scaling by is continuous; [F7] gives , which is the u-projection of G by [F2]. If v=0, [F4] applied to gives almost surely for each k. For each fixed n the union of the finitely many exceptional null sets is null, so the projected row sum is zero almost surely and has exactly the law N(0,0). Thus the same projection convergence holds without dividing by v.
The preceding convergence holds for every fixed u to the projections of the same specified probability G. [F3] therefore proves the vector conclusion. AC is inherited in the scalar CLT, Gaussian construction and Cramer–Wold theorem. A common null set across all u is neither claimed nor needed. When Sigma=0 all projections are in the zero-variance case. For d=1 this agrees with the scalar result after scaling.
Central-limit convergence is only in distribution
Remarks
The conclusions of Lindeberg-Levy iid central limit theorem, Lindeberg-Feller central limit theorem: sufficiency and Multivariate iid central limit theorem compare the row-sum laws with a Gaussian law. In Convergence in distribution of random elements, the limit is specified by tests on laws; the theorem does not construct a Gaussian random vector jointly with the original summands. The AC hypotheses of the cited theorems remain in force when applying them.
By contrast, Convergence in probability requires the probabilities of distance-from-the-limit events on a common probability space to tend to zero. Almost-sure convergence of real random variables requires pointwise convergence outside a null set on such a space. Neither conclusion is supplied merely by identifying the Gaussian limit law; a coupling or further argument is required. These statements do not deny that a suitable coupling can sometimes give stronger convergence.
The special result Convergence in distribution to a constant is convergence in probability concerns a point-mass limit. It applies to scalar degenerate normal limits on a common space, but does not turn a positive-variance normal limit into convergence in probability to a newly chosen Gaussian. A multivariate fully degenerate Gaussian limit requires the corresponding vector argument, which that real-valued theorem does not supply; a singular covariance with some positive-variance directions still gives a nonconstant law. No convergence rate or almost-sure version is asserted here.
5 · Examples, counterexamples and false statements
None yet.
Sources
- Durrett, Probability: Theory and Examples, Example 3.3.5
- Norris, Probability and Measure, Section 8
- Aldous and Chewi, Probability Theory notes, Lemmas 5.4-5.5
- Varadhan, Probability Theory, Chapter 3, Section 3.6
- Durrett, Probability: Theory and Examples, proof of Theorem 3.4.10
- Aldous and Chewi, Probability Theory notes, proof of Theorem 5.3
- Durrett, Probability: Theory and Examples, Theorem 3.4.1
- Aldous and Chewi, Probability Theory notes, Theorem 5.2
- Jech, The Axiom of Choice, §2.4, pp22–23; elementary AC-to-DC restriction proof
- Durrett, Probability: Theory and Examples, Section 3.1
- Aldous and Chewi, Probability Theory notes, Corollary 6.1
- Durrett, Probability: Theory and Examples, Section 3.4.2
- Billingsley, Probability and Measure, Section 27
- Durrett, Probability: Theory and Examples, Theorem 3.4.10
- Billingsley, Probability and Measure, Theorem 27.2
- Billingsley, Probability and Measure, Theorems 28.1-28.4 and Example 28.4
- Durrett, Probability: Theory and Examples, Exercise 3.4.12
- Aldous and Chewi, Probability Theory notes, Corollary 6.2
- Norris, Probability and Measure, Sections 8.1-8.2
- Durrett, Probability: Theory and Examples, Section 3.10
- Norris, Probability and Measure, Theorem 8.2.1
- Aldous and Chewi, Probability Theory notes, Theorem 8.2
- Durrett, Probability: Theory and Examples, Theorem 3.10.7
- Aldous and Chewi, Probability Theory notes, Theorem 8.4
- Durrett, Probability: Theory and Examples, Sections 3.2 and 3.4
- Aldous and Chewi, Probability Theory notes, Lectures 5-6