How statement and proof provenance work
The first chip identifies the source of the statement or construction; the second identifies the source of its local proof or verification.
- Literature-sourced: the exact statement appears in a cited source; only wording and notation differ.
- AI-adapted: a semantically identical restatement of literature-sourced material, modulo indexing, notation, and boundary cases adopted by the library.
- AI-generated: a genuinely novel statement formulated by AI, with no source for the claim itself.
These labels describe origin, not correctness: citations and verification chips remain separate evidence.
Probability Spaces Random Variables and Expectation
1 · Prerequisites
- Absolute and Conditional Convergence; Rearrangement; Products
- Binary Operations, Monoids, Groups and Subgroups
- Compactness in Metric Spaces
- Completeness, Completion, and Uniform Continuity
- Construction of the Natural Numbers
- Construction of the Real Numbers via Cauchy Sequences
- Construction of the Real Numbers via Dedekind Cuts
- Continuity, IVT, EVT, and Uniform Continuity
- Convexity
- Cosets, Index and Lagrange's Theorem
- Countability and Uncountability
- Divisibility, Euclidean Domains, Principal Ideal Domains and Unique Factorisation
- Dual Spaces, Bilinear and Quadratic Forms, and Sylvester's Law of Inertia
- Finite Counting, Factorials and Binomial Coefficients
- Finite Probability Spaces and Random Variables
- Foundations of the Real Numbers for Analysis
- Ideals, Quotient Rings and the Isomorphism Theorems for Rings
- Inner Product Spaces, Gram-Schmidt, Projections and Adjoints
- Lebesgue-Stieltjes Measures and Distribution Functions
- Limits of Real Functions
- limsup, liminf, and Subsequential Limits
- Linear Independence, Bases and Dimension
- Measurable Functions and Simple Approximation
- Measures and Their Basic Properties
- Metric Spaces
- Monotone Functions, Discontinuities, and Continuity Sets
- Monotone Sequences, Bolzano-Weierstrass, and Cauchy Completeness
- Normal Subgroups and Quotient Groups
- Order, Zorn's Lemma, and the Axiom of Choice
- Outer Measure and the Caratheodory Extension Theorem
- Polynomial Rings, the Division Algorithm and Roots
- Power Series and Real-Analytic Functions
- Product Measures and the Fubini Tonelli Theorems
- Properties of the Integral and the Working FTC
- Relations, Functions, and Quotients
- Rings, Subrings, Integral Domains and Fields
- Roots, Rational Powers, and Classical Inequalities
- Sequences and Limits
- Sequences and Series of Functions; Uniform Convergence
- Series: Convergence and the Nonnegative Tests
- Sigma Algebras and Borel Sets
- Simple Field Extensions and the Construction of the Complex Numbers
- Subspaces, Products, and Quotients
- Suprema and Infima
- The Derivative and the Mean Value Theorems
- The Exponential Function
- The Lebesgue Integral and the Convergence Theorems
- The Logarithm and General Powers
- The Lᵖ Spaces Holder Minkowski and Riesz Fischer
- The Riemann Integral: Definition and Integrability
- The Topology of Euclidean Space
- The ZFC Axioms and the Basic Set Constructions
- Topological Spaces and Continuity
- Topology of ℝ
- Triangularisation, Generalised Eigenspaces and Jordan Canonical Form
- Vector Spaces, Linear Subspaces, Span and Direct Sums
2 · Summary
This page moves from measure spaces of total mass to the standard probability toolkit built on them. The first block identifies the finite probability model with probability measures on finite full power sets, then defines random elements, laws, and distribution functions in a way that matches the earlier finite page exactly.
Expectation is defined as the ambient Lebesgue integral, so change of variables, layer-cake formulas, Jensen, Markov, Chebyshev, Holder, and Cauchy-Schwarz all become probability-space corollaries of the measure-theory backbone. The page closes with moments, variance, covariance, and the normal equations for best affine prediction.
3 · Logical flowchart
4 · Definitions, theorems and proofs
Basic identities for a probability measure
Statement
Let be a probability space, let , and let be events.
- .
- If , then .
- If , then .
- , and for every natural ,
- .
- If , then If , then
Facts & Assumptions
Given: A probability space , events , and an event sequence .
A probability measure is a measure of total mass (Probability measures and probability spaces).
Measures are monotone, set differences subtract when the smaller set has finite measure, subadditivity holds, continuity from below holds, continuity from above holds once one set has finite measure, and finite inclusion-exclusion holds for finite-measure sets (Measures are monotone, Measure of a set difference when the smaller set has finite measure, Finite and countable subadditivity of measures, Continuity from below for measures, Continuity from above when one set has finite measure, Inclusion-exclusion for a nonempty finite family of finite-measure sets).
Proof
Because by [L1] and , [L2] gives so . Monotonicity in [L2] also gives .
Since every probability is at most , the finite-measure hypotheses in [L2] apply to , , and . Thus if , then subadditivity gives the countable and finite union bounds, finite inclusion-exclusion gives and continuity from below and from above give the two monotone-limit formulas, because a decreasing probability sequence always has finite first term.
Steps 1.1 and 1.2 are exactly the stated probability identities.
Finite probability spaces are exactly finite full-power-set probability spaces
Statement
Let be a finite set.
- If is a finite probability space in the sense of Finite probability spaces, outcome weights, events, and event probabilities, then is a probability measure on .
- Conversely, if is a probability measure on , then makes a finite probability space and
These two constructions are inverse to each other. In particular, zero-weight outcomes remain genuine outcomes in both descriptions.
Facts & Assumptions
Given: A finite set .
A finite probability space is a finite set with nonnegative weights summing to , every subset is an event, and event probabilities are the corresponding sub-weight sums (Finite probability spaces, outcome weights, events, and event probabilities).
A probability measure is a measure of total mass (Probability measures and probability spaces).
On a finite sigma-algebra, the atoms partition the space, every measurable set is the union of the atoms it contains, and a measure is the sum of the atom masses over those atoms (A measure on a finite sigma-algebra is a finite weighted sum over its atoms).
Proof
If is a finite probability space, then [L1] already states that every subset of is an event and that is its probability. Therefore is a probability measure on by [L2].
Conversely, let be a probability measure on and put . Each singleton is an atom of the full power-set sigma-algebra, and every is the union of the singletons it contains. Thus [L3] gives Taking yields , and nonnegativity of comes from the measure axioms inside [L2]. So is a finite probability space.
Step 1.1 constructs a full-power-set probability measure from any finite weight model, and step 1.2 recovers exactly those singleton weights from any full-power-set probability measure. Hence the two descriptions are equivalent, including the boundary case of outcomes with weight .
Agreement with the published finite probability-space definition
The earlier item Finite probability spaces, outcome weights, events, and event probabilities and the present measure-theoretic formulation describe the same finite object. The theorem above does not replace the published finite definition; it identifies it with the probability-measure language on the full power set.
This agreement keeps two boundary features visible. First, every subset of a finite sample space remains measurable. Second, an outcome of weight is still part of the sample space, so a nonempty event may still have probability .
Random elements and real random variables
Definition
Let be a probability space and let be a measurable space. A random element of is a measurable map in the sense of A measurable function between measurable spaces.
When and from The Borel sigma-algebra of a topological space, the map is a real random variable.
Finite random variables are measurable
Statement
Let be a finite probability space, and regard it as the probability space from Finite probability spaces are exactly finite full-power-set probability spaces. Then every function is a real random variable.
In particular, the published finite definition Real random variables on finite probability spaces and their finite distributions is exactly the measure-theoretic definition on that full-power-set probability space.
Facts & Assumptions
Given: A finite probability space and a function .
The theorem on finite probability spaces identifies with a probability measure on (Finite probability spaces are exactly finite full-power-set probability spaces).
A real random variable is a measurable map from the sample-space sigma-algebra to (Random elements and real random variables).
On a finite probability space, a real random variable is simply a function (Real random variables on finite probability spaces and their finite distributions).
Proof
By [L1], every subset of is measurable. Hence for every Borel set , the preimage is a subset of , so it lies in . Therefore is measurable.
Step 1.1 proves that every finite random variable in the sense of [L3] is a real random variable in the sense of [L2], so the two notions agree exactly on finite full-power-set probability spaces.
Law or distribution of a random element
Definition
Let be a random element. Its law or distribution is the set function
Thus the law of records the probability of each measurable target set by pulling it back to an event in the original probability space.
The law of a random element is a probability measure
Statement
Let be a random element. Then is a probability measure on .
Facts & Assumptions
Given: A random element .
The law is defined by (Law or distribution of a random element).
A random element is measurable, so measurable target sets have measurable preimages (Random elements and real random variables).
A probability measure is a measure with total mass (Probability measures and probability spaces).
Proof
By [L2], every has , so [L1] is well defined. Also and , so
If is a pairwise disjoint sequence in , then the preimages are pairwise disjoint and Therefore
Steps 1.1 and 1.2 show that is a measure of total mass , hence a probability measure by [L3].
Laws commute with measurable maps
Statement
Let be a random element, and let be measurable. Then is a random element and for every ,
Facts & Assumptions
Given: A random element and a measurable map as in the Statement.
Composition of measurable maps is measurable (Composition with a Borel measurable outer map preserves measurability).
The law of a random element is defined by pullback of measurable target sets (Law or distribution of a random element).
The law of any random element is a probability measure (The law of a random element is a probability measure).
Proof
By [L1], the composite is measurable, hence a random element.
For every , Therefore [L2] gives This right-hand side is defined because [L3] makes a probability measure on .
Steps 1.1 and 1.2 prove the measurable-map compatibility of laws.
Cumulative distribution function of a real random variable
Definition
Let be a real random variable. Its cumulative distribution function is the function
The second expression is the same quantity written in terms of the law Law or distribution of a random element of .
Probability laws correspond to distribution functions
Statement
Assume the Axiom of Countable Choice.
- Let be a real random variable, let be its law, and let . Then is nondecreasing and right-continuous, satisfies and obeys
- Conversely, if is nondecreasing and right-continuous with then there is a unique Borel probability measure on such that equivalently
Facts & Assumptions
Given: Countable Choice, a real random variable , its law , and a function as in part 2.
The law is a probability measure on (Law or distribution of a random element, The law of a random element is a probability measure).
For measures, monotonicity, set-difference subtraction, continuity from below, and continuity from above are available (Measures are monotone, Measure of a set difference when the smaller set has finite measure, Continuity from below for measures, Continuity from above when one set has finite measure).
Assuming Countable Choice, finite-on-compacts Borel measures on correspond to nondecreasing right-continuous functions modulo constants, and the interval increments determine the measure (Assuming countable choice, finite-on-compacts Borel measures on correspond to nondecreasing right-continuous functions modulo constants).
Proof
If , then , so [L1] and [L2] give . Also so the finite-measure difference formula from [L2] yields
For fixed , the sets decrease to , and . Hence [L2] gives right continuity of . Likewise and , so continuity from below and from above give
Put . Then is still nondecreasing and right-continuous, so [L3] gives a unique Borel measure finite on compact sets with
For fixed , the sets increase to . Hence [L2] and step 1.3 give Applying continuity from below once more to shows , so is a probability measure. If is another Borel probability measure with for all , then for every , and [L3] gives .
Steps 1.1 and 1.2 prove part 1, and steps 1.3 and 2.1 prove part 2.
Atoms and continuity points of a law
Definition
Let be a real random variable with law and cumulative distribution function .
- A point is an atom of the law of when it is an atom of the Borel measure , equivalently when
- A point is a continuity point of the law of when is continuous at .
The atom language belongs to the measure An atom of a measure on , while the continuity-point language belongs to the distribution function Cumulative distribution function of a real random variable.
Expectation of a nonnegative or integrable random variable
Definition
Let be a probability space.
- If is measurable, its expectation is the extended-valued integral
- If or is integrable in the sense of Integrable real and complex functions, and their integrals, its expectation is again now a finite real or complex number.
For an integrable real random variable, where and are its positive and negative parts.
Expectation depends only on the almost-everywhere class
Statement
If and are integrable real or complex random variables on one probability space and almost surely, then
Thus expectation is a function of the -equivalence class.
Facts & Assumptions
Given: Integrable random variables on one probability space.
Expectation is the Lebesgue integral with respect to the underlying probability measure (Expectation of a nonnegative or integrable random variable).
Two integrable functions are equal almost everywhere exactly when all of their integrals over measurable sets agree (Two integrable functions are equal almost everywhere exactly when all of their indefinite integrals agree).
Proof
Since almost surely and both are integrable, [L2] applied to the measurable set gives
Rewriting the two integrals as expectations by [L1] yields .
Change of variables for expectation
Statement
Let be a random element, let be its law, and let or be measurable.
- If , then
- If is integrable, then is integrable with respect to and the same formula holds:
Facts & Assumptions
Given: A random element , its law , and a measurable map as in the Statement.
The law is a probability measure on (Law or distribution of a random element, The law of a random element is a probability measure).
Measurable outer maps preserve measurability under composition (Composition with a Borel measurable outer map preserves measurability).
Every nonnegative measurable function is the increasing limit of nonnegative simple functions, monotone convergence holds, and the nonnegative integral agrees with the simple integral on simple functions (Every nonnegative measurable function is the increasing limit of simple measurable functions, Monotone convergence for the integral, The nonnegative integral agrees with the simple integral on simple functions, The integral of a nonnegative simple function).
Expectation is integration against , and real or complex integrability uses the positive-negative and real-imaginary decompositions (Expectation of a nonnegative or integrable random variable, Integrable real and complex functions, and their integrals).
Measurable functions are closed under the elementary operations used by the integral, and the Lebesgue integral is linear on (Closure properties of measurable functions used by the integral, The Lebesgue integral is linear on ).
Proof
By [L2], the composite is measurable. If is a nonnegative simple function on with the pairwise disjoint, then is a nonnegative simple function on . Using [L1], [L3], and [L4],
Assume now that . By [L3], choose nonnegative simple functions on . Then by step 1.1, so monotone convergence on both spaces and step 1.1 give
Assume is integrable. Then , so step 2.1 applied to gives Hence is -integrable. For real-valued , [L4] and [L5] give with both parts nonnegative, so step 2.1 and linearity yield
If is complex-valued, write with real measurable parts . The inequality and step 3.1 show that and are integrable, so the real-valued case applied to and , followed by complex-linearity from [L5], gives
Step 2.1 proves the nonnegative case, while steps 3.1 and 4.1 prove the integrable real and complex cases.
Expectation agrees with the published finite weighted sum
Statement
Let be a finite probability space and let . After identifying with the full-power-set probability space of Finite probability spaces are exactly finite full-power-set probability spaces, the expectation defined by Expectation of a nonnegative or integrable random variable agrees with the published finite formulas:
Facts & Assumptions
Given: A finite probability space and a real-valued function .
Finite probability spaces are exactly full-power-set probability spaces, and every finite real random variable is measurable there (Finite probability spaces are exactly finite full-power-set probability spaces, Finite random variables are measurable).
Change of variables for expectation identifies with the integral of the identity function against the law of (Change of variables for expectation).
The published finite expectation is , and it is also the sum over attained values weighted by their probabilities (Expectation of a real random variable on a finite probability space, Expectation is the sum of each attained value times its probability).
Proof
By [L1], the general expectation and the law of are defined on the same full-power-set probability space attached to .
Applying [L2] to the identity map on gives For a finite random variable, [L3] identifies this quantity with both
Thus the general expectation is exactly the published finite weighted-sum expectation and its finite-distribution reformulation.
The expectation of an indicator is the probability of the event
Statement
Let be a probability space and let . Then the indicator satisfies
Facts & Assumptions
Given: A probability space and an event .
Expectation of a nonnegative random variable is its integral with respect to (Expectation of a nonnegative or integrable random variable).
The complement identity gives (Basic identities for a probability measure).
On nonnegative simple functions, the nonnegative integral is the simple integral (The nonnegative integral agrees with the simple integral on simple functions, The integral of a nonnegative simple function).
Proof
The function is the simple function Therefore [L1] and [L3] give
Using [L2] only to note that is the complementary event in the same probability space, step 1.1 simplifies to .
Layer-cake formulas for random variables
Statement
Let be a probability space.
- If is measurable, then where the right-hand side may be .
- If is an integrable real random variable, then
Facts & Assumptions
Given: A probability space and a random variable in the relevant clause.
Expectation is integration against , and with for real (Expectation of a nonnegative or integrable random variable, The positive and negative parts of a function).
The layer-cake formula with gives for measurable (For 0 < p < infinity, the layer-cake formula computes the integral of |f|^p from the distribution function).
The Lebesgue integral is linear on (The Lebesgue integral is linear on ).
Proof
Apply [L2] with and . Because , one has and , so [L1] gives
If is integrable and real, then and [L1] gives . By step 1.1 applied to and , Subtracting these identities and using [L3] proves the second formula.
Linearity, monotonicity, and the modulus bound for expectation
Statement
Let be integrable real or complex random variables on one probability space.
- For scalars ,
- If and are real-valued and almost surely, then
Facts & Assumptions
Given: Integrable random variables .
Expectation is the Lebesgue integral against the probability measure (Expectation of a nonnegative or integrable random variable).
The Lebesgue integral is linear on , the nonnegative integral is monotone, and the modulus of an integral is bounded by the integral of the modulus (The Lebesgue integral is linear on , Monotonicity and nonnegative homogeneity of the nonnegative integral, The modulus of an integral is bounded by the integral of the modulus).
Proof
Rewriting expectation as the integral by [L1], linearity in [L2] gives
Applying the integral triangle inequality from [L2] after [L1] gives
If almost surely, then almost surely. Hence [L1] and [L2] give so .
Steps 1.1, 1.2, and 2.1 prove the three assertions.
Moments, variance, and covariance on a probability space
Definition
Let be a real random variable on a probability space.
- For , the th absolute moment of exists when is integrable, and is then
- If is integrable, its mean is .
- If is square-integrable, its variance is
If and are square-integrable real random variables on the same probability space, their covariance is
Variance and covariance identities for random variables
Statement
Let be square-integrable real random variables on one probability space. Then Moreover, covariance is symmetric and bilinear on finite linear combinations. On finite full-power-set probability spaces these formulas reduce to the published finite identities.
Facts & Assumptions
Given: Square-integrable real random variables .
Variance and covariance are the expectations of the centered square and centered product (Moments, variance, and covariance on a probability space).
Expectation is linear on integrable random variables (Linearity, monotonicity, and the modulus bound for expectation).
Finite probability spaces agree with the full-power-set probability-space formalism (Finite probability spaces are exactly finite full-power-set probability spaces).
Proof
Expanding and applying [L2] gives Likewise,
The covariance formula in step 1.1 is symmetric in and , so . If and are finite linear combinations of square-integrable real random variables, expanding and using [L2] gives
On a finite probability space, [L3] identifies the general formulas above with the already-published finite ones, so the finite and general identities agree exactly.
Jensen's inequality for expectation
Statement
Let be a probability space, let be an integrable real random variable, let be an interval containing for almost every , and let be convex.
Assume additionally that is either integrable or nonnegative, so its expectation is defined by Expectation of a nonnegative or integrable random variable. Then In the nonnegative case the right-hand side may be .
Facts & Assumptions
Given: A probability space, an integrable real random variable , an interval containing its almost-everywhere range, and a convex such that is integrable or nonnegative.
Expectation is integration against the underlying probability measure (Expectation of a nonnegative or integrable random variable).
Jensen's integral inequality holds for a probability measure whenever the composed function is integrable (Jensen's integral inequality for a probability measure).
Proof
If is nonnegative and , then the claimed inequality is automatic, because is a real number while the right-hand side is .
In every remaining case, is integrable: this is assumed directly, or follows from nonnegativity and finite expectation. Thus [L1] and [L2] applied to the probability measure give
Steps 1.1 and 1.2 cover the infinite and finite expectation cases.
Markov's inequality for random variables
Statement
If is a nonnegative random variable on a probability space and , then
Facts & Assumptions
Given: A nonnegative random variable and a real number .
Expectation is integration against the probability measure (Expectation of a nonnegative or integrable random variable).
The integral Markov inequality states for nonnegative measurable and (Chebyshev-Markov inequality for the integral).
Proof
Apply [L2] to the probability measure , the function , and the threshold . Rewriting the integral by [L1] gives
Step 1.1 is exactly Markov's inequality for random variables.
Chebyshev's inequality for random variables
Statement
If is a square-integrable real random variable and , then
Facts & Assumptions
Given: A square-integrable real random variable and a real number .
Variance is the expectation of the squared centered variable (Moments, variance, and covariance on a probability space).
Markov's inequality applies to every nonnegative random variable (Markov's inequality for random variables).
Proof
The random variable is nonnegative, and
Applying [L2] to and using [L1] gives
Holder's inequality for random variables
Statement
Let be conjugate exponents. If and are real random variables in the spaces named by the corresponding clause of Holder's inequality for integrals, including the endpoint cases, then
In particular, is integrable.
Facts & Assumptions
Given: Real random variables and conjugate exponents as in the Statement.
Expectation is integration against the probability measure (Expectation of a nonnegative or integrable random variable).
Holder's integral inequality, including the endpoint cases, holds on every measure space (Holder's inequality for integrals, including the endpoint cases).
Proof
Apply [L2] to the measure space . Rewriting the left-hand side with [L1] gives
The same theorem [L2] already states that the right-hand side is finite in every allowed case, so is integrable.
Cauchy-Schwarz for random variables
Statement
If are real random variables, then
Equality holds if and only if at least one of is zero almost surely, or there is a constant with
Facts & Assumptions
Given: Real random variables .
Holder's inequality on a probability space specializes to (Holder's inequality for random variables).
The Cauchy-Schwarz equality criterion is already proved for general measure spaces (Cauchy-Schwarz inequality for ).
Proof
Step [L1] at gives
The equality clause is exactly the probability-measure specialization of [L2].
Lyapunov's moment inequality on a probability space
Statement
Let be a probability space and let .
- If and , then and
- If and , then and
Facts & Assumptions
Given: A probability space and exponents .
A probability measure has total mass (Probability measures and probability spaces).
On a finite measure space, includes into with factor for finite , and includes into with factor (Finite-measure includes into for ).
Proof
If , the inequality is equality. If , apply [L2] with and use [L1] to collapse the factor to .
If , the same specialization of [L2] and [L1] gives
Steps 1.1 and 1.2 are exactly the finite- and cases of Lyapunov's moment inequality on a probability space.
The second-moment lower bound for positive probability
Statement
Let be a nonnegative square-integrable real random variable.
- If , then
- If , then almost surely and .
Facts & Assumptions
Given: A nonnegative square-integrable real random variable .
The expectation of an indicator is the probability of its event (The expectation of an indicator is the probability of the event).
Cauchy-Schwarz holds for square-integrable random variables (Cauchy-Schwarz for random variables).
Square-integrability means that the second moment is finite (Moments, variance, and covariance on a probability space).
Proof
The identity holds pointwise because . Applying [L2] to and gives By [L1], this is
If , divide the inequality in step 1.1 by that positive number. If , then step 1.1 forces as well, and since , zero second moment means almost surely, hence almost surely and .
Step 2.1 proves both the positive-second-moment case and the zero boundary case.
The general inequalities compare cleanly with the published finite ones
The probability-space inequalities above recover the published finite statements exactly in the Markov, Chebyshev, and Cauchy-Schwarz cases, and they isolate the precise extra nonnegativity needed for the positive-probability bound.
- Markov's inequality for random variables restricts to Markov's inequality on a finite probability space.
- Chebyshev's inequality for random variables restricts to Chebyshev's inequality on a finite probability space.
- Cauchy-Schwarz for random variables restricts to Cauchy-Schwarz for finite random variables: .
- The second-moment lower bound for positive probability gives the nonnegative specialization of The finite second-moment bound when , with in place of .
The earlier finite proofs remain the canonical finite arguments. The present page packages them as consequences of the general integral theory on a probability space.
Normal equations for best affine prediction
Statement
Let be square-integrable real random variables on one probability space, and write where denotes the almost-everywhere class from The space as the quotient by null functions.
Then there is a unique class minimizing over . This minimizing class has an affine representative , and its coefficients satisfy
The class is unique. If the covariance matrix is singular, the coefficient vector need not be uniquely determined by the normal equations, but any two solutions yield the same predictor almost surely.
Facts & Assumptions
Given: Square-integrable real random variables .
Variance and covariance are given by centered expectations, satisfy the identities and are bilinear on finite linear combinations (Moments, variance, and covariance on a probability space, Variance and covariance identities for random variables).
Cauchy-Schwarz gives integrability of products of square-integrable random variables (Cauchy-Schwarz for random variables).
is the quotient by almost-sure equality, and the span of a set is exactly the set of finite linear combinations (The space as the quotient by null functions, Linear combination of a finite list, and the span as the smallest linear subspace containing , is exactly the set of linear combinations of finite lists of elements of , and ).
In a finite-dimensional inner product space, the orthogonal projection onto any subspace is the unique nearest point in that subspace (The orthogonal projection is the unique nearest point in the subspace).
A nonnegative measurable function has integral exactly when it vanishes almost everywhere (A nonnegative measurable function has integral exactly when it vanishes almost everywhere).
Proof
Let This space is finite-dimensional because it is spanned by vectors, and is a subspace of . Equip with the usual inner product. Applying [L4] to the vector and the subspace gives a unique class minimizing over .
By [L3], every class in has an affine representative of the form . Put and . Then Expanding the square and using [L1] shows that the mixed term with the final constant vanishes, so Hence every affine representative of the minimizing class must satisfy
After imposing the intercept from step 2.1, define Using [L1] and [L2],
Let . Substituting into step 3.1 and subtracting gives Therefore, if satisfies the normal equations, then for every , so minimizes .
Conversely, suppose minimizes . Fix and take , where is the th standard basis vector. Then step 4.1 gives for every real . The coefficient of must therefore be , since the same inequality holds for both signs of . This is exactly the th normal equation, and was arbitrary.
Let and be two coefficient vectors satisfying the normal equations, and let Subtracting the two linear systems gives for every . Multiplying the th equation by and summing over , then using bilinearity from [L1], yields Because , [L5] gives almost surely. Thus the corresponding predictors agree almost surely, so the minimizing class in is unique even when the coefficient vector is not.
Step 1.1 gives existence and uniqueness of the minimizing class, step 2.1 identifies the optimal intercept, steps 4.1 and 5.1 characterize the optimal centered coefficients by the covariance normal equations, and step 6.1 proves that all coefficient solutions yield the same predictor almost surely. This is exactly the best affine prediction statement.
Best affine prediction from one random variable
Statement
Let and be square-integrable real random variables.
- If , the unique best affine predictor of from is
- If , then every best affine predictor is almost surely equal to the constant .
Facts & Assumptions
Given: Square-integrable real random variables .
Best affine predictors are characterized by the normal equations (Normal equations for best affine prediction).
Variance is (Moments, variance, and covariance on a probability space).
Proof
With one predictor variable, the normal equation from [L1] is By [L2], this is
If , then step 1.1 makes the normal equation Hence every real solves it, and [L1] says that all corresponding affine predictors yield the same optimal class. Taking gives the constant predictor , so every best affine predictor is almost surely equal to that constant.
If , step 1.1 gives Substituting this into the intercept formula from [L1] yields the displayed predictor
Steps 2.2 and 2.1 give the positive-variance and zero-variance cases.
5 · Examples, counterexamples and false statements
None yet.
Sources
- Rick Durrett, Probability: Theory and Examples, 5th ed., Section 1.1
- J. R. Norris, Probability and Measure, Section 1.9
- S. R. S. Varadhan, Probability Theory, Definition 1.10
- J. R. Norris, Probability and Measure, Section 2.1
- Rick Durrett, Probability: Theory and Examples, 5th ed., Section 1.3
- Rick Durrett, Probability: Theory and Examples, 5th ed., Section 1.2
- J. R. Norris, Probability and Measure, Section 2.2
- S. R. S. Varadhan, Probability Theory, Section 1.6
- S. R. S. Varadhan, Probability Theory, Section 1.4
- J. R. Norris, Probability and Measure, Section 3.3
- J. R. Norris, Probability and Measure, Section 2.3
- Jean-Francois Le Gall, Integration, Probabilities and Stochastic Processes, Section 8.1.6
- Rick Durrett, Probability: Theory and Examples, 5th ed., Sections 1.4 and 1.6
- Rick Durrett, Probability: Theory and Examples, 5th ed., Section 1.5
- Rick Durrett, Probability: Theory and Examples, 5th ed., Section 1.6.3
- Rick Durrett, Probability: Theory and Examples, 5th ed., Section 1.6
- J. R. Norris, Probability and Measure, Section 4.2
- Jean-Francois Le Gall, Integration, Probabilities and Stochastic Processes, Section 8.2.1
- Rick Durrett, Probability: Theory and Examples, 5th ed., Section 1.6.1
- J. R. Norris, Probability and Measure, Theorem 4.3.2
- J. R. Norris, Probability and Measure, Theorem 4.4.1
- J. R. Norris, Probability and Measure, Theorem 4.4.2 discussion
- Jean-Francois Le Gall, Integration, Probabilities and Stochastic Processes, Section 8.2.2