Motivations for Measure Theory
(This post is based on some of the notes that I took while watching the excellent Todd Kemp lectures on Probability Theory)
Suppose you have a circle and I ask you to find the proportion of the length of the arc corresponding to the points of a subset E of a set denoting a circle.
a common approach would be:
\[ proportion = \frac{1}{2\pi} \int_{\theta \in E} d\theta \]
This computation in the riemann sense, however, is problematic as we shall see.
Let us try to make our definition of proportion more rigorous. A proportion is a number between 0 and 1 therefore introducing a function that assigns a number \(\in [0,1]\) makes much more sense. Formally speaking:
Suppose you have a circle,
\[ S = \{ z \in \mathbb{C} : |z| = 1 \} \]
and you take any \(E \subseteq S\), then define a function \(P_r:2^{S} \rightarrow [0,1]\) (for proportion). Intuitively, we would wish for the function to exhibit certain properties such as -
(1) \(P_r(S) = 1~ \&~ P_r(\phi) = 0\).
(2) If \(E_1,E_2 \subseteq S~ \&~ E_1 \cap E_2 = \phi\) then \(P_r(E_1 \cup E_2) = P_r(E_1) + P_r(E_2)\).
(3) And while we are at it, let us also have the function be countably additive. Let I be an index set and let \(E_i, i \in I\), be disjoint \((E_i \cap E_j = \phi, \forall~ i \neq j, i,j \in I)\) then: \[ P_r(\bigcup_{i=1}^{\infty} E_i) = \sum_{i=1}^{\infty} P_r(E_i).\]
Is there anything else that we might like the function to have? Take \(E_1 = \{ z \in S : 0 \leq arg(z) \leq \frac{\pi}{4} \}\) and \(E_2 = \{ z \in S : \frac{\pi}{2} \leq arg(z) \leq \frac{3\pi}{4} \}\) and observe that both of these subsets essentially have the same arc length (even though, we haven't yet defined a notion of length but since we are intuitively introducing a function, it makes sense for similar or rather, more structurally speaking, congruent* sets to have the same arc length). Therefore, we get our 3rd property for our function.
(4) If \(E_1,E_2 \subseteq S~\) are congruent*, then \(P_r(E_1) = P_r(E_2)\).
note * = two sets A & B are said to be congruent to each other if \(\exists f:A \rightarrow B \ni f\) is bijective.
Seems like our function definition is now complete and we now have a notion of proportions! Unfortunately, the above properties are not logically consistent with each other, especially 1,3 and 4 as we are going to prove.
Theorem = Properties 1,3 and 4 are logically inconsistent with each other considering the axiom of choice*.
Proof: Let \(E \subseteq S\) & let \(\mu \in S\).
claim: The subset E and the set \(\mu E = \{ \mu z: z\in E \}\) are congruent.
proof of claim: Consider a function \(f:E \rightarrow \mu E \ni f(z) = \mu z~ \forall~ z \in E\). Proving that f is bijective is enough to prove the congruence.
Injectivity: - suppose not, let \(y \in \mu E\) and let \[ x_1,x_2 \in E \ni f(x_1) = f(x_2) = y \] \[\rightarrow \mu x_1 = \mu x_2 = y\] \[\rightarrow x_1 = x_2 \] Surjectivity: \(\forall~ x \in \mu E, x \in \{ \mu z : z \in E \} \rightarrow x = \mu z\) for some \(z \in E, \forall~ x \in \mu E\). \[ \therefore E \cong \mu E \]
hence by our 3rd property, \(P_r(E) = P_r(\mu E)~ \forall~ \mu \in S, E \subseteq S\)
let \(\mu = e^{i\theta}\), we see that \(\mu E\) is the set E, rotated by an angle \(\theta\).
Now consider T ⊂ S, \[ T = \{ e^{2\pi i t} : t \in \mathbb{Q} \} \] and observe that the set T is countable.
now in order to proceed, we must learn a new mathematical object, equivalence classes.
Equivalence classes: Suppose you have an equivalence relation '\(\sim\)' between sets A and B then \(\forall~ a \in A\), the equivalence class of a i.e. \([a]\) is given as: \[ [a] = \{ b \in B : a \sim b \} \] equivalence classes are mainly used to denote some sort of equality between 2 mathematical objects.
let S\T = {equivalence classes in S where \(z \sim w \leftrightarrow z = \mu w\) for some \(\mu \in T\)} for example: Suppose \(z = e^{2\pi i \frac{1}{5}}\) then \[ \Rightarrow [z] = \{ \mu z : \mu \in T \} \]
\[ \Rightarrow [z] = \{ e^{2\pi i q} e^{2\pi i \frac15} : q \in \mathbb Q \} \]
\[ \Rightarrow [z] = \{ e^{2\pi i (q+\frac15)} : q \in \mathbb Q \} \] \(\forall~ z \in S\)
Choose* exactly 1 representative element \(\psi\) from each equivalence class & let \(\Phi = \{ \psi \} \subset S\) be the collection of all representatives.
note * = This is exactly where we make use of the axiom of choice. The axiom of choice says that given a family of non-empty sets, \(E_i\) and an index set \(I\), there exists a function \(f:I \rightarrow \bigcup \limits_{i \in I}E_i \ni f(i) \in E_i~ \forall~ i\).
Note that this is not so obvious because the number of sets in the family that we have considered may not be countable at all! and as a matter of fact, the indexing set is also not required to be countable (as someone who never thought that the index set could be something other than the set of natural numbers, this came as a shock)
A simple way of proving the existence of such uncountable index sets is to just let \(I = E\) where E is uncountable and let f be the identity function.
claim: \(S = \bigcup \limits_{\mu \in T} \mu \Phi\)
proof of claim: Let \(x \in \Phi \ni x = e^{2\pi i r}, r \in \mathbb{R}\) then x lies in some equivalence class containing elements of the form \(\{ e^{2\pi i (q + r)} : q \in Q \}\) thus \(\forall\) such x, \(\mu x\) gives us all the points achievable by rotating by \(\mu\) i.e. \(\mu \Phi\). \(\bigcup \limits_{\mu \in T} \mu \Phi\) just gives all the points achievable by rotating by any value of \(\mu \in T\) i.e. S.
claim: if \(\mu_1, \mu_2 \in T, \mu_1 \Phi \cap \mu_2 \Phi = \phi\)
proof of claim: Suppose \(q_1 \neq q_2, q_1,q_2 \in \Phi \ni \mu_1 q_1 = \mu_2 q_2\) \[ \Rightarrow q_1 = \mu_1^{-1} \mu_2 q_2 \]
\[ \Rightarrow q_1 = \mu ' q_2 \leftrightarrow q_1 \sim q_2 \] but \(q_1,q_2 \in \Phi\), contradiction!
\[ \therefore \mu_1 \Phi \cap \mu_2 \Phi = \phi \] so now we can make use of the (3) property of our proportion function - \[ P_r(S) = P_r(\bigcup \limits_{\mu \in T} \mu \Phi) = \sum_{\mu \in T} P_r(\mu \Phi) \]
\[ \Rightarrow 1 = P_r(S) = \sum_{\mu \in T} P_r(\Phi) \] (as \(\Phi \cong \mu \Phi\))
claim: The set of all the equivalence classes i.e. \(\Phi\) is uncountable.
proof of claim: Suppose not then: \[ S = \bigcup \limits_{i=1}^{\infty} E_i, \] thus, \(\forall~ i, E_i \subset S \rightarrow S\) is countable but \(S = \{ e^{2\pi i r} : r \in \mathbb{R} \}\) which is clearly not countable (otherwise we would able to find a bijection between \(\mathbb{N}\) and \(\mathbb{R}\)) therefore, \(\Phi\) is uncountable.
and since we have \(P_r:2^{S} \rightarrow [0,1], \sum_{\mu \in T} P_r(\Phi)\) is either 0 or \(\infty\), in any case, not 1. \[ \therefore 1 = \sum_{\mu \in T} P_r(\Phi) \neq 1 \] thus the properties (1),(2) and (3) are logically inconsistent.
How does it affect us? Recall the problem at the beginning of this post and notice that in the following definition:
\[ P_r(E) = \frac{1}{2\pi} \int_{\theta \in E} d\theta \] not all sets \(E \subseteq S\) works as we just proved! So where exactly did things went wrong? Notice that, while defining the proportion measure (the proportion function basically), the domain that we considered was \(2^{S}\) which, by definition, allowed us to work with all possible subsets of S, freely without any restriction and that's what led us to encounter such bizarre things! We clearly need to do something about the domain of such measures, either restrict them or provide certain conditions. Such domains are actually a mathematical object called as \(\sigma\) -fields/algebras. The whole idea is that there are some sets, that are subsets of S, that cannot be measured i.e. they are unmeasurable sets.
note - Even though 2S itself is a \(\sigma -\) field, the problem is that it contains every imaginable (and unimaginable, funnily enough!) subset of S which also includes sets like the ones we used while proving the theorem above.
If we allow ourselves to be content with the fact that given a sample space, (\(\Omega\)), there may be some subsets, \(E \subset S \ni~ E\) is unmeasurable, we find some pretty strange theorems in mathematics for example, the infamous, Banach-Tarski theorem which states that -
Banach-Tarski Theorem = Given any two subsets \(E,F \subset \mathbb{R}^{d}, d \geq 3\), with non-empty interior*, there exists finite disjoint partitions, \[ E = E_1 \cup E_2 \cup ... \cup E_n \] \[ F = F_1 \cup F_2 \cup ... \cup F_n \] \[ \ni E_j \cong F_j, \forall~ 1 \leq j \leq n. \] note - Given a set \(A \subseteq S\), the interior of A is defined as follows: \[ int(A) = \{ x \in S : \exists~ r > 0 \ni B_{d}(x,r) \subset A \} \] where \(d\) is the defined metric. Thus non-empty interior would just mean that \(int(A) \neq \phi\).