In the definition of measures, regarding complements you say “elements that are in F but not in F” and I think the second F there should be E?
Couple questions:
- when you talk about the alternative construction of the Borel algebra, do the operations stop after the last one you mention, or do you also continue that process ad infinitum?
- regarding the sets not included in the Borel algebra, do you mean unmeasurable sets like the Vitali set or are there more measurable subsets like the cantor set that are also not included
2. You do need to keep going ad infinitum. I will clarify this.
3. It depends on whether you assume the axiom of Choice or not. If you do, then there is a strict containment of Borel, measureable, and then all subsets of the reals. The Cantor set is Borel though, and in fact any subset that you can write down remotely constructively will be Borel. This is because it is consistent with the negation of the axiom of Choice that *all* subsets of the real numbers are Borel, because it is consistent with the negation of the axiom of Choice that the reals are a countable union of countable sets!
Yep. Cohen's *Set Theory and the Continuum Hypothesis* has a nice (albeit difficult) exposition of this and similar results.
Less exotically, the axiom of determinacy—a (relatively) popular alternative to the axiom of choice—implies that all subsets of the real numbers are Lebesgue measurable (though not necessarily Borel).
"Then define fn(x)=1ℤ(2^-nx)—this is 1 if x is an integer multiple of 2^-n, and 0 otherwise." Should that be $2^n x$ (no minus sign) in the definition of $f_n(x)$?
> I personally prefer formalizing the Dirac delta in a different way, but that is a topic for another time.
Hm, really? I know there's "distributions"[0] aka "generalized functions", but to my mind it's fundamentally a measure. After all, distributions are even more general, so you can just always pass to that when you need to differentiate or whatever; but it's nicer to think of it as an honest measure when you don't need to do that, right? Since after all it is one. Much like 2 is fundamentally a whole number, rather than fundamentally a rational number, even though sometimes (indeed, often) you might need to invert it!
[0]This name annoys me to no end, since to my mind "distribution" should mean "probability distribution" which is to say a particular sort of measure, but instead in this context it doesn't!
As I said, if you want to do things like differentiate the Dirac delta function, then I think that working with distributions is more convenient. And I don't really need the Dirac delta function if I'm not going to be doing one of those things...
Alex, I will almost certainly eventually post lecture notes on complex analysis and measure theory. But my backlog is long---I wouldn't expect to get to the latter earlier than in a year. I might be able to get to the former in... six months? But that might be optimistic.
This is a very nice write-up! It doesn't get at one thing for me though (I've seen other treatments and it was always the same; this may not be an answerable question). I understand the shortcomings of Riemann integration and I understand the sort of axiomatic approach of measure theory. But I don't have a good feel for why the second approach works better than the first. The story as you tell it, and as it's often told, is: here's Riemann integration, this is where it sucks. Here's a totally different approach; look, it doesn't suck in that way. But I don't get a sense of why the new approach might be expected to not suck that way. The usual domain to co-domain flip doesn't explain it, or at least I'd need someone to connect those dots for me.
I'm not sure that I agree that the domain-to-codomain flip doesn't explain things. After all, once we slice up the codomain, we are still slicing up the domain, but we are slicing it up into *pre-images* of nice things like intervals. And even if you take something nice like a continuous function, the pre-image of a point can still be a Cantor set... or worse!
An example can perhaps help. Let f(x) = x², and g(x) = x²1_{R\Q}. The integral of these two functions on, say, [0,1] should agree, since they are equal almost everywhere. But “they are equal almost everywhere” is a statement about the coordination of domain and codomain, so we need to use a method for selecting the intervals \Delta x_i that sees this coordination between domain and codomain. When we set up the standard Riemann integral we ignore the codomain till we have already partitioned the domain into intervals, and thus our partitioning, and especially our selection of f(x_i*) is blind to any coordination between domain and co-domain. We are thus incapable of using the fact that f and g agree almost everywhere when setting up the finite sums, and we may pick x_i* such that g [maps] them all to zero. (We could pick them randomly with a uniform probability distribution over the interval, and then we almost surely get the Lebesgue Integral—but picking them randomly requires measure theory.) What the Lebesgue Integral allows us to do is look at the values of the codomain up front, and so to coordinate the treatment of domain and codomain.
[edit: “maps” added as its omission was an editorial mistake.]
Lebesgue measure being the complete version of Borel measure is a big deal. Glad to see your footnote on it.
When higher dimensional analogs of Lebesgue measure are defined, we can't just take the product measure of Lebesgue measures. We have to also take the completion each time. Speaking of which, being able to express iterated integration with a single Lebesgue integration with a different measure is very useful. The Fubini and Tonelli theorems are some of my favorites.
Excellent followup to your articles on functional analysis, thank you! Very helpful to someone starting to study this material!
In the definition of measures, regarding complements you say “elements that are in F but not in F” and I think the second F there should be E?
Couple questions:
- when you talk about the alternative construction of the Borel algebra, do the operations stop after the last one you mention, or do you also continue that process ad infinitum?
- regarding the sets not included in the Borel algebra, do you mean unmeasurable sets like the Vitali set or are there more measurable subsets like the cantor set that are also not included
In order:
1. Yes, and I will fix this in a moment.
2. You do need to keep going ad infinitum. I will clarify this.
3. It depends on whether you assume the axiom of Choice or not. If you do, then there is a strict containment of Borel, measureable, and then all subsets of the reals. The Cantor set is Borel though, and in fact any subset that you can write down remotely constructively will be Borel. This is because it is consistent with the negation of the axiom of Choice that *all* subsets of the real numbers are Borel, because it is consistent with the negation of the axiom of Choice that the reals are a countable union of countable sets!
> it is consistent with the negation of the axiom of Choice that the reals are a countable union of countable sets
huh, that's news to me!
Yep. Cohen's *Set Theory and the Continuum Hypothesis* has a nice (albeit difficult) exposition of this and similar results.
Less exotically, the axiom of determinacy—a (relatively) popular alternative to the axiom of choice—implies that all subsets of the real numbers are Lebesgue measurable (though not necessarily Borel).
"Then define fn(x)=1ℤ(2^-nx)—this is 1 if x is an integer multiple of 2^-n, and 0 otherwise." Should that be $2^n x$ (no minus sign) in the definition of $f_n(x)$?
Indeed! I will fix it in a moment. Thank you.
> I personally prefer formalizing the Dirac delta in a different way, but that is a topic for another time.
Hm, really? I know there's "distributions"[0] aka "generalized functions", but to my mind it's fundamentally a measure. After all, distributions are even more general, so you can just always pass to that when you need to differentiate or whatever; but it's nicer to think of it as an honest measure when you don't need to do that, right? Since after all it is one. Much like 2 is fundamentally a whole number, rather than fundamentally a rational number, even though sometimes (indeed, often) you might need to invert it!
[0]This name annoys me to no end, since to my mind "distribution" should mean "probability distribution" which is to say a particular sort of measure, but instead in this context it doesn't!
As I said, if you want to do things like differentiate the Dirac delta function, then I think that working with distributions is more convenient. And I don't really need the Dirac delta function if I'm not going to be doing one of those things...
Hi, do you or can provide university undergraduate level notes on Measure Theory? If so I'm happy to pay.
Similarly, can you also provide worksheets and annotated solutions also?
I also like your material on Complex analysis linked to fluid mechanics. The same question above, please?
Best,
Alex
Alex, I will almost certainly eventually post lecture notes on complex analysis and measure theory. But my backlog is long---I wouldn't expect to get to the latter earlier than in a year. I might be able to get to the former in... six months? But that might be optimistic.
This is a very nice write-up! It doesn't get at one thing for me though (I've seen other treatments and it was always the same; this may not be an answerable question). I understand the shortcomings of Riemann integration and I understand the sort of axiomatic approach of measure theory. But I don't have a good feel for why the second approach works better than the first. The story as you tell it, and as it's often told, is: here's Riemann integration, this is where it sucks. Here's a totally different approach; look, it doesn't suck in that way. But I don't get a sense of why the new approach might be expected to not suck that way. The usual domain to co-domain flip doesn't explain it, or at least I'd need someone to connect those dots for me.
I'm not sure that I agree that the domain-to-codomain flip doesn't explain things. After all, once we slice up the codomain, we are still slicing up the domain, but we are slicing it up into *pre-images* of nice things like intervals. And even if you take something nice like a continuous function, the pre-image of a point can still be a Cantor set... or worse!
An example can perhaps help. Let f(x) = x², and g(x) = x²1_{R\Q}. The integral of these two functions on, say, [0,1] should agree, since they are equal almost everywhere. But “they are equal almost everywhere” is a statement about the coordination of domain and codomain, so we need to use a method for selecting the intervals \Delta x_i that sees this coordination between domain and codomain. When we set up the standard Riemann integral we ignore the codomain till we have already partitioned the domain into intervals, and thus our partitioning, and especially our selection of f(x_i*) is blind to any coordination between domain and co-domain. We are thus incapable of using the fact that f and g agree almost everywhere when setting up the finite sums, and we may pick x_i* such that g [maps] them all to zero. (We could pick them randomly with a uniform probability distribution over the interval, and then we almost surely get the Lebesgue Integral—but picking them randomly requires measure theory.) What the Lebesgue Integral allows us to do is look at the values of the codomain up front, and so to coordinate the treatment of domain and codomain.
[edit: “maps” added as its omission was an editorial mistake.]
Lebesgue measure being the complete version of Borel measure is a big deal. Glad to see your footnote on it.
When higher dimensional analogs of Lebesgue measure are defined, we can't just take the product measure of Lebesgue measures. We have to also take the completion each time. Speaking of which, being able to express iterated integration with a single Lebesgue integration with a different measure is very useful. The Fubini and Tonelli theorems are some of my favorites.