This section is devoted to proving the entropy ergodic theorem for the special case of stationary ergodic sources. The result was originally proved by Breiman [19]. The original proof first used the martingale convergence theorem to infer the convergence of conditional probabilities of the formm(X0|X−1, X−2,· · ·, X−k)
tom(X0|X−1, X−2,· · ·). This result was combined with an an extended form of the ergodic theorem stating that ifgk→g ask→ ∞and ifgk isL1-dominated
(supk|gk|is inL1), then 1/nPnk=0−1gkTk has the same limit as 1/n
Pn−1
k=0gTk. Combining these facts yields that that
1 nlnm(X n) = 1 n n−1 X k=0 lnm(Xk|Xk) = 1 n n−1 X k=0 lnm(X0|X−kk)Tk
has the same limit as 1
n
nX−1
k=0
lnm(X0|X−1, X−2,· · ·)Tk which, from the usual ergodic theorem, is the expectation
E(lnm(X0|X−)≡E(lnm(X0|X−1, X−2,· · ·)).
As suggested at the end of the preceeding chapter, this should be minus the conditional entropy H(X0|X−1, X−2,· · ·) which in turn should be the entropy rate ¯HX. This approach has three shortcomings: it requires a result from mar-
tingale theory which has not been proved here or in the companion volume [50], it requires an extended ergodic theorem which has similarly not been proved here, and it requires a more advanced definition of entropy which has not yet been introduced. Another approach is the sandwich proof of Algoet and Cover [7]. They show without using martingale theory or the extended ergodic theo- rem that 1/nPni=0−1lnm(X0|X−ii)Ti is asymptotically sandwiched between the
entropy rate of akth order Markov approximation: 1 n nX−1 i=k lnm(X0|X−kk)Ti→Em[lnm(X0|X−kk)] =−H(X0|X−kk) and 1 n n−1 X i=k lnm(X0|X−1, X−2,· · ·)Ti→Em[lnm(X0|X1,· · ·)] =−H(X0|X−1, X−2,· · ·).
By showing that these two limits are arbitrarily close ask → ∞, the result is proved. The drawback of this approach for present purposes is that again the
3.2. STATIONARY ERGODIC SOURCES 51 more advanced notion of conditional entropy given the infinite past is required. Algoet and Cover’s proof that the above two entropies are asymptotically close involves martingale theory, but this can be avoided by using Corollary 5.2.4 as will be seen.
The result can, however, be proved without martingale theory, the extended ergodic theorem, or advanced notions of entropy using the approach of Ornstein and Weiss [117], which is the approach we shall take in this chapter. In a later chapter when the entropy ergodic theorem is generalized to nonfinite alphabets and the convergence of entropy and information densities is proved, the sandwich approach will be used since the appropriate general definitions of entropy will have been developed and the necessary side results will have been proved.
Lemma 3.2.1: Given a finite alphabet source {Xn} with a stationary er-
godic distributionm, we have that lim
n→∞
−lnm(Xn)
n =h; m−a.e.,
whereh(x) is the invariant function defined by
h(x) = ¯Hm(X). Proof: Define hn(x) =−lnm(Xn)(x) =−lnm(xn) and h(x) = lim inf n→∞ 1 nhn(x) = lim infn→∞ −lnm(xn) n .
Sincem((x0,· · ·, xn−1))≤m((x1,· · ·, xn−1)), we have that
hn(x)≥hn−1(T x).
Dividing by n and taking the limit infimum of both sides shows that h(x) ≥
h(T x). Since the n−1h
n are nonnegative and uniformly integrable (Lemma
2.3.6), we can use Fatou’s lemma to deduce that h and hence also hT are integrable with respect tom. Integrating with respect to the stationary measure
myields Z
dm(x)h(x) =
Z
dm(x)h(T x) which can only be true if
h(x) =h(T x);m−a.e.,
that is, if his an invariant function with m-probability one. If h is invariant almost everywhere, however, it must be a constant with probability one since
m is ergodic (Lemma 6.7.1 of [50]). Since it has a finite integral (bounded by ¯
We now proceed with steps that resemble those of the proof of the ergodic theorem in Section 7.2 of [50]. Fix ² >0. We also choose for later use aδ >0 small enough to have the following properties: IfA is the alphabet of X0 and ||A|| is the finite cardinality of the alphabet, then
δln||A||< ², (3.7)
and
−δlnδ−(1−δ) ln(1−δ)≡h2(δ)< ². (3.8) The latter property is possible since h2(δ)→0 asδ→0.
Define the random variable n(x) to be the smallest integer n for which
n−1hn(x)≤h+². By definition of the limit infimum there must be infinitely manyn for which this is true and hence n(x) is everywhere finite. Define the set of “bad” sequences by B = {x : n(x) > N} where N is chosen so large that m(B)< δ/2. Still mimicking the proof of the ergodic theorem, we define a bounded modification ofn(x) by ˜ n(x) = ½ n(x) 1 x6∈B x∈B
so that ˜n(x)≤N for allx∈Bc. We now parse the sequence into variable-length blocks. Iteratively definenk(x) by
n0(x) = 0 n1(x) = ˜n(x) n2(x) =n1(x) + ˜n(Tn1(x)x) =n1(x) +l1(x) .. . nk+1(x) =nk(x) + ˜n(Tnk(x)x) =nk(x) +lk(x),
wherelk(x) is the length of thekth block:
lk(x) = ˜n(Tnk(x)x).
We have parsed a long sequence xL = (x0,· · ·, xL−1), where L >> N, into blocksxnk(x),· · ·, xnk+1(x)−1=x
lk(x)
nk(x)which begin at timenk(x) and have lengthlk(x) fork= 0,1,· · ·. We refer to this parsing as theblock decomposition
of a sequence. The kth block, which begins at time nk(x), must either have
sample entropy satisfying
−lnm(xlk(x)
nk(x))
lk(x) ≤h+² (3.9)
or, equivalently, probability at least
3.2. STATIONARY ERGODIC SOURCES 53 or it must consist of only a single symbol. Blocks having length 1 (lk = 1)
could have the correct sample entropy, that is, −lnm(x1
nk(x))
1 ≤
¯
h+²,
or they could be bad in the sense that they are the first symbol of a sequence withn > N; that is,
n(Tnk(x)x)> N, or, equivalently,
Tnk(x)x∈B.
Except for these bad symbols, each of the blocks by construction will have a probability which satisfies the above bound.
Define for nonnegative integersnand positive integerslthe sets
S(n, l) ={x:m(Xnl(x))≥e−l(h+²)},
that is, the collection of infinite sequences for which (3.2.2) and (3.2.3) hold for a block starting at n and having lengthl. Observe that for such blocks there cannot be more than el(h+²) distinct l-tuples for which the bound holds (lest the probabilities sum to something greater than 1). In symbols this is
||S(n, l)|| ≤el(h+²). (3.11) The ergodic theorem will imply that there cannot be too many single symbol blocks with n(Tnk(x)x) > N because the event has small probability. These facts will be essential to the proof.
Even though we write ˜n(x) as a function of the entire infinite sequence, we can determine its value by observing only the prefixxN ofxsince either there is ann≤N for which n−1lnm(xn)≤h+²or there is not. Hence there is a function ˆn(xN) such that ˜n(x) = ˆn(xN). Define the finite length sequence event
C={xN : ˆn(xN) = 1 and−lnm(x1)> h+²}, that is,Cis the collection of all
N-tuplesxN that are prefixes of bad infinite sequences, sequences xfor which
n(x)> N. Thus in particular,
x∈B if and only ifxN ∈C. (3.12) Now recall that we parse sequences of lengthL >> N and define the setGL
of “good”L-tuples by GL={xL : 1 L−N L−XN−1 i=0 1C(xNi )≤δ},
that is,GLis the collection of allL-tuples which have fewer thanδ(L−N)≤δL
the ergodic theorem for stationary ergodic sources we know thatm-a.e. we get anxfor which lim n→∞ 1 n nX−1 i=0 1C(xNi ) = limn→∞ 1 n nX−1 i=0 1B(Tix) =m(B)≤ δ 2. (3.13) From the definition of a limit, this means that with probability 1 we get anx
for which there is anL0=L0(x) such that 1
L−N
L−XN−1
i=0
1C(xNi )≤δ; for allL > L0. (3.14) This follows simply because if the limit is less thanδ/2, there must be anL0so large that for largerLthe time average is at least no greater than 2δ/2 =δ. We can restate (3.14) as follows: with probability 1 we get anxfor whichxL∈GL
for all but a finite number ofL. Stating this in negative fashion, we have one of the key properties required by the proof: IfxL∈GLfor all but a finite number
ofL, thenxL cannot be in the complementGcLinfinitely often, that is,
m(x:xL∈GcL i.o.) = 0. (3.15)
We now change tack to develop another key result for the proof. For each
L we bounded above the cardinality ||GL|| of the set of good L-tuples. By construction there are no more than δLbad symbols in an L-tuple inGL and these can occur in any of at most
X k≤δL µ L k ¶ ≤eh2(δ)L (3.16)
places, where we have used Lemma 2.3.5. Eq. (3.16) provides an upper bound on the number of ways that a sequence inGLcan be parsed by the given
rules. The bad symbols and the final N symbols in the L-tuple can take on any of the||A||different values in the alphabet. Eq. (3.11) bounds the number of finite length sequences that can occur in each of the remaining blocks and hence for any given block decomposition, the number of ways that the remaining blocks blocks can be filled is bounded above by
Y
k:Tnk(x)x6∈B
elk(x)(h+²)=ePklk(x)(h+²)
=eL(h+²), (3.17)
regardless of the details of the parsing. Combining these bounds we have that ||GL|| ≤eh2(δ)L× ||A||δL× ||A||N ×eL(h+²)=eh2(δ)L+(δL+N) ln||A||+L(h+²)
or
3.2. STATIONARY ERGODIC SOURCES 55 Sinceδsatisfies (3.7)–(3.8), we can chooseL1large enough so thatNln||A||/L1≤
²and thereby obtain
||GL|| ≤eL(h+4²); L≥L1. (3.18) This bound provides the second key result in the proof of the lemma. We now combine (3.18) and (3.15) to complete the proof.
Let BL denote a collection of L-tuples that are bad in the sense of having
too large a sample entropy or, equivalently, too small a probability; that is if
xL∈BL, then
m(xL)≤e−L(h+5²)
or, equivalently, for anyxwith prefixxL hL(x)≥h+ 5².
The upper bound on||GL|| provides a bound on the probability ofBLTGL:
m(BL \ GL) = X xL∈BLTGL m(xL)≤ X xL∈GL e−L(h+5²) ≤ ||GL||e−L(h+5²)≤e−²L.
Recall now that the above bound is true for a fixed² >0 and for all L≥L1. Thus ∞ X L=1 m(BL \ GL) = LX1−1 L=1 m(BL \ GL) + ∞ X L=L1 m(BL \ GL) ≤L1+ ∞ X L=L1 e−²L<∞
and hence from the Borel-Cantelli lemma (Lemma 4.6.3 of [50]) m(x : xL ∈ BLTGL i.o.) = 0. We also have from (3.2.8), however, that m(x : xL ∈
GcL i.o. ) = 0 and hence xL ∈ GL for all but a finite number of L. Thus
xL ∈ BL i.o. if and only if xL ∈ BLTGL i.o. As this latter event has zero
probability, we have shown thatm(x:xL∈BLi.o.) = 0 and hence
lim sup
L→∞ hL(x)≤h+ 5².
Since ²is arbitrary we have proved that the limit supremum of the sample entropy−n−1lnm(Xn) is less than or equal to the limit infimum and therefore that the limit exists and hence withm-probability 1
lim
n→∞
−lnm(Xn)
n =h. (3.19)
Since the terms on the left in (3.19) are uniformly integrable from Lemma 2.3.6, we can integrate to the limit and apply Lemma 2.4.1 to find that
h= lim n→∞ Z dm(x)−lnm(X n(x)) n = ¯Hm(X),
which completes the proof of the lemma and hence also proves Theorem 3.1.1 for the special case of stationary ergodic measures. 2