Finite Expectations
Abstract
We propose a simple statistical test for determining whether a sample of extreme values has been drawn from a probability distribution with finite or infinite expectation.
Full text
Finite Expectations Bryan S. Todd 14th December 2025 Abstract When it takes a long time for a random process to terminate it is tempting to wonder whether waiting is futile. We propose a statistical test to determine from a few outlier waiting times whether or not the expected waiting time is actually finite. Introduction Imagine that we have a process inside a black box that runs each time we press “start”. It takes a random amount of time to finish and the times are statistically independent. The process is guaranteed always to terminate eventually, but it can take an extremely long time to do so. After it has been running for a seemingly endless length of time it is tempting to wonder whether it is worth waiting. Past behaviour provides some insight into what to expect. We can measure statistics such as the average and median of the previous run times, but it would be helpful to know that the expected run time is actually finite. For example, a symmetric one-dimensional random walk always eventually returns to its starting point but, paradoxically, the expected time to do so is infinite. Let the run time expressed in suitable units (say hours) be random variable X. Let its probability density function be pwhose support is the non-negative reals. Thus ∫∞ 0 p(x)dx =1(1) We do not know what pis without looking inside the box, which we are forbidden to do. However, we would at least like to know whether the expected waiting time Ep[X]is finite. Ep[X]=∫∞ 0 x p(x)dx (2) Now if we can find a known probability density function qthat is a strict upper bound on p for all values of xgreater than some value s ∀x•x≥s⟹p(x)≤q(x)(3) and has finite expectation Eq[X] Eq[X]=∫∞ 0 x q(x)dx (4) then Ep[X]is finite too. Thus, in order to make progress we must assume that pis a member of an appropriate family of probability distributions range from long-tailed to short-tailed varieties. 1
Lomax Distribution A suitable type of probability distribution is the Lomax which has been widely used for modelling failure times [1]. It has two positive real parameters, aand b, and probability density function qwhere q(x)=b a(a x+a)b+1 (5) This has expectation Eq[X]=⎧ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎩ ∞if b≤1 a b−1if b>1(6) Thus for b>1the expectation is finite. Notice that given any a, the probability distribution q for which b=1has a longer tail than that of every distribution q′with corresponding parameters αand βwhere β<1. q(x)=a (x+a)2(7) q′(x)=βαβ (x+α)β+1(8) The denominator on the right-hand side of Equation 7 increases with xfaster than that of Equation 8 because it is raised to a strictly larger power. Therefore, for sufficiently large x, q(x)<q′(x). So the Lomax distribution that lies at the boundary between those with infinite expectation and those with finite expectation is that for which b=1. We refer to this as the Lomax1distribution; it has only one parameter a. Therefore if we can show that our sample is drawn from a probability distribution whose tail is strictly bounded above by Lomax1then its expectation is finite. An Example Suppose that L0is the list of six longest waiting times in decreasing order: L0=[10,5,4,3,2,1](9) The null hypothesis is that these data are the largest six of a random sample of some unspecified size from a Lomax distribution that has a tail at least as long as that of the Lomax1distribution (Equation 7). The latter has cumulative distribution function Q(x)=∫x 0 q(x)dx (10) =x x+a(11) If we can show that the sample is from a distribution with shorter tail than that of Lomax1 then its expectation is finite. 2
Normalization Our sample has been conditioned on the values being greater than some chosen threshold. Therefore, let us normalize it by subtracting and discarding every instance of the smallest value s. It follows that the remaining values have an unconditional Lomax1distribution. Let Xbe drawn from a Lomax1distribution with parameter a′and conditioned on X>s. For any positive real x, we have from the definition of conditional probability Pr(X−s≤x∣X>s)=Pr(X≤x+s∣X>s)(12) =Pr(X≤x+s, X >s) Pr(X>s)(13) =Q(x+s)−Q(s) 1−Q(s)(14) =(x+s x+s+a′−s s+a′)/(1−s s+a′)(15) =(x+s)(s+a′)−s(x+s+a′) a′(x+s+a′)(16) =a′x a′(x+s+a′)(17) =x x+a(18) where a=s+a′. So the normalized random variable X−sis equivalently drawn from a Lomax1 distribution with parameter s+a′. Thus after normalization our list L0becomes [9,4,3,2,1] and we have to show that this is a random sample from a distribution with shorter tail than that of Lomax1. Parameter Determination Let the list of nnormalized values (n=5in our example) in decreasing order be x0,x1, …xn−1. We assume that n>1, all the values are positive, and they are not all identical. Let us determine the maximum likelihood value for parameter a. This entails minimizing the total surprise S associated with the sample. S=− n−1 ∑ i=0 log q(xi)(19) = n−1 ∑ i=0(2 log(xi+a)−log a)(20) 3
Calculating the first and second derivatives we obtain dS da =1 a n−1 ∑ i=0 a−xi a+xi(21) d2 S da2=2 a(n−1 ∑ i=0 xi (xi+a)2)−1 a(dS da )(22) Minimum Sis easily computed because a=x0⟹dS da >0(23) a=xn−1⟹dS da <0(24) The second derivative is strictly positive at the turning point and so the solution is the unique minimum. Therefore optimum ais found rapidly and accurately by binary search for zero first derivative. For our example a=2.9229 to four decimal places and calculation took just 12 iterations. Parameter ais also the median of the distribution and close to that of the sample itself. A Test Statistic According to the null hypothesis, the cumulative Lomax probabilities Q(xi)for our sample should be drawn from the standard uniform distribution. If instead they appear to be clustered towards the mid-point (1 2) it implies that our sample is actually drawn from a distribution with a shorter tail. A useful measure of deviation from the mid-point is the log-odds. log Q(x) 1−Q(x)=log x−log a(25) The calculations for our example are Index (i) Value (xi) Probability Q(xi)log xi−log a 0 9 0.7549 1.1247 1 4 0.5778 0.3137 2 3 0.5065 0.0261 3 2 0.4063 -0.3794 4 1 0.2549 -1.0726 Since the direction of deviation from the mid-point is immaterial we are interested only in the absolute value of the log-odds. Our test statistic is thus v, suitably shifted and scaled: v=1 √nc n−1 ∑ i=0(∣log xi−log a∣−log 4)(26) 4
where the normalizing constant is c=π2 3−(log 4)2(27) ≈1.3681 (28) It follows that the v-statistic has zero mean under the null hypothesis in which Q(x)has a standard uniform distribution. The expectation of each term in the sum is thus ∫1 0(»»»»»»»log p 1−p»»»»»»»−log 4)dp =2∫1 1 2 log p 1−pdp −∫1 0 log 4 dp (29) =2 log 2 −log 4 (30) =0(31) It also follows that each term in the sum has variance c. ∫1 0(»»»»»»»log p 1−p»»»»»»»−log 4)2 dp =∫1 0(log p 1−p)2 dp −2 log 4∫1 0 »»»»»»»log p 1−p»»»»»»»dp +∫1 0(log 4)2 dp (32) =π2 3−(log 4)2(33) =c(34) Therefore after division by √cthe term has unit variance. Critical Values For large n, the v-statistic has a standard normal distribution (central limit theorem). Critical values are provided in Table 1. They were obtained by randomly generating 106samples from the Lomax1distribution. For our example, v=−1.5352 which is statistically significant at the p<0.05 level and allows us to reject the null hypothesis. We can be reasonably confident that list L0relates to a process with finite expected duration. For comparison consider list L1in which the longest duration is increased by a factor of ten with respect to that of L0. L1=[100,5,4,3,2,1](35) In this case v=−0.5706 and we cannot reject the null hypothesis even at the p<0.1level. List L1is thus consistent with a process that has infinite expected duration. More generally, suppose that our sample were actually drawn from the standard exponential distribution with probability density function p(x)=e−x(36) This has unit mean. Table 2 shows the probabilities of rejection of the null hypothesis if we draw samples of various sizes from this distribution. Compilation of the table entailed Monte 5
Table 1: Lower and upper critical values of vstatistic. n p≤0.1p≤0.05 p≤0.02 p≤0.01 4 -1.17, 1.33 -1.40, 1.83 -1.63, 2.44 -1.75, 2.87 5 -1.19, 1.33 -1.43, 1.81 -1.67, 2.39 -1.81, 2.80 6 -1.20, 1.33 -1.45, 1.80 -1.70, 2.37 -1.86, 2.77 7 -1.21, 1.33 -1.47, 1.80 -1.73, 2.35 -1.90, 2.74 8 -1.21, 1.32 -1.48, 1.78 -1.75, 2.33 -1.92, 2.72 9 -1.22, 1.32 -1.48, 1.78 -1.77, 2.31 -1.94, 2.70 10 -1.22, 1.32 -1.50, 1.77 -1.78, 2.31 -1.97, 2.68 12 -1.23, 1.32 -1.51, 1.76 -1.81, 2.29 -2.00, 2.65 15 -1.23, 1.32 -1.53, 1.75 -1.83, 2.26 -2.03, 2.61 25 -1.25, 1.31 -1.55, 1.73 -1.89, 2.22 -2.10, 2.55 50 -1.25, 1.30 -1.58, 1.70 -1.94, 2.17 -2.17, 2.48 100 -1.26, 1.30 -1.60, 1.69 -1.97, 2.14 -2.21, 2.43 200 -1.27, 1.29 -1.61, 1.67 -1.99, 2.12 -2.25, 2.41 400 -1.28, 1.29 -1.62, 1.67 -2.01, 2.10 -2.27, 2.38 >400 -1.28, 1.28 -1.64, 1.64 -2.05, 2.05 -2.33, 2.33 Carlo simulation with104iterations. The results show that a sample size of about 25 is needed in order to achieve equal odds of detecting that the expected duration of the process is finite at the p<0.05 level of significance. A sample size of a little over 50 is sufficient to achieve equal odds at the p<0.01 level of significance. Surprisingly there is still a reasonable chance (26%) of rejection at the p<0.05 level of significance with just four values. Since Lomax1lies on the boundary between those Lomax distributions with finite expectation and those with infinite, the test statistic also allows us to reject the null hypothesis that the sample was drawn from a distribution with finite expectation. For example consider L2: L2=[1000,100,4,3,2,1](37) For this sample v=1.8581. This exceeds the upper critical level at the p<0.05 level of significance. We can therefore conclude with reasonable confidence that this sample relates to a process with infinite expected duration. Thus is the process were currently still running and had been running for more than 1000 hours, we could reasonably conclude that it was not worth waiting any longer for it to terminate. More generally, suppose that our sample were actually from the standard log-Cauchy distribution with probability density function p(x)=1 πx(1+(log x)2)(38) This is a heavy-tailed distribution with infinite mean. Table 3 shows the probability of rejection of the null hypothesis if we draw samples of various sizes from this distribution. Here we 6
Table 2: Probability of rejecting null hypothesis that expectation is infinite at significance level pfor sample of size ndrawn from the exponential distribution. Size (n)p<0.10 p<0.05 p<0.02 p<0.01 4 0.39 0.27 0.15 0.09 5 0.41 0.27 0.15 0.10 6 0.41 0.28 0.16 0.10 7 0.43 0.29 0.17 0.10 8 0.43 0.30 0.17 0.11 9 0.45 0.32 0.19 0.12 10 0.45 0.32 0.19 0.12 12 0.48 0.35 0.21 0.14 15 0.54 0.40 0.26 0.18 25 0.61 0.50 0.35 0.27 50 0.75 0.67 0.57 0.49 100 0.82 0.78 0.74 0.70 200 0.85 0.83 0.81 0.79 400 0.86 0.85 0.83 0.82 compare the test statistic to the upper critical value and reject the null hypothesis in favour of the alternative that the distribution has infinite expectation. The results show that sample sizes as small as seven and nine achieve equal odds of detecting at the p<0.05 and p<0.01 levels of significance, respectively, that the process has infinite expected duration. There is still a reasonable chance (36%) of rejection at the p<0.05 level of significance with just four values. Generalized Extreme Value Distribution An alternative to using Lomax1is to model our data with the generalized extreme value (GEV) distribution. This describes the maxima of a set of independent and identically distributed random variables. We use the standardized variable s=(x−µ)/σwhere µis the location and σthe scale. The cumulative GEV distribution function is F(s)= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ e−e−s if c=0 e−(1−cs)1/c if c≠0and cs<1 0if c<0and cs≥1 1if c>0and cs≥1 (39) where cis the shape parameter. A variety of methods are available for rejecting the null hypotheses that c=0,c≤0, and c≥0[2]. However, we are interested here in testing whether c≤−1. This is because the mean of the GEV distribution is finite precisely when c>−1. 7
Table 3: Probability of rejecting null hypothesis that expectation is finite at significance level pfor sample of size ndrawn from the log-Cauchy distribution. Size (n)p<0.10 p<0.05 p<0.02 p<0.01 4 0.41 0.37 0.32 0.30 5 0.46 0.42 0.37 0.35 6 0.52 0.47 0.42 0.40 7 0.55 0.51 0.46 0.42 8 0.60 0.56 0.51 0.47 9 0.63 0.59 0.54 0.51 10 0.66 0.61 0.56 0.53 12 0.71 0.67 0.62 0.59 15 0.77 0.73 0.69 0.66 25 0.88 0.86 0.83 0.80 50 0.97 0.97 0.95 0.94 100 1.00 1.00 1.00 1.00 200 1.00 1.00 1.00 1.00 400 1.00 1.00 1.00 1.00 A maximum-likelihood estimator of the three parameters (µ,σ, and c) for any given sample has been implemented in the Python programming language by the SciPy group [3]. We therefore compared the estimate of the shape parameter cwith the v-statistic for discriminating between the exponential and log-Cauchy distributions. For each of three different sample sizes (5, 10, and 25) we generated 104samples from the exponential distribution and an identical number from the log-Cauchy distribution. From these samples we derived the receiver operating characteristic (ROC) curve for the v-statistic. We repeated the procedure for the c-statistic. Figures 1, 2, and 3 plot the corresponding ROC curves. The v-statistic appears to be a significantly better discriminant than the c-statistic for small samples. For a sample of size 25 there appears to be no significant difference. Tail Index Estimation The tail index estimator due to Hill [4] is another well established method of assessing the tail of the distribution from which a sample has been drawn. The index αis estimated by α=(( 1 n−1 n−2 ∑ i=0 log xi )−log xn−1)−1 (40) The mean of the underlying distribution is finite precisely when α>1. The ROC curves for the α-statistic were generated similarly to those of the other two statistics and are included 8
in Figures 1, 2, and 3. For the particular discrimination task the α-statistic performs less well than the v-statistic with all three sample sizes. Conclusion The apparent effectiveness of the v-statistic and the relative ease with which it can be reliably calculated suggest that it is a useful method for determining whether a sample is from a distribution with finite or infinite expectation. B.S. Todd 14/12/2025 Figure 1: ROC curves (n=5) for v-statistic,cstatistic, and α-statistic showing ability to discriminate exponential distribution from log-Cauchy. False positive rate True positive rate 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 9