Showing posts with label Epsilon. Show all posts
Showing posts with label Epsilon. Show all posts

Wednesday, August 29, 2012

Dead Worshipers Tell No Tales

[caption id="attachment_1371" align="aligncenter" width="355"]Cicero Cicero, Roman philosopher and statesman.[/caption]

Cicero recounted the following story:

A non-believer named Diagoras was shown painted tablets depicting some worshipers who had prayed and then survived a shipwreck.

Diagoras was not impressed. He wondered about the portraits of those worshipers who had prayed but still drowned.

Apparently, drowned worshipers told no tales.

Think of this story the next time

  • you think the supermarket queue you join always moves slowly. Is it because you tend not to remember the times when queues move fast?

  • someone tells you that dropping out of university is a smart career move ever since Bill Gates left Harvard to found Microsoft. People and the media do not seem to talk a lot about university dropouts who end up with mediocre careers

  • an astrological prediction seems to come true

  • you meet a rude person

  • you read the story of how the self-exiled New Castle mining magnate Nathan Tinkler literally punted his house on a risky mining venture and amassed a fortune within the span of a few years

Tuesday, August 28, 2012

Meaning of Life and Color of Jealousy

[caption id="attachment_1350" align="aligncenter" width="500"]Richard Dawkins Richard Dawkins, biologist, author and atheist.[/caption]

Who has not quested and pondered on the meaning of life?

My own misguided, adolescent quest for the meaning of life led me to the ashram of a yogi, the sermons of J. Krishnamurti, the glib utterances of the charlatan Osho Rajneesh, a university course in Shada Darshana (the Six Schools of Indian Philosophy), and the popular works of Bertrand Russell and other Western thinkers.

One of these thinkers was an eccentric Austrian named Ludwig Wittgenstein, who had become a cult figure by the time he died in Cambridge, England at the age of 62 in 1951.

When Wittgenstein first came to Cambridge in 1911 to study the foundations of mathematics with Russell, his lordship could not decide if Wittgenstein was a crank or a genius but eventually settled for the latter.

In 1929, Wittgenstein returned to Cambridge for a Ph D, occasioning the economist Keynes’ letter to his wife in which he noted: “Well, God has arrived. I met him on the 5:15 train.”

Wittgenstein’s philosophy of logical positivism inspired a group of thinkers called the Vienna Circle. Its central tenet was distilled in the aphorism “The meaning of a sentence is its method of verification.”

According to this "method of verification", a sentence had to satisfy one of the following conditions to be valid:

1. True by definition. “A triangle has three sides.”
2. Empirically verifiable: “Mt. Everest is taller than Mt. Druit.”

Conversely, the following sentences are not valid as they are neither true by definition nor empirically verifiable:

1. A thing of beauty is a joy forever
2. Jesus was born of immaculate conception
3. God is great
4. In my End is my Beginning

Logical positivism had a grand ambition: To smash metaphysics and, with it, all the “ultimate” questions. It ran into a familiar epistemological hurdle: itself.

By the yardstick of logical positivism itself, the sentence “The meaning of a sentence is its method of verification” is a nonsense. Perhaps, this is why, despite Wittgenstein’s modest belief that he had solved all philosophical problems by analysing language, we keep asking the ultimate questions.

Fast forward to 2012 and an epiphany of sorts!

A recent ABC TV’s Q & A panel discussion pitted Cardinal George Pell, Archbishop of Sydney, against the British evolutionary biologist and outspoken atheist Richard Dawkins.

At one point, Dawkins, arguing that life has no meaning beyond itself, said that just because you can ask a question does not mean it is a valid question.

“What is the color of jealousy?” is one such invalid question, according to the renowned biologist and author of The Selfish Gene and The God Delusion.

Dawkins’ argument resonated deeply with me. I wish I had come across such an insight during my adolescent meanderings. Then, perchance, if not abandon my futile search for the miraculous, I might at least not have given up the study of calculus in favor of canard.

Sunday, August 26, 2012

Ambushed by Outlier



On Saturday morning, I reached a radiology in a neighbouring suburb for a dental x-ray at 9:30, hoping to wrap up the visit in 15 minutes. I ended up waiting for an hour - twice punctuated by my inquiries about expected waiting time  - before an amiable, bespectacled radiologist materialised and led me to the x-ray room.

When asked why I had to wait so long when this type of x-ray should be a fairly short affair, he only murmured, "I don't know. It was a misunderstanding."

Back at the main waiting room, I handed the slip of paper that the radiologist had asked me to hand in to the reception, repeating to one of the receptionists what the radiologist had told me, that I had waited an hour as a result of a "misunderstanding".

The receptionist conferred with a colleague to her right and said something like, "You'd to wait as long as you did because the x-ray room has another machine that was being used by another patient for a procedure that takes time."

"You're unlucky," she added as she apologized and wished me a great day.

***


To begin with, what us lay folks call being "unlucky", experts from quantitative disciplines such as economics, social sciences and statistics may call being victims of "outliers", which are nothing but rare, out-of-the-ordinary events.

Examples of outliers include winning a lottery, the volcanic eruption of Mt. Vesuvius in 79 AD that buried Pompeii, air crashes in developed nations, Australia's national rugby team Wallabies posting a win against New Zealand's All Blacks ... and, apparently, waiting for an hour to get one's "missing/crowded" teeth x-rayed.

Since my dentist indicated that I will have to make a number of visits to the radiology, I am mainly concerned with two questions.

First, what is the probability of my being "unlucky" in the same radiology's waiting room in my next visit?

For the sake of argument, let's assume that, on average, 1 in 100 dental x-ray patients on a Saturday morning gets "unlucky" in that particular radiology, ending up waiting an hour. This means the probability of getting unlucky is 1 percent or 0.01.

From this, assuming that two visits to the radiology for dental x-rays are independent, i.e. the first visit does not influence the waiting time of the second visit, the probability that I will again be "unlucky" on the second visit is still 0.01.

A slightly different question is this. What is the probability that a patient like myself will be "unlucky" in two consecutive visits? Intuition tells us that it has to be lower than 1 in 100. In fact, it is 0.01 x 0.01 = 0.0001 or 1 in 10,000.

On the other hand, I may reason that since I was already unlucky in my last visit to the radiology, the probability of getting unlucky in the next visit is lower than 1 in 100. If I reason like this, assuming that the previous visit has no effect on the waiting time of my next visit, I have just fallen victim to the gambler's fallacy.

The next question I am interested is this. If I am not unlucky in my next visit to the radiology, what should I expect the waiting time to be? To put it another way, what is the average waiting time for a dental x-ray on a Saturday morning?

Assuming that waiting times are normally distributed and a waiting time of 1 hour lies in the upper 1 percent of the distribution, i.e. 99 percent of waiting times are less than 1 hour, the average waiting time is around 35 minutes, with the standard deviation of around 11 minutes.

This means that my initial hope of wrapping up the visit in 15 minutes was a forlorn hope. From the properties of the normal distribution, it can be estimated that there is less than 3 percent likelihood of a waiting time to be 15 minutes or less.

Tuesday, July 17, 2012

Luhn Algorithm in Teradata SQL

Luhn algorithm is used, among others, to calculate the checksum digit of credit cards and mobile handset IMEIs. The following is my attempt to implement this algorithm in Teradata sql. It flags each IMEI as valid or not. Needless to say, IMEIs would typically be read from a table rather than hard-coded as in this example.
SELECT dt3.IMEI
,CASE
WHEN (dt3.dig1 + dt3.dig2 + dt3.dig3 + dt3.dig4
+ dt3.dig5 + dt3.dig6 + dt3.dig7
+ dt3.dig8 + dt3.dig9 + dt3.dig10 + dt3.dig11
+ dt3.dig12 + dt3.dig13
+ dt3.dig14 + dt3.dig15) MOD 10 = 0 THEN 'Y'
ELSE 'N'
END AS VALID_IMEI
FROM
(
SELECT dt2.IMEI
,dt2.dig1
,CASE
WHEN dt2.dig2 = 0 THEN 0
WHEN dt2.dig2 MOD 9 = 0 THEN 9
ELSE dt2.dig2 MOD 9
END AS dig2
,dt2.dig3
,CASE
WHEN dt2.dig4 = 0 THEN 0
WHEN dt2.dig4 MOD 9 = 0 THEN 9
ELSE dt2.dig4 MOD 9
END AS dig4
,dt2.dig5
,CASE
WHEN dt2.dig6 = 0 THEN 0
WHEN dt2.dig6 MOD 9 = 0 THEN 9
ELSE dt2.dig6 MOD 9
END AS dig6
,dt2.dig7
,CASE
WHEN dt2.dig8 = 0 THEN 0
WHEN dt2.dig8 MOD 9 = 0 THEN 9
ELSE dt2.dig8 MOD 9
END AS dig8
,dt2.dig9
,CASE
WHEN dt2.dig10 = 0 THEN 0
WHEN dt2.dig10 MOD 9 = 0 THEN 9
ELSE dt2.dig10 MOD 9
END AS dig10
,dt2.dig11
,CASE
WHEN dt2.dig12 = 0 THEN 0
WHEN dt2.dig12 MOD 9 = 0 THEN 9
ELSE dt2.dig12 MOD 9
END AS dig12
,dt2.dig13
,CASE
WHEN dt2.dig14 = 0 THEN 0
WHEN dt2.dig14 MOD 9 = 0 THEN 9
ELSE dt2.dig14 MOD 9
END AS dig14
,dt2.dig15
FROM
(
SELECT dt1.IMEI
,SUBSTR(dt1.IMEI, 1, 1) AS dig1
,SUBSTR(dt1.IMEI, 2, 1) * 2 AS dig2
,SUBSTR(dt1.IMEI, 3, 1) AS dig3
,SUBSTR(dt1.IMEI, 4, 1) * 2 AS dig4
,SUBSTR(dt1.IMEI, 5, 1) AS dig5
,SUBSTR(dt1.IMEI, 6, 1) * 2 AS dig6
,SUBSTR(dt1.IMEI, 7, 1) AS dig7
,SUBSTR(dt1.IMEI, 8, 1) * 2 AS dig8
,SUBSTR(dt1.IMEI, 9, 1) AS dig9
,SUBSTR(dt1.IMEI, 10, 1) * 2 AS dig10
,SUBSTR(dt1.IMEI, 11, 1) AS dig11
,SUBSTR(dt1.IMEI, 12, 1) * 2 AS dig12
,SUBSTR(dt1.IMEI, 13, 1) AS dig13
,SUBSTR(dt1.IMEI, 14, 1) * 2 AS dig14
,SUBSTR(dt1.IMEI, 15, 1) AS dig15
FROM
(
SELECT '999999999999999' AS IMEI
FROM SYS_CALENDAR.CALENDAR
WHERE calendar_date = CURRENT_DATE

UNION

SELECT '352651010278244' AS IMEI
FROM SYS_CALENDAR.CALENDAR
WHERE calendar_date = CURRENT_DATE

) AS dt1
) AS dt2
) AS dt3
;

It returns the following answerset:

Sunday, July 1, 2012

Summation Rules

Here is a list of summation rules, with 'k' denoting a constant:

  1. \[\sum_{i=1}^{n}{(x_i + y_i)} = \sum_{i={1}}^{n}{x_i} + \sum_{i=1}^{n}{y_i}\]

  2. \[\sum_{i=1}^{n}{(x_i - y_i)} = \sum_{i=1}^{n}{x_i}  - \sum_{i=1}^{n}{y_i}\]

  3. \[\sum_{i=1}^{n}{x_iy_i} \neq \sum_{i=1}^{n}{x_i} \times \sum_{i=1}^{n}{y_i}\]

  4. \[\sum_{i=1}^{n}{x_i}^2 \neq (\sum_{i=1}^{n}{x_i})^2\]

  5. \[\sum_{i=1}^{n}{k} = nk\]

  6. \[\sum_{i=1}^n{(x_i + k)} = \sum_{i=1}^{n}{x_i} + \sum_{i=1}^{n}{k} = \sum_{i=1}^{n}{x_i}+nk\]

  7. \[\sum_{i=1}^{n}{(x_i-k)} = \sum_{i=1}^{n}{x_i} - nk\]

Simulating Central Limit Theorem

In this post, the Central Limit Theorem (CLT) will be simulated using Python, SciPy and matplotlib. The CLT gives the following two theorems:

Theorem 1: If the sampled population is normally distributed with population mean = $\mu$ and standard deviation = $\sigma$, then for any sample size n, sampling distribution of the mean for simple random samples is normally distributed, with mean ($\mu_\overline{x}$) = $\mu$ and standard deviation ($\sigma_\overline{x}$) = $\frac{\sigma}{\sqrt {n}}$.

Theorem 2: For large sample sizes $(n\geq 30)$, even if the sampled population is not normally distributed, the sampling distribution of the mean for simple random samples is approximately normally distributed, with mean ($\mu_\overline{x}$) = $\mu$ and standard deviation ($\sigma_\overline{x}$) = $\frac{\sigma}{\sqrt {n}}$.

The standard deviation of sampling mean ($\sigma_\overline{x}$) is also known as the standard error of mean, standard error of estimate or simply as standard error as the sampling standard deviation gives the average deviation of the sample means from the actual population mean.

The following Python script simulates Theorem 1.
#----------------------------------------------------------------------------
# By Ram Limbu @ ramlimbu.com
# Copyright 2012 Ram Limbu
# License: GNU GPLv3 http://www.gnu.org/licenses/gpl.html
#----------------------------------------------------------------------------

# import required packages
import random
import matplotlib.pylab as pylb

def plotDist(t, val='Values'):
'''plot histogram of distribution'''
pylb.hist(t, bins=50, color='R')
pylb.title('Central Limit Theorem Simulation')
pylb.ylabel('frequency')
pylb.xlabel(val)
pylb.show()

def simulateSampDist(t_pop):
'''simulate sampling distributions'''
samp_sizes = (5,15,25)
t_samp_mean = []
for i in range(0,len(samp_sizes)):
for j in range(0,1000):
t_samp_mean.append(pylb.mean(random.sample(t_pop, samp_sizes[i])))

# plot the population distribution
samp_mean = round(pylb.mean(t_samp_mean), 2)
samp_stddev = round(pylb.std(t_samp_mean), 2)
val = 'mean = ' + str(samp_mean) + ' stddev = ' + str(samp_stddev) \
+ ' n=' + str(samp_sizes[i])
plotDist(t_samp_mean, val)

def main():
'''simulate central limit theorem'''

# generate a population of 10,000 normally distributed random numbers
# with mean = 50 and standard deviation = 10
t_pop = []
mu = 50
sigma = 10
pop_size = 10000

for i in range(0,pop_size):
t_pop.append(random.gauss(mu, sigma))

# plot a histogram of the population
plotDist(t_pop)

# simulate sampling distributions by drawing and replacing
# samples of various sizes from this population
simulateSampDist(t_pop)

if __name__ == '__main__':
main()

First, it creates a normally distributed population of 10,000 pseudo-random numbers with $\mu$ = 50 and $\sigma$ = 10. Then, it takes a sample of size 5, calculates its mean and appends it to a list, repeating this process 1,000 times. Finally, it plots the histogram of the sample means. Then, it repeats the whole sampling process with samples of size 15 and 25.

Histogram of normally distributed population of 10,000 random numbers.


The following histogram shows the distribution of sampling means of size 5. It has mean of 49.92, which is close to the population mean of 50, and the standard error of 4.51. The latter figure decreases as the sample size increases.


The next two figures show the distribution of means of samples of size 15 and 25. Note that in each case, the distribution has a mean close to 50, with the standard error decreasing as the sample size increases.

histogram of sample means of size 15


 


 The following Python script simulates Theorem 2, generating means, standard errors and histograms of samples of size 30, 50 and 100 from a population of exponentially distributed pseudo-random numbers with $\mu$=50.
#----------------------------------------------------------------------------
# By Ram Limbu @ ramlimbu.com
# Copyright 2012 Ram Limbu
# License: GNU GPLv3 http://www.gnu.org/licenses/gpl.html
#----------------------------------------------------------------------------

# import required packages
import random
import matplotlib.pylab as pylb

def plotDist(t, val='Values'):
'''plot histogram of distribution'''
pylb.hist(t, bins=30, color='R')
pylb.title('Central Limit Theorem Simulation')
pylb.ylabel('frequency')
pylb.xlabel(val)
pylb.show()

def simulateSampDist(t_pop):
'''simulate sampling distributions'''
samp_sizes = (30,50,100)
t_samp_mean = []
for i in range(0,len(samp_sizes)):
for j in range(0,1000):
t_samp_mean.append(pylb.mean(random.sample(t_pop, samp_sizes[i])))

# plot the population distribution
samp_mean = round(pylb.mean(t_samp_mean), 2)
samp_stddev = round(pylb.std(t_samp_mean), 2)
val = 'mean = ' + str(samp_mean) + ' stddev = ' + str(samp_stddev) \
+ ' n=' + str(samp_sizes[i])
plotDist(t_samp_mean, val)

def main():
'''simulate central limit theorem'''

# generate a population of 10,000 exponentially distributed random numbers
# with mean = 10
t_pop = []
mu = 50.00
pop_size = 10000

for i in range(0,pop_size):
t_pop.append(random.expovariate(1/mu))

# plot a histogram of the population
plotDist(t_pop)

# simulate sampling distributions by drawing and replacing
# samples of various sizes from this population
simulateSampDist(t_pop)

if __name__ == '__main__':
main()

The following figure shows the histogram of a population of exponentially distributed 10, 000 pseudo-random numbers. The distribution is centred on 50, and is positively skewed.



The next three figures show the distributions of means of samples of size 30, 50 and 100. Even though the samples were drawn from a non-normal distribution, the sample distributions approximate normal distribution as the sample size increases.







The importance of the CLT lies in the fact that given normally distributed populations or sufficiently large sample sizes ($n\geq 30$), it shows that (a) the sample statistic ($\mu_\overline{x}$) approximates population parameter ($\mu$) and (b) sampling distributions approximate normal distribution. Once a distribution approximates normality, the properties of normal distribution can be used to make inferences about the sampled population.