Wednesday, September 28, 2016

DESCRIPTIVE VS. INFERENTIAL STATISTICS

Introduction to inferential statistics


In the 1991 hit movie "What About Bob?", psychiatric patient Bob Wiley tells his therapist something like, "There are only two kinds of people in the world, those who love Neil Diamond and those who don't."

What About Bob film.jpg
Some might call Bob's statement a "false dichotomy", but there really are only two kinds of statistics: Those that DESCRIBE and those the INFER. 








 
So far, we have been dealing with those that describe. Now, we are going to start talking about statistics that infer! (It is a great day!)

Descriptive statistics are very nice to us...they don't care about anything other than describing the data that we have in front of us. 

The real fun comes when we use statistics to scientifically predict things. When I was a child, I always wanted to become a weather person...I dreamed about how everyone would love me because of my ability to predict the weather. Almost like magic, I would predict impending weather disasters and save entire cities. Well, I became a statistician, and you are becoming one too! So, we can do similar things (and, sadly, sometimes get it just as wrong as the weather people :(  More on this later...)

Samples: Because collecting data from everyone is too costly!


The entire group of people* you want to study is called the "population". If you can get information from everyone in your population, you are all set! There is nothing else to do besides the descriptive statistics we have been learning: mean, median, mode (measures of central tendency), standard deviation, variance, skew, kurtosis, and maybe some nice visuals: bar chart, pie chart, histogram, frequency table and so on.

Those things tell you all about--or describe--your population.

So, where do inferential statistics come in? We need inferential statistics when we are no longer able to gather data from the entire population!


Gathering data from every person in the entire population is called a "census". Hence the name of the U.S. Census that happens every 10 years--they are trying to gather data from everyone!

Q. So, why not gather data from the whole population?
A. It is too costly!

If you can afford to collect data from every member of your population (the group of people that you want to study), then do it! Usually, that is not the case.

{Quiz}

If you need to know how many people in your office prefer cheese pizza, pepperoni, or black olive, then what is your population?

{Quiz}

Thanks for taking that quiz! Hopefully by now you realize that your population is the group of people you want to study!

After you ask everyone in your office, you might have a frequency table like this:

Fx %
Cheese 5 35.7%
Pepperoni 7 50.0%
Black Olive 2 14.3%
TOTAL 14 100.0%

Now, there is nothing else to do! You know you will need 35.7% of your pizza to be cheese, 50% to be pepperoni and so on (hopefully the pizza place is good at math!).

But what if your population is on a bigger scope? Sociologists, doctors and others are often interested in national trends (or even global).

So, what if your question is: "What are the pizza preferences OF AMERICANS?"

Imagine that. Really imagine that. If your next assignment said, "Find out the pizza preference of each and every American." How would you do it? A survey would cost you THOUSANDS of dollars even if you had already identified every American. Even then, perhaps only 10% will respond without an incentive. Maybe if you gave everyone a $25 gift card for participating, you could get 50% of people to respond. That would cost you OVER $4 BILLION! Let's make it a $5 gift card and assume we still get a 50% response rate (which we probably won't). That is only (*sarcasm alert*) going to cost you $750,000,000. All of this isn't even taking into account the time it will take to do the surveys and analyze the data or what it will cost to hire staff to do all of those things.

Bottom line: We need a better way!

It turns out that we can estimate (key word alert!) ESTIMATE these preferences by taking a sample, a smaller number of people from our total population.

YOU WILL NOW BE LEARNING 2 RULES OF SAMPLING IN DETAIL:

1) If you use a random sample, you can use it to estimate the "true" statistics of your population (The "true" statistics of your population are called parameters. So we have Sample Statistics and Population Parameters! SS, PP).

2) The larger your sample is, the more confident you are that it is a good reflection of the population parameters.

Now, allow me to translate into English from Statistics language using the pizza example:

When you cannot get information about pizza preferences for every single person in the group you are interested in studying, you can use a smaller number of people to estimate pizza preferences for the large group! Using this smaller group means that we may or may not get numbers that match the large group. But we can improve our chances of it by doing two things. First, choose your smaller group without any "rhyme or reason". More specifically, you need to use a method of selection that gives every person in the population a chance to be chosen, and that gives everyone an equal change of being chosen. Rolling dice is a good example (as long as it is not a weighted di from Las Vegas!). The second thing is the get the biggest number of people that you can! The more people you have, the more likely it will be that your numbers match what is really going on in the population!

Whew! See why Statistics is a helpful language? Instead of the whole paragraph above, once you know Statistics language we can just say: You may use a random sample to estimate population parameters and those estimates will be less biased the larger the sample is.

Don't worry! You will get the hang of all of this as we go...The key is to remember that Samples have Statistics and Populations have Parameters.

*Here, we will use the term "people" but it is not limited to that. If you are studying Redwood Trees in a certain forest, your population is the group of trees. It might also be fish, amoebas or anything else that you want to study!

Thursday, January 7, 2016

FREQUENCY TABLES: PART II

Hopefully FREQUENCY TABLES: PART I is permanently emblazoned in your mind and heart. If not, here is the recap of the major points:

POINT #1: Frequency tables are all about summarizing COUNTS or the frequency with which something occurs, BUT NOT ALL NUMBERS IN A FREQUENCY TABLE REFER TO COUNTS! **Be sure you take the time to differentiate between numbers that represent COUNTS or FREQUENCIES and other numbers.

POINT #2: The column on the left is a list of VALUES that someone in the dataset provided. (They are NOT counts even if they are numbers).

POINT #3: FOCUS FIRST ON THE COUNT! Whatever you are doing with the frequency table, make sure you first recognize which column refers to the counts, and which columns do not. This is especially essential if the VALUES are also numerical. 

POINT#4: Interval/ratio variables are terrible candidates for frequency tables! This is especially true when they are continuous variables. (If you need a refresher on levels of measurement, click here). However, people can and commonly do make frequency tables by changing your variable (for example, to ordinal or nominal variables). 

POINT #5: Counts can tell you where the mode is, but counts are NEVER the mode. Students, repeat: COUNTS ARE NEVER THE MODE. They just tell you which category is the most frequent response, but the category itself is the mode. 

The mean and median can also be uncovered through frequency tables, but we have to expand them a little first. 

STEP 1 (VITAL!!): Determine your level of measurement. You may remember that NOMINAL VARIABLES DO NOT HAVE A MEAN OR A MEDIAN. 





add, subtract, multiply, divide…
(MEAN)
Greater than/less than
(MEDIAN)
Difference
(MODE)
Ratio
X
X
X
Interval
X (w/ caution)
X
X
Ordinal
-
X
X
Nominal
-
-
X




Because the mean requires addition and division, it only applies to ratio and interval variables. The median is the middle number when the values are arranged from lest to greatest, so it is only possible to compute it for ratio, interval and ordinal variable because nominal variables cannot be arranged from least to greatest. Mode applies to all levels of measurement because it is simply the most common response.

STEP 2: If you have a nominal variable, make the frequency table as shown in part I, find the most common category (the mode) and you are done!

If not, expand your table.

STEP 3: The first step in expanding the table is to find the cumulative frequency.

CUMULATIVE FREQUENCY=the total count up to a given point. Here is an example that builds on the GPA frequency table from PART I.

Picture #1


So, Column B hold the counts for each value, and Column C (our new column) holds the counts up to a certain point. Here it is crystal clear:

VERY LOW: There are 0 people in this category and it is the first category so it is everything up to that point.
VERY LOW through LOW: There are 3 people from "Very low" through "Low". From the beginning through the "Low" category, there are only 3 people (0 in "very low" plus 3 in "low").
VERY LOW through HIGH: There are 5 people from "very low" through "High". From the beginning through the "High" category, there are 5 people (0 in "very low" plus 3 in "low" plus 2 in "high").
VERY LOW through VERY HIGH: There are 10 people from "very low" through "very high". From the beginning through the "very high" category, there are 10 people (0 in "very low" plus 3 in "low" plus 2 in "high" plus 5 in "very high").

STEP 4: Add a new column, "cumulative percent". It is just like it sounds--the percent of people that the cumulative frequency represents.

Picture #2

The calculations appear in Column D, Cumulative Percent, but usually only the percentage appears. The calculation is a courtesy to make it more clear. Because there are 10 total observations (0+3+5+2=10, or just look at the biggest/last number in the Cumulative Frequency column) we divide each number in the Cumulative Frequency column (Column C) by the total (in this case 10).

This means that from the beginning through "very low" we have 0% of the total observations. From the beginning through "low" we have 30% of the total observations and so on. This is how we can find the median.

The median occurs at the 50% mark. (We call these "percent marks" percentiles). What is the median in Picture #2?

REMEMBER, THE MODE IS NEVER THE COUNT! THE COUNT TELLS YOU WHERE THE MODE IS, BUT IT IS NEVER THE MODE!

You don't have to pass it like this though...
Also, the 50% mark (50th percentile) after all the "HIGH" responses are accumulated, so we need to pass it! We do not PASS the 50% mark until the VERY HIGH category. BUT, the same rules apply here as in the past and because we have an even dataset, the median can technically be between two different values. Arrange them in order like you are familiar with and you will see:

LOW, LOW, LOW, HIGH, HIGH, VERY HIGH, VERY HIGH, VERY HIGH, VERY HIGH, VERY HIGH.

So our median here is between HIGH and VERY HIGH. And that's exactly how you report it with ordinal data: Between HIGH and VERY HIGH.

STORE IN LONG-TERM MEMORY: The same even numbered dataset rule applies to the median as before. If you see the exact 50th percentile in your frequency table, the median will be between the corresponding value and the next value. If it is an odd numbered set, you will not see the exact 50th percentile, and the median is in the category corresponding to the point where you have passed the 50th percentile.

Picture #3
Here is an odd numbered, similar version of the last dataset. What would the median be here?

Now the median is firmly within the "HIGH" category because that is the point where we have PASSED the 50th percentile.

STEP 5: Calculating the mean. Remember this figure of a frequency table of a ratio-level continuous variable (GPA)?

Picture #4


These are the kinds of frequency tables that are not very helpful. They are also the only kind for which you could calculate the mean. It is good practice to learn how to do it, and it may appear ON THE TEST in your stats class.

Let's go back to the "Slices of Pizza Eaten" frequency table from the last unit.

Picture #5

Here is the expanded table that also shows us how to compute the mean:

Picture #6

Again, focus on the COUNT! Look at the 1 slice of pizza row. The count is 4. So 4 people said that they ate one slice of pizza. Similarly, 1 person said they ate 2, 1 said they ate 3, 2 said they ate 4 and 2 said they ate 5. So we really have this:

1 1 1 1 2 3 4 4 5 5

You can see how you could simply add this to get: 1+1+1+1+2+3+4+4+5+5=27.

However, you could also make it easier by doing 4(1)+2+3+2(4)+2(5)=4+2+3+8+10=27.

Basically, the frequency table does the second. We simply multiple the value of each response by the count for that response. Then, we add up all of those multiplied values and divide by the total. Remember, the biggest value in the cumulative frequency column is the total number of observations. (Here the red arrow is pointing it out).


MAJOR TAKEAWAYS:

  • Frequency tables can be expanded to include new columns: cumulative frequency (the total frequency up to a given point), and cumulative percent. (You could also add a category percentages column just be dividing the frequency of each response/category by the total number of observations--not discussed here).
  • The level of measurement of the variable being used in the table determines the measures of central tendency that you can compute:
    • Ratio/interval=mean, median, mode
    • Ordinal=median, mode
    • Nominal=mode only (discussed only in PART I)
    • The median is the CATEGORY (not the COUNT) that corresponds to the point where you pass the 50th percent mark (percentile). If there is an exact 50th percent mark in the table, the media is right between that and the next category. 
  • As in PART I, KEEP YOUR EYE ON THE COUNT!


FREQUENCY TABLES: PART I

A frequency table is a simple thing. Here is one:

PICTURE #1
...but BE CAREFUL it can be deceptively simple! It has trapped and deceived many statisticians over the years. You will be its next victim--or would have been except for this mini lesson. In just a few minutes, you will be Master of the Frequency Table.

A few keys to avoid deception:

POINT #1: Frequency tables are all about summarizing COUNTS or the frequency with which something occurs, BUT NOT ALL NUMBERS IN A FREQUENCY TABLE REFER TO COUNTS! **Be sure you take the time to differentiate between numbers that represent COUNTS or FREQUENCIES and other numbers.

Look at PICTURE #1 again. Which column of numbers are counts and which are not?


COUNT RECOGNITION


  1. WHICH COLUMN IN PICTURE #1 ARE COUNTS?

  2. The left
    The right


ONLY the column on the RIGHT represents counts. So what are the numbers in the column on the left?


LEFT COLUMN CHECK


  1. What do the numbers on the left in Picture #1 represent?

  2. The number of times each person ate pizza
    A total number of slices that was eaten by at least one person in the dataset
    The number of people that ate pizza
    A value of all theoretically possible numbers of slices that might be eaten by people
    None of these



See how it can get tricky?

The column on the left is a number of pieces that at least one person ate. The column on the right is the NUMBER OF PEOPLE that ate that many pieces. MORE ON THIS LATER. For now just let it incubate as you move on to the next point.

POINT #2: The column on the left is a list of VALUES that someone in the dataset provided. It is clearer if we use values that are not numbers. Watch:

Picture #2


Now there is little (to no) confusion about which column contains the counts (frequencies) and which contains the values of the responses.


  • 4 people said "bus" is their preferred mode of transportation
  • 1 said "car (driving myself)"
  • 1 said "car (getting a ride)"
  • 2 said "skateboard"
  • 3 said "bike"
  • 5 said "walk" 
  • and 1 said "other"
Easy to see. Now use the same eyes to look again at Picture #1:

Picture #1
How many people answered with each of the possible responses?

FIND THE COUNTS:

  1. How many people answered with each of the possible responses?

  2. 1 person ate 4 slices, 2 ate 1, 3 ate 1, 4 ate 2 and 5 ate 2
    4 people ate 1 slice, 1 ate 2, 1 ate 3, 2 ate 4 and 2 ate 5
    1 slice ate 4 people, 2 slices ate 1 person, 3 slices ate 1 person, 4 ate 2 and 5 ate 2
    None of these is correct
    All are correct, in a sick way...

Not this count...


Starting to see it now? If not, here is tip #3 to master frequency tables:

POINT #3: FOCUS FIRST ON THE COUNT!


You simply will NOT fail if you focus first on the counts. In the first row of Picture #1, where is the COUNT? It is the column on the right. This is all a little facetious because it is labelled "Count". Don't worry, this is your chance to master the skill of "focusing on the COUNT". Surprisingly, this column often will be labeled either as "counts" or "frequency" or just and "f". Your job, first and foremost is to find that column. It is the absolute ground zero of the frequency table. That is why it is called a frequency table.

NOW YOU TRY>

Picture #3


Make a frequency table for the GPA data by filling in the worksheet below. Your grade will appear on the right.



If you figured that one out--A+ to you! If not, you are probably in really good company at this point. Why? Because frequency tables can be terribly tricky and deceptive as simple as they appear to be.

REMEMBER, FOCUS ON THE COUNT!
Not this count...


And the COUNT lives in Column B! Column A is just telling us all the values that were actual responses. And in this case, *THERE IS NO DUPLICATE VALUE* You notice this if you are focused on the count because you would have noticed that there is only 1 total COUNT of any given response!

In other words, everyone has a different GPA in this case, so there is one response for each value and each value gets its own row:

Here is the answer key. If you didn't get it, no sweat! This was a tough one--as long as you learned to focus on the COUNT, you came out with what you needed to learn!

Picture #4
This brings up the next point to avoid being tricked by frequency tables:

POINT#4: Interval/ratio variables are terrible candidates for frequency tables! This is especially true when they are continuous variables. (If you need a refresher on levels of measurement, click here). However, people can and commonly do make frequency tables by changing your variable. For example, what if we made this more of an ordinal variable by having just 4 categories instead of the exact GPA?

  • 0-0.9999 (Let's call it "VERY LOW GPA")
  • 1.0-1.9999 (Let's call it "LOW GPA")
  • 2.0-2.9999 (Let's call it "HIGH GPA")
  • 3.0-3.9999 (Let's call it "VERY HIGH GPA")
(Notice that the ".9999" endings make it so there is no overlap between categories. Otherwise even numbers would be included in both. E.g. 2.0 would be included in both "1.0-2.0" and "2.0-3.0" categories.)

Now let's look at the new table next to the old table (the colors illustrate the way we combined the categories):

Picture #5

Notice how we start to notice some trends now. For example, a lot of people have "very high" GPAs and no one has a "very low" GPA (grade inflation at work!). you may also notice that it tells a story about measures of central tendency (MEAN, MEDIAN and MODE). Let's quickly revisit those terms in a way that is so simple it almost doesn't do justice to them:

MEAN: All the numbers added up, divided by the total number of observations.
MEDIAN: The middle number when they are all sorted from smallest to largest (or the average of the two middle numbers if it there is an even number of observations).
MODE: The most common number.

You can find all three from this table! The MODE is the CATEGORY (>AHEM< THE "***!!!CATEGORY!!!***" )with the biggest COUNT. Keep your eye on the COUNT!

In this case, what is the mode in our new table with the new categories?

    2
    3
    5
    None of these is correct



KEEP YOUR EYE ON THE COUNT! If you answered 2 or 3 or 5, you were WRONG!! Why? Because THOSE ARE ALL COUNTS! Is the median GPA at a school the COUNT of some category? If you ask me the median GPA at Frantuckanilly State University and I told you 2,342 (the COUNT of people with an average GPA) would it make an sense to you?

POINT #5: A count (frequency) is NEVER the mean, median or mode!!

Read that sentence again 1,822 times. A count (frequency) is NEVER the mean, median or mode!

The count tells us where the mode is, but it is NOT the mode. Think of someone telling you the average GPA at their university is 9,286 and you will get the point.

So, in this case, what is the mode? Hopefully, if you got the mini quiz wrong you answered 5, because 5 is the count that indicates the mode--very high.

The mode is "very high" because it was the most frequent response.

you can also find out the mean and median, but to do this, we need to expand our frequency tables. You will learn how to do that in FREQUENCY TABLES: PART II. 



Tuesday, January 5, 2016

WHO IS REY?

If you saw Star Wars Episode 7 you may be wondering, Who is Rey? Rey herself said, "I am no one." Bu there is little doubt we will find out more about her past in the coming two movies. However, with release dates deep into the future, you may feel too anxious to wait! The Google search "Who is Rey?" Generates over 120 MILLION hits. The internet is ablaze with conversations and debates about what will be revealed in the next two movies.

If you are lucky enough to have some skills in statistics, you may be able to get ahead of the game. Remember that show "Who Wants to be a Millionaire?"? It was amazing that the poll the audience answers seemed to yield the correct response so often. Perhaps you could use statistics to "poll the audience" and it will give us the correct answer.

The problem is that with 129,000,000 matches and multiple possible theories being expressed on each matching web page, the task is formidable. Random sampling makes it possible to take a smaller number of those webpages and still end up with the same answer--at least it gives us a range that we are somewhat confident about (see other posts on sampling and confidence intervals on this cite).

First, it can be helpful to do some pre-research so that we have some idea what we are looking for. I have done it for us this time around and found nine theories (some related to others) and twelve criteria that people say should be satisfied by the theory.

The theories:


First, have a look at the nine theories:


  • The Obi Wan's granddaughter theory (daughter of Luke and Obi Wan's daughter)
  • The daughter of Luke Skywalker theory
  • The daughter of Luke and a "Mara Jade" type Jedi theory
  • The daughter of Luke and a "Mara Jade"-turned-evil theory
  • The daughter of Leia and Han Solo theory
  • The daughter of Leia-turned-evil and Han Solo theory
  • The daughter of Obi Wan Kanobi theory
  • The conceived by the force theory
  • The reincarnated Anakin Skywalker ("Chosen one") theory

These were gathered by a purposive selection of articles that seemed most relevant. Purposive sampling means that articles are chosen based on the information they can provide. The results of purposive samples are not generalizable (able to be applied to the whole population) but can be an excellent choice in exploratory research or in pre-research because it helps the researcher get their bearing with major themes relevant to the study. In this case, it helped us uncover nine theories that we kept hearing over and over again.

There could be millions of potential theories, but we can know when to stop doing our pre-research when we start hearing only the same major theories over and over again. (We sometimes call this saturation).

The criteria:

Next, there were also twelve criteria that people kept talking about. According to the articles and comments read, these are things that the theory should satisfy for the theory to be chosen by those making the movie. A good theory should:

  • Have the ability to explain Rey's advanced abilities with the force
  • To explain her advanced pilot skills
  • To explain her advanced mechanical skills
  • To explain why she was abandoned (or placed) on Jakku as a young child
  • To explain Maz Kanata's statement that Rey's family is not coming back
  • To explain the draw Rey has to Luke's/Anakin's lightsaber
  • Keep the movies Skywalker family-centric (to meet a statement made by creators of the movies)
  • Create an interesting plot twist
  • Be true to the Star Wars feel (and perhaps parts of even the expanded universe EU)
  • Explain Obi Wan's "first steps" line in Rey's vision
  • Explain cinematographic allusions or foreshadowing about who Rey is (like the look of her clothing or her upbringing on a dust planet).
  • Explain the apparent connection between Rey and Leia toward the end of the movie
There are two ways to look at the "worthiness" of a theory vis-a-vis the criteria: How well a theory meets each criterion, and how many of the criteria it meets. One theory may provide a fascinating and excellent explanation about Rey's ability as a pilot but completely fail to explain Rey hearing Obi Wan's voice during her vision. Similarly, one theory may provide a fair explanation of all these criteria, but another may provide phenomenal explanations for half of the criteria. 

This will have to be sorted out later. But, just keep in mind that our evaluation of these theories will be some combination of how good the theory is at explaining each criterion, and how many of the criteria the theory explains. These could be called "depth" and "breadth" respectively. 

Preliminary results

There are many possible ways to go about this but one is to make a crosstabulation (crosstab). This simply means that we put the categories of one thing as column headers, and the categories of the other as row headers, and they we will in the frequencies we observe at the intersection. (In practice we usually make two variables and the "intersect" or "cross" them using statistics software). 

For now, I have filled in each cell with my subjective analysis based on my readings. 

The table is below:



This is simply the product of me rating each theory (in rows ->) as + , ++ , or +++ where + means "the theory would provide a decent explanation" and +++ means "the theory would provide a very good explanation". Notice there is also a - rating, meaning "this would provide a bad explanation of the criterion".

Now, back to the two ways of assessing the results: depth and breadth. If you scroll over to the right, you can see the total number of points each theory got. This is its overall strength and is simply the total number of + that it has, minus the total number of -. So + gets 1 point, ++ gets 2 points, +++ gets 3 points and - gets -1 point.

To the right of that is another column that give each theory a point for each criterion that it satisfies with at least one + .

Results

Rank by total points (depth):

  1. Reincarnated Anakin/Chosen one
  2. Daughter of Luke and "Mara Jade"-turned-evil
  3. Daughter of Han Solo and Leia-turned-evil
  4. Daughter of Luke and "Mara Jade"
  5. Luke's daughter
  6. Force conceived
  7. Daughter of Luke and the daughter of Obi Wan's
  8. Daughter of Han Solo and Leia
  9. Daughter of Obi Wan
As we see here, the more simple versions of theories are less able to provide strong explanations for the different criteria on average. The top theories involve more details and usually, a woman that has turned evil. The internet world tends to find some appeal in the idea that there will be a "Rey, I am your mother" moment over the next two movies at some point. The idea is that either "Mara Jade"/Luke's Jedi wife turned evil in a sort of Darth Sidius reversal. She feels that the only way to conquer the dark side is to make one's way into it and then destroy it from within. Luke, disagreeing with this philosophy parts ways with his wife and has to hide their young daughter (Rey) and wipe her memory. Nevertheless, she has some Jedi training from Luke (and possibly ghost Obi Wan) that resurfaces later on, making Rey as powerful as we see in the movie. So, Luke's estranged wife will turn out to be either Snoke (a disguise) or Phasma, and, upon learning of Rey, try to bring her to the dark side. Some in this camp even think it may be Ben's intention to do the same, thus the scene where he looks at the Darth Vadar mask and says, "I will finish what you started"--not referring to destroying all Jedi, but to restoring balance to the force by passing through the dark side. 

A similar vibe runs through the Leia-turned-evil theory--that she is not pleased with the inability of the republic to put down the first order and decides to take a small band of resistance fighters to do it. Thus, the resistance is portrayed as a small movement with rudimentary spacecraft rather than more elaborate ships and equipment. This could play out with some similar "Rey, I am your mother" moments, and many believe that Leia might also be behind Snoke (as a disguise). Other possibilities behind Leia's turn to the dark side might be related to her inability to face the darkside like Luke did with Darth Vadar, her lack of training in the force by the light side, or, possibly the same thing that would turn a "Mara Jade" character to the dark side--fighting it from within. 

The #1 theory, however, does not have this tone. Instead, it portrays Rey as a reincarnation of Anakin, or the Chosen one. In this theory, the "Chosen one" is not a single person, but takes on many different personas over time through reincarnation. Thus, Rey is, in a sense, Anakin, out to undo the mistakes of his past life. We see Rey on a dusty run-down planet, good at flying and fixing things, and some argue that Rey bears a remarkable resemblance to Shmi Skywalker. This theory obviously explains a lot of criteria because it is sort of the catch-all--instead of answering how she is related to Anakin (a point that seems to be pretty overtly made by the movie makers), it simply asserts that she is him. However, this theory provides great depth of explanations for a lot of the criteria, but not as much breadth as others. For example, it does not offer a ready explanation of why she was abandoned as a young child or hears Obi Wan Kenobi's voice in her vision. 

Let us look at the ranking by breadth--percent of criteria satisfied:

  1. Daughter of now-evil "Mara Jade" and Luke
  2. (tie for 1st) Daughter of Han and now-evil Leia
  3. Reincarnated chosen one 
  4. Daughter of "Mara Jade" and Luke
  5. (tie for 4th) Luke's daughter
  6. Daughter of Obi Wan's daughter and Luke
  7. (tie for 6th) Obi Wan's daughter
  8. Daughter of Leia and Han
  9. Force conceived
Once again, there is the general trend that more complex theories involving women turned evil are at the top! In fact, the same theories occupy the top three places, but the "reincarnated chosen one"theory has fallen to 3rd. This is because, while it satisfies many of the criteria very well, it does not satisfy as many of the criteria as some other theories.

Conclusion

Short of drawing up and conducting a full survey to a representative sample of the Star Wars fan universe, we conclude that three of the theories seem to land at the top:

  1. Daughter of now-evil "Mara Jade" and Luke
  2. Daughter of Han and now-evil Leia
  3. Reincarnated chosen one 
So, which should be crowned the best? Here, we might want to weight depth or or breadth differently--implying that one is more important than the other. But, assuming we weight them equally, we just average them and end up with these final standings:

  1. Daughter of now-evil "Mara Jade" and Luke
  2. Daughter of Han and now-evil Leia
  3. (tie for 2nd) Reincarnated chosen one 
So, in episode 8 when most of the world finds out that Rey is the daughter of Luke and a "Mara Jade" type Jedi-turned Phasma who was trained by Luke and ghost Obi Wan but had her memory wiped and was sent to Jakku to protect her from the dark side and her mother, maybe you will be able to say you heard it here first thanks to the power of analysis.

Saturday, June 6, 2015

Statistics Superheroes

Super heroes that help you remember all those Greek characters! The first two have been rolled out, hopefully with more to come!

Is there a Greek or other statistics symbol you just can't seem to remember? Comment below!


SIGMA MAN:


statistics superhero

FASTER AND STRONGER THAN 99.99% OF THE POPULATION!


BETA MAN:


statistics superheroes

HEROIC, BUT CAN'T BRING HIMSELF TO REJECT THINGS THAT AREN'T TRUE!

Thursday, May 7, 2015

MEASURES OF VARIANCE

HOW FAR SPREAD OUT SOMETHING IS 

This unit explains different measures of variance. Measures of variance refer to how spread out a dataset is. Measures of variance include:

  • The range
  • Interquartile range
  • Standard deviation
  • Variance (this one seems obvious!)
After you complete this page and the quizzes on it, you should have a pretty solid foundation for understanding measures of variance. 

If you think you may already be a pro at this, just skip the explanations and go right to the quizzes. If you pass all the quizzes, you may be ready to move on! 

Range = The biggest number - the smallest ("Max-Min")

Pretty easy...Why don't you have a go at it:

What is the range of the following dataset?

2,3,3,3,4,5,5,6,7,9,10

    2 to 10
    9
    8
    7
    6



How did you do?







Interquartile Range=The 3rd quartile-the 1st quartile. 

So, what's a quartile? It is literally a quarter (like 25 cents). So the first thing to do is to identify what 1 quarter of the data is. 

Look at this little dataset: 1,1,2,3,3,3,3,4,5,6,8,9. 

There are 12 observations, so a quarter (or 1/4) of 12 is 3. 

This means that 1,1,2 are the first quarter (or 4th or first 25%) of the dataset.

3,3,3 are the next (2nd) quarter.

3,4,5 are the 3rd quarter.

And 6,8,9 are the fourth or last quarter. 

However, in statistics we often talk about quartiles instead of quarters. Quartile means that little infinitesimally small point "between" one quartile and the next. So the first quartile is right between 2 and 3. In this case, we average the two numbers (2 + 3 / 2=2.5). 

The 1st quartile is 2.5

The 2nd quartile is 3

The 3rd quartile is 5.5 (right between 5 and 6)

And...wait for it...There is no "4th quartile"! At least not that we talk about in statistics. Some people challenge this, and I suppose there is a theoretical 4th quartile right after the last number, but in this case, we don't know what the number after that is, so we can't average it anyway...

Despite a few people that want to talk about a "4th" quartile you will never really see it pop up--so no worries!

Now, we have the 1st, 2nd and 3rd quartiles. 


SIDE NOTE: It turns out that the 2nd quartile (sometimes called the "middle" quartile) is also the median. (Remember the median is the number right in the middle? So is the 2nd quartile!)
Now that you know the quartiles, the interquartile range is very straightforward: Find the 3rd quartile and the 1st quartile, then subtract the 1st from the 3rd. 

Interquartile range = the 3rd quartile - the 1st quartile. 

NOTE OF CAUTION: In the example above we had 12 observations and 4 divided nice and evenly into it. But that is not always the case. Consider this mini dataset:

4,5,6,6,6,7,7,8,9,9,10

Here we have 11 observations. So 11/4=2.75. So it is a little harder to brake it up into quartiles (4ths). To do this, first find the median:


4,5,6,6,6,7,7,8,9,9,10

Median=7. 

Now divide the dataset into two smaller datasets, including the median in each:

4,5,6,6,6,7
              7,7,8,9,9,10

Now, find the median of each half (remember to include the median in each):

4,5,6,6,6,7 (6+6)/2=6

So Q1=6.

7,7,8,9,9,10 (8+9)/2=8.5

So Q3=8.5

Now we have 
Q1=6
Q2=7 (the median)
Q3=8.5

Can you compute the interquartile range? Remember it is just Q3-Q1 :)

IQR: 8.5-6=2.5
The interquartile range is often shown visually through a graphic known as the boxplot or box and whisker plot. 

bow and whisker plot explained

Notably, the "min" and "max" exclude outliers so they may not match up with the very smallest and very largest numbers. It is different with each software package you use, but it is often 3 times the IQR above the mean (for the "max" line) and 3 times the IQR below the mean (for the "min" line). 

Here the boxplot is turned sideways, but it can be shown vertically as well:





Monday, April 20, 2015

Z SCORES FOR OBSERVATIONS IN A SAMPLE

OVERVIEW:

In this unit, we will discuss how to calculate and use Z scores for observations in a sample or population. By the end of this lesson you will better understand how to use Z scores. These are the main points we will cover:
  • Z scores tell us how far from the mean an observation is in standard deviations. 
  • If you know a Z score for an observation, you can look up the corresponding percentile based on the area under the curve.
  • You can turn raw scores to Z scores by finding the raw distance from the mean, and then dividing by standard deviation. (raw-mean)/standard deviation
  • You can turn Z scores into raw scores by multiplying Z times the standard deviation and adding that to the mean. (Z * standard deviation)+mean.


Odds are you have taken a standardized test at some point, like the SAT or ACT. At some point, you see your score in a percentile instead of the actual score. Percentiles are useful because they let us compare from one kind of thing to another.

Some people take the ACT in high school but others take the SAT. While certain colleges may only take one or the other, many will take either. Each test has a different scoring system so it would be like comparing apples to oranges when deciding which students to admit. The percentile ranking gives them some basis for comparison. For example, one student may have scored in the 95th percentile on the ACT and another in the 90th percentile on the SAT, so it seems like the first student performed better, even though we could never know that based on the raw scores alone because of the way the tests use different scoring systems.

A Z SCORE DOES THE SAME THING!

Z scores "standardize" how big or small a raw score is compared to all the other scores. Your percentile means that you score higher than that percent of people.


UNIT 1

A Z Score is: THE NUMBER OF STANDARD DEVIATIONS AWAY FROM THE MEAN.

Just like that sounds, a Z score of 1 means that the observation is 1 standard deviation from the mean, a Z score of 2 means that the observation is 2 standard deviations from the mean and a z score of 1.28695 means that the observation is 1.28695 standard deviations from the mean. 

Now you try...

  1. Bill is 182 cm tall which is 1.2 standard deviations from the mean. What is the Z score?

  2. 182
    218.2
    1.2
    1.82
    None of these

  3. A leaf is 22 cm across--a Z score of 2.9. How many standard deviations is this from the mean?

  4. 22
    63.8
    2.2
    2.9
    None of these

  5. Fahid scored 11 on a test. The average was 10 and standard dev. was 1. What is the Z score?

  6. 11/1=11
    11/10=1.1
    11-10=1/1=1
    Not enough information
    Too much information

  7. Sara's shoe size of 11 has a Z score of 2. The average shoe size is 10. Compute stand. dev.:

  8. 11-10=1/2=0.5
    11-2=9/10=0.9
    11/10=1.1
    10+2=12/11=1.09
    Not enough information

  9. Will you ever get the hang of this?

  10. Yes!
    Yes!
    Yes!
    Yes!!
    YES!!!!!!!!!!


How did the quiz go? The idea is that Z SCORE = THE NUMBER OF STANDARD DEVIATIONS FROM THE MEAN!


UNIT 2


So if you know that a score was 32 points from the mean, and the standard deviation is 32, this score is 1 standard deviation from the mean. Because Z score means "the number of standard deviations from the mean", this score of 32 is not only 1 standard deviation from the mean, but we can also say that it has a Z score of 1!

Likewise, if we know that the average IQ is 100, and Michael's IQ of 115 has a Z score of 1, we now know the standard deviation--BECAUSE Z SCORE MEANS "THE NUMBER OF STANDARD DEVIATIONS FROM THE MEAN!" Because Michael's IQ of 115 has a Z score of 1, we know that a distance of 15 points from the mean (115-100) is a standard deviation of 1! 15 points from the mean is not only a Z score of 1, but 1 standard deviation from mean--BECAUSE Z SCORE IS THE NUMBER OF STANDARD DEVIATIONS FROM THE MEAN. 

Z score and number of standard deviations from the mean are the same thing!

NOTE OF CAUTION: Z score and number of standard deviations are NOT the same thing! Let's hear that again: Z score and number of standard deviations are NOT the same thing! It has to be number of standard deviations from the mean! 

FOR EXAMPLE: We know from the example above that standard deviation is 15. But a Z score of 2 is not 30 (2*15), but 130 (the mean of 100 + 2*15). 


So, why is this useful?

As mentioned above, Z scores can be used in a similar way to percentiles--to compare observations with different raw scores. 

This is because each Z score goes with a certain percentile. There  is a catch though. When we compute a Z score, we are assuming normal distribution of our data, which means that most cases are in the middle with fewer and fewer as we taper off to either side. This is shown by the red dotted line in the picture below:

z-score normal distribution



See how there are the most cases in the middle and there are fewer and fewer as you taper off to either side?

This is covered in more depth in the lesson on distribution. (Review it now if needed). 

The green line shows the mean. The other colored lines show the different standard deviations. Now for the first time, we can really see what is "standard" about the standard deviation--the raw distance from one standard deviation to the next is the same! 

Remember that in the IQ example we said that standard deviation was equal to 15 IQ points and the mean was 100. So in that case the green line is at IQ=100. 1 standard deviation (the blue line) is at IQ=115 (100+15); 2 standard deviations (the golden line) is at 130 (100+15*2); and 3 standard deviations (the brown line) is at 145 (100+15*3). Forget about the negative numbers for now...

In short, in the IQ example, the green line (the mean) is at 100, and there is a space of 15 IQ points from one line to the next. 

So in this dataset, 15 is our standard, and "deviation" means how far away we are from something (in the case of Z scores--how far from the mean). 

So standard deviation literally means a unit of fixed size away from the mean. But the size of that unit varies from one dataset to another. You have to figure out the raw distance of your standard deviation for each new dataset, but from there, it will be your "standard" way of referring to Z scores and percentiles. 

The reason is because of calculus. Even though the thickness of our red dotted line curve (the distribution) changes at each point along the distribution, we can use calculus to know what percent of the total area falls under the curve at any given point. This lesson assumes that you are not required to understand the calculus behind that, and in fact, you don't necessarily "need to". This is because those areas are always the same for any given Z score in a normal distribution. 

Let's keep it simple for now though...

Let's back up and review the major points so far before we go on:
  1. Z score means "the number of standard deviations from the mean"
  2. The standard deviation in one dataset is different than that of another
  3. The thing that is "standard" about standard deviation is that it is your "ruler" for any given dataset. That set distance from the mean can tell you the percentile of any observation in the dataset
We have covered point 1 fairly exhaustively at this point, so let's elaborate on point 2 and 3 a little. 

2. For IQs, the mean is 100 and the standard deviation is 15. Those are national averages. But let's go back to the standardized test example we mentioned at the beginning of the lesson. 

The mean on the SAT is usually around 1500, with standard deviation of 250. The mean for the ACT is around 20.8 with standard deviation around 4.8. 

So the SAT and ACT have very different standard deviations (250 and 4.8), but this makes sense once we delve into point #3.

3. The standard deviation is set at the point where 34.1% of the cases are between the mean and that point. So for the SAT, a standard deviation of 250 means that 34.1% of test takers scored between the mean, and the mean + 250 (1 standard deviation). So if you score a 1750 on the SAT, you outscored 34.1% of everyone with an above-average score. This will obviously be different for the ACT. On the ACT, you would accomplish the same feat of outscoring 34.1% of those with above-average scores by getting a 25.6 (mean plus 1 standard deviation=20.8+4.8=25.6). 

So a student with an SAT score of 1750 scored in the same percentile as a student with a 25.6 ACT score. NOW WE CAN COMPARE APPLES TO ORANGES (in a way...)!

NOTICE, THOUGH THAT THE CHANGE FROM 0 to 1 standard deviation in terms of percent  is NOT the same from one standard deviation to the next! This is simply because the curve is getting thinner and thinner, so going out 1 standard unit covers less and less area as you move toward the extremes (edges). 

Let's do some examples before we move on to the 2nd quiz!


Use this picture to visualize the examples below.
area under curve z scores

Let's pretend the picture above shows the distribution of IQ scores in the United States.  
The green line is 100, because that is the mean.

Each colored line is 15 IQ point, because that is our "ruler" or "standard" for this dataset. We know that because 15 is the standard deviation for IQs in the United States. In other words, 34.1% of people have an IQ between 100 and 115. (We will show how to calculate this later...)

Now you try. Use the picture above to answer these questions...

  1. What is the IQ of someone at the point of the brown line to the LEFT of the mean (-3 sd)?

  2. 100-3*15=55
    100-15=85
    3*15=-45
    -1*15=-15
    None of these

  3. What is the IQ of someone at the point of the golden line to the LEFT of the mean (-2 sd)?

  4. 15*2=30
    100+15*2=130
    -2*15=-30
    100-15*2=70
    None of these

  5. What is the IQ of someone at the point of the blue line to the LEFT of the mean (-1 sd)?

  6. -1*15=-15
    100-15*1=80
    100+15=115
    15*1=15
    None of these

  7. What is the IQ of someone at the point of the blue line to the RIGHT of the mean (1 sd)?

  8. 100-15=85
    1*15=15
    100-15*2=70
    100+15=115
    None of these

  9. What is the IQ of someone at the point of the brown line to the RIGHT of the mean (3 sd)?

  10. 100-3*15=55
    100+2*15=130
    100+3*15=145
    2*15=30
    3*15=45


How did you do? If you missed any answers, consider reviewing the major points and taking the quiz again before moving on. 

So, we know that a standard deviation is the raw value from the mean that gives us 34.1% of the area under the curve. (Again, we are not reviewing calculus here so you will have to take my word for it...). 

We also know that whatever that raw score ends up being (for IQ it is 15) is our "standard" for measuring how far things are from the mean. We can also use this to compute percentiles. This is how we do it:


HOW TO COMPUTE PERCENTILES:

Look again at the percentages in the picture. 

percentages for z scores


Let me point out something that is very useful: The TOTAL area is 100%, and either half is 50%. This means that we can use the other percentages to figure out percentiles. For example, to compute the percentile of a score 2 standard deviations from the mean, start with 100% (the total area) and subtract the percentage of people with scores greater than 2 standard deviations (2.1% between 2 and 3 s.d. and another 0.1 beyond 3 s.d.). So we have 100%-2.1%-0.1%=97.8%. So if your score is 2 standard deviations from the mean for an exam, you scored in the 97.8 percentile! Good job! 

You can also use 50% instead of 100% to do this. So, if you scored 2 standard deviations from the mean, you outscored the 50% in the left half of the distribution, but also the people between the mean and 1 s.d., and the people between 1 s.d. and 2 s.d. So we have 50%+34.1%+13.6%=97.7%. This tells us that you scored in the 97.7th percentile! This is the same answer we got by subtracting from 100, except for some small rounding error caused by only giving the areas out to the tenths place. 

Now you try:

Remember not to worry too much about rounding; just choose the closest answer!

  1. What is the percentile of someone whose test score is -2 standard deviations from the mean?

  2. 50-2.1-0.1=47.8 percentile
    100-2.1-0.1=97.8 percentile
    2.1+0.1=2.2 percentile
    Not enough information
    Too much information!

  3. What is the percentile of someone whose test score is 3 standard deviations from the mean?

  4. 100-0.1=99.9 percentile
    50-0.1=49.9 percentile
    0.1 percentile
    Not enough information
    Too much information!

  5. What is the percentile of someone whose test score is -1 standard deviations from the mean?

  6. around the 34.2 percentile
    around the 15.8 percentile
    around the 84.1 percentile
    Not enough information
    TMI!

  7. What is the percentile of someone whose test score is -3 standard deviations from the mean?

  8. around the 0.1 percentile
    around the 49.8 percentile
    around the 99.9 percentile
    Not enough information
    TMI!

  9. What percent of people scored between -1 s.d. and 2 s.d.?

  10. around 97.7%
    around 72.3%
    around 81.1%
    Not enough information
    TMI!


Once again, be sure to review anything you missed before moving on!

UNIT 3

If you have passed all the quizzes to this point, congratulations, we are ready to get a little more advanced!

So you may be thinking, my IQ is 128, not 100, 115, 130 or any other number e have talked about--how do I know my Z score and associated percentile? 

It is actually painlessly easy! (Sorry to disappoint you...). Remember to think of standard deviation as our "ruler" for our dataset. So, in the case of IQs, the "ruler" is 15 IQ points long, and we have discussed the percentiles associated with 1 ruler, 2 rulers, 3 rulers, and even the negative versions of those. BUT WE CAN HAVE PARTIAL RULERS TOO! So if your IQ is 122.5, you know that you are 1.5 "ruler" lengths from the mean (15+7.5=22.5. 100+22.5=122.5). So how would we calculate your Z score (how may standard deviations you are from the mean) if your IQ is 128? 


How to calculate Z score:


  1. First, find your distance from the mean: 128-100=28, so you are 28 points from the mean.
  2. Now find out how many standard deviations there are in your distance from the mean. In this case, we just divide 28 by 15 (the standard deviation). So 28/15=1.867.
  3. YOU FOUND THE Z SCORE!
It is just like converting feet to inches, miles to feet--you just divide! You are just finding out how many standard deviations are in your difference from the mean.

In terms of an equation, you may see something like this in your textbook:

Z=(X-Xbar)/s

TRANSLATION: "Subtract the mean from the raw score, then divide by standard deviation".

CAVEAT: MAKE SURE TO DO RAW SCORE - MEAN, NOT MEAN - RAW SCORE OR YOU WILL GET THE OPPOSITE SIGN (negative vs. positive) AND A VERY WRONG (OPPOSITE) ANSWER!

You can also turn a Z score into a raw score this way:

Raw score=s*Z+Xbar

In our example above, we know a Z of 1.867 goes with a raw score of 128. We can see this using the equation:

128=15*1.867+100.

Because the Z score is really the number of standard deviations from the mean, both equations make sense. If you want a raw score from a Z score, you just multiply Z (the number of standard deviations from the mean) by the standard deviation, which is a raw score. Here, Z of 1.867 means "the number that is 1.867 standard deviations from the mean". So 15 (the standard deviation) *  1.867 gives us 28. In other words, 28 is 1.867 "fifteens" from the mean. We then add that to 100 to get 128. 

Now you have a go...

Suppose that the average number of points per game for professional basketball teams is 96, with standard deviation of 12 points. Based on this information, answer the questions below:

  1. What percent of teams scored between 84 and 108 points per game?

  2. 34.1%
    84.1%
    18.7%
    68%
    None of these

  3. What is the Z score for a team that scores 112 points per game?

  4. 1.3
    4.5
    0.8
    1.5
    None of these

  5. What is the raw score that corresponds to a Z score of 3.4?

  6. 137
    145
    108
    99.4
    84.2

  7. What is the Z score for a team that scores 87 points per game?

  8. 0.75
    0.3598
    -1.28
    1.01
    -0.75

  9. What is the raw score that corresponds to a Z score of-0.54

  10. 68.73
    89.52
    133.04
    102.48
    140.05


How did you do? If you scored 5/5, you probably have a pretty good understanding of Z scores! If you scored below 4/5, you may need to go back for some review!

Finally, if the Z score you are interested in is not 1, 2, 3 or -1, -2, or -3, like in this picture:


z score distribution

you can still find the percentile by using a table like this one from UT Dallas. You just go down the left column until you find the score to the tenths place, and then over to the number in the hundredths place. So if your Z score is -2.67, you go down to -2.6, and then scroll across to the number in the 0.07 column. This table gives the area under the curve (to the left of the Z score) so Z of -2.6 is greater than .0038, or .38% of the observations. You could also say this is the 0.38 percentile. That's not where you want to score on the ACTs!



Summary
Often, it can be useful to compare similar things that have different "scoring" systems. Percentiles are one way to do that. Z scores are another, similar way. 


  • Z scores tell us how far from the mean an observation is in standard deviations. 
  • If you know a Z score for an observation, you can look up the corresponding percentile based on the area under the curve.
  • You can turn raw scores to Z scores by finding the raw distance from the mean, and then dividing by standard deviation. (raw-mean)/standard deviation
  • You can turn Z scores into raw scores by multiplying Z times the standard deviation and adding that to the mean. (Z * standard deviation)+mean.
GOOD WORK! YOU NOW HAVE THE TOOLS TO USE AND UNDERSTAND WHAT Z SCORES MEAN! For more on Z scores, check out the lesson on finding Z scores for a sample mean...