From Skewed to Normal: Visualizing the Central Limit theorem and Confidence Intervals
Summary
This activity is designed to provide an overview of the concept of the Central Limit Theorem and the Confidence Interval. The code walks students through the practical implications of the CLT and how re-sampled data always approaches normality despite the fact that the initial sample suggests otherwise.
Learning Goals
1 By the end of this activity, students will be able to:
Learn how to create a random population of observations
Learn how to sample repeatedly from that random population
Generate a sampling distribution that looks normal compared to the original population.
2 Matlab is used to create the population, and generates visualizations based on student input. Students can vary the size of the population that they start with.
3 The activity does include higher-order skills such as critical thinking because they have to understand how the initial skewed population approaches normal. They learn data analysis because the activity includes the ability to discern whether data is skewed or normally distributed. Finally, students learn about model development because they need to generate a mental model for how data is distributed.
Learn how to create a random population of observations
Learn how to sample repeatedly from that random population
Generate a sampling distribution that looks normal compared to the original population.
2 Matlab is used to create the population, and generates visualizations based on student input. Students can vary the size of the population that they start with.
3 The activity does include higher-order skills such as critical thinking because they have to understand how the initial skewed population approaches normal. They learn data analysis because the activity includes the ability to discern whether data is skewed or normally distributed. Finally, students learn about model development because they need to generate a mental model for how data is distributed.
Context for Use
The activity is designed primarily for upper-division/graduate level students. This training was delivered as part of a stats boot-camp at a regional comprehensive University. The cohort who worked through this initial simulation was small (n=12) but could easily scale up to a lab activity.
There are no technical skills or experience with Matlab required to complete the activity. A working understanding of statistics particularly descriptive students should suffice for this activity. This activity was the 4th in a series of 10 stats activities for an interested group of students and faculty. It would not be difficult to adapt it to other settings; any discipline in which statistics is necessary to understand, analyze, and interpret data would be a fine venue for this activity.
There are no technical skills or experience with Matlab required to complete the activity. A working understanding of statistics particularly descriptive students should suffice for this activity. This activity was the 4th in a series of 10 stats activities for an interested group of students and faculty. It would not be difficult to adapt it to other settings; any discipline in which statistics is necessary to understand, analyze, and interpret data would be a fine venue for this activity.
Description and Teaching Materials
The activity is a follow-up to a video that I recorded outlining the different statistical distributions commonly used in statistics. The mechanics are as follows:
Students watch an introductory video found here: https://youtu.be/mgsXHBMYTv0
Because the video is a follow-along with students working on the items as-is, the students annotate the file "CLT Student Scaffold"
The faculty member can check the work using the pdf file "CLT Instructor Scaffold"
After the video lecture, the lab (students and faculty) walk through the simulation together. The code is setup in cells, so faculty can choose to do parts of the simulations, or run all of them.
The code does the following:
Produces a skewed population
Draws n random samples (default is 75 per sample)
Plots the individual samples to demonstrate that the raw data distribution is skewed
Computes the mean, sd, and 68/95% CIs.
It then sees how many of those intervals include the mean and how many do not. The key insight is that the confidence interval determines the number of times the average is included , for a 95% interval, it should "miss" 5 times (32 times for the 68% CI). The students will see that it's not always 5 nor 32, whic is the point.
Central Limit Theorem Simulation (Matlab File 11kB Sep30 26)
CLT Lecture Student Scaffold (Acrobat (PDF) 1.8MB Sep30 26)
CLT Instructor Scaffold (Acrobat (PDF) 2.3MB Sep30 26)
Students watch an introductory video found here: https://youtu.be/mgsXHBMYTv0
Because the video is a follow-along with students working on the items as-is, the students annotate the file "CLT Student Scaffold"
The faculty member can check the work using the pdf file "CLT Instructor Scaffold"
After the video lecture, the lab (students and faculty) walk through the simulation together. The code is setup in cells, so faculty can choose to do parts of the simulations, or run all of them.
The code does the following:
Produces a skewed population
Draws n random samples (default is 75 per sample)
Plots the individual samples to demonstrate that the raw data distribution is skewed
Computes the mean, sd, and 68/95% CIs.
It then sees how many of those intervals include the mean and how many do not. The key insight is that the confidence interval determines the number of times the average is included , for a 95% interval, it should "miss" 5 times (32 times for the 68% CI). The students will see that it's not always 5 nor 32, whic is the point.
Central Limit Theorem Simulation (Matlab File 11kB Sep30 26)
CLT Lecture Student Scaffold (Acrobat (PDF) 1.8MB Sep30 26)
CLT Instructor Scaffold (Acrobat (PDF) 2.3MB Sep30 26)
Teaching Notes and Tips
* The code is divided into "cells" to facilitate inspection of the code and to run smaller parts to build the theory. It is also commented very extensively to illustrate the concept and the method of the code. There's nothing fancy in the code per se, but you should have the Statistics Toolbox installed just in case.
* Students may not be used to the probability distribution plots. The idea is for them to see that skewed data may still hue close to the normal probability for values around the 50% percentile, the skew puts values above and below the line at the extremes. After generating the sampling distribution, the data hue much more closely to the normal line.
* The last two cells are helpful tools to show students how to do two critical inferential statistic calculations:
1 How to generate the probabiity for a statistic for Z, t, and F
2 What is the critical value for a test statistic, in this case, they get the Z, t, F. This can be useful for a quick check if the statistical value you calculated has a p-value below 0.05
* Students may not be used to the probability distribution plots. The idea is for them to see that skewed data may still hue close to the normal probability for values around the 50% percentile, the skew puts values above and below the line at the extremes. After generating the sampling distribution, the data hue much more closely to the normal line.
* The last two cells are helpful tools to show students how to do two critical inferential statistic calculations:
1 How to generate the probabiity for a statistic for Z, t, and F
2 What is the critical value for a test statistic, in this case, they get the Z, t, F. This can be useful for a quick check if the statistical value you calculated has a p-value below 0.05
Share your modifications and improvements to this activity through the Community Contribution Tool »
Assessment
A good way to assess student mastery of the code would be to have them generate non-standard confidence intervals, such as the 75% confidence interval. To do this, they would
1) Determine the value of Z to multiply by the SEM. A good way to have them conceptualize this is to note that 25% of the distribution is split in half, so rather than looking for 1.00 which corresponds to 16% in one tail (thus 0.8419), they need to look for 0.875, which is 1.15
2) They could then update the code (line 136). This could also be a question. Give them the multiplier and ask them to identify how to modify the code to produce the "waterfall" plot that is generated.
1) Determine the value of Z to multiply by the SEM. A good way to have them conceptualize this is to note that 25% of the distribution is split in half, so rather than looking for 1.00 which corresponds to 16% in one tail (thus 0.8419), they need to look for 0.875, which is 1.15
2) They could then update the code (line 136). This could also be a question. Give them the multiplier and ask them to identify how to modify the code to produce the "waterfall" plot that is generated.
References and Resources
A standard Bio/Statistical text should provide sufficient background, such as Sokahl and Rolf or Zar
Below is a good explainer that doesn't use R or Python. It uses Minitab to generate the output, but doesn't obscure things by incorporating code into the webpage.
https://statisticsbyjim.com/basics/central-limit-theorem/
Below is a good explainer that doesn't use R or Python. It uses Minitab to generate the output, but doesn't obscure things by incorporating code into the webpage.
https://statisticsbyjim.com/basics/central-limit-theorem/