Chapter 2 data set
Frequency Distributions
After collecting data, the first task for a researcher is to organize & simplify the data so that it is possible to get a general overview of the results
This is the goal of descriptive statistical techniques
One method for simplifying and organizing data is to construct a frequency distribution
Frequency Distribution: an organized tabulation showing exactly how many individuals are located in each category on the scale of measurement(Table shows each possible x value/scores and how frequently they occur)
Can be structured either as a table or as a graph, and presents the same two elements
The set of categories that make up the of measurement scale
A record of the frequency or number of individuals in each category
Frequency distribution tables
A frequency distribution table consists of at least 2 columns
Lists the categories on the scale of measurement (X)
Values are listed from highest to lowest without skipping any
Frequency (f)
Tallies are determined for each value (how often each x value occurs in the data set)
The sum of the frequencies should equal N
Frequency Distribution Tables
3. A third column can be used for the proportion (p) for each category:
p = f/N
• Because the proportions describe the frequency (f) in relation to the total
number (N), they often are called relative frequencies
4. A fourth column can display the percentage of the distribution corresponding to each X value
• The percentage is found by multiplying p by 100
• The sum of the percentage column is 100%
𝑝 = 𝑓 ÷ 𝑁
% = 𝑝×100
Grouped Frequency Distribution
Sometimes a set of scores covers a wide range of values
A list of all the X values would be too long to be a “simple” presentation of the data
In a grouped table, the X column lists groups of scores, called class intervals,rather than individual values
The grouped frequency distribution table should have about 10 class intervals
The width of each interval should be a relatively simple number such as 2, 5, 10, or 20
The bottom score is each class interval should be a multiple of the width
All intervals should be the same width
Frequency Distribution Graphs
X axis: Contains the score categories (X)
Y axis: Contains the frequencies
• When the score categories consist of numerical scores from an interval or ratio scale, the graph should be either a histogram or a polygon
Histograms
• In a histogram, a bar is centered above each score (or class interval)
The height of the bar corresponds to the frequency
The width extends to the real limits, so that adjacent bars touch
Polygons
In a polygon, a dot is centered above each score
The height of the dot corresponds to the frequency
A continuous line is drawn from dot to dot to connect the series of dots
The graph is completed by drawing a line down the x-axis (zero frequency) at each end of the range of scores
Bar Graphs
Used when the score categories (X values) are measurements from a nominal or an ordinal scale
A bar graph is just like a histogram, except that gaps or spaces are left between adjacent bars
Nominal Scale: The space emphasizes that the scale consists of separate, distinct categories
Ordinal Scale: Separate bars are used because you cannot assume that the categories are all the same size
Categories - bar graph
Numerical - histogram/polygon
Graphs for Population Distributions
• When you can obtain an exact frequency for each score in a population, you can construct frequency distribution graphs that are exactly the same as the histograms, polygons, and bar graphs that are typically used for samples
Many population are so large that it is impossible to know the exact number of individuals (frequency) for any specific category
Relative Frequencies
When the exact number of individuals is not known, population distributions can be shown using relative frequency instead of the absolute number of individuals for each category
Smooth Curve
• If the scores in a population are measured on an interval or ratio scale, it is customary to present the distribution as a smooth curve rather than a jagged histogram or polygon
• The smooth curve emphasizes the fact that the distribution is not showing the exact frequency for each category
Normal Curve: one commonly occurring population distribution
The word normal refers to a specific shape that can be precisely defined by an equation
Describing Frequency Distributions
• Researchers often simply describe a distribution by listing its characteristics
Characteristics:
1. Central Tendency: measures where the center of the distribution is located
2. Variability: measures the degree to which the scores are spread over a wide range or are clustered together
3. Shape
Shape
A graph shows the shape of the distribution
• Symmetrical: the left side of the graph is (roughly) a mirror image of the right side
• Skewed: the scores tend to pile up toward one end of the scale and taper off gradually at the other end
• Tail: the section where the scores taper off toward one end of a distribution
Positively and Negatively Skewed Distributions
Positively Skewed: the scores tend to pile up on the left side of the
distribution with the tail tapering off to the right
• Example: Wealth among citizens
Negatively Skewed: the scores tend to pile up on the right side and the
tails points to the left
• Example: Life Expectancy
Stem-and-Leaf Displays
Stem-and-Leaf: provides an efficient method for obtaining and
displaying a frequency distribution
Each score is divided into a stem consisting of the first digit/digits, and a leaf consisting of the final digit
Then, go through the list of scores, one at a time, and write the leaf for each score beside its stem
The resulting display provides an organized picture of the entire distribution
The number of leaves beside each stem corresponds to the frequency, and the individual leaves identify the individual scores
Identify the unique stems - 2,3,4,5,6,7
The leaves are all of final values go in second column
Central Tendency
A statistical measure to determine a single score that defines the centre of a distribution
Goal: to find the single score that is most typical or most representative of the entire group
“Average” or “Typical” score
This average value can be used to provide a simple description of an entire population or a sample
Measures of central tendency are also useful for making comparisons between groups of individuals or between sets of data
There is no single, standard procedure for determining central tendency
The problem is that no single measure produces a central, representative value in every situation
There can be problems defining the “centre” of a distribution
To deal with these problems, statisticians have developed 3 different methods for measuring central tendency
• Mean
• Median
• Mode
Funky letters = population
Normal letters = sample
Alternative Definitions of the Mean
Dividing the total equally:
• Think of the mean as the amount each individual received when the total (Σ𝑋) is divided equally among all the individuals (N) in the distribution
The mean is a balance point:
• Think of the mean as a balance point for the distribution
• The total distance below the mean is the same as the total distance above the mean
1. The overall sum of the scores for the combined group (Σ𝑋), and
2. The total number of scores in the combined group (n)
The Weighted Mean
Example 1: Two Samples with the same n
Professor L wants to examine the weighted mean of exam scores across sections 01 and 02 of Psyc*1010 at U of G. She collects a sample of 10 students from each section and has them report their exam grade. Calculate the weighted mean for the two samples.
Section 01:87, 49, 78, 59, 66, 42, 59, 52, 69, 44
Σ𝑋! = 605
𝑀! = 60.50
Section 02:61, 54, 43, 48, 67, 84, 48, 70, 89, 65
Σ𝑋" = 629
𝑀" = 62.90
Weighted Mean (M W ) = 61.70
When the two samples are the same size, the weighted mean will be halfway between the original two sample means
Unless there are the same number of scores for each group, the
overall mean will not be halfway between the original two sample
means
When the samples are not the same size, one makes a larger contribution to the total group and therefore carries more weight in determining the overall mean
Characteristics of the Mean
In general, the characteristics of the mean result from the fact that every score in the distribution contributes to the value of the mean
Specifically, every score adds to the total (Σ𝑋) and every score contributes one point to the number of scores (n)
1. Changing the value of any score will change the mean
2. Adding a new score to a distribution, or removing an existing score, will usually change the mean
The exception is when the new score (or the removed score) is exactly equal to the mean
Original X values: 5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10
𝑀 = Σ𝑋 ÷ 𝑛
= 5.42
Adding X values:5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10, 15
𝑀 = 6.15
Removing X values:5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10,
𝑀 = 5
3. If a constant value is added to every score in a distribution, the same constant will be added to the mean
• Similarly, if you subtract a constant from every score, the same constant will be subtracted from the mean
4. If every score in a distribution is multiplied by (or divided by) a constant value, the mean will change in the same way
The Median
Goal: To locate the midpoint of the distribution
• If the scores in a distribution are listed in order from smallest to largest, the median is the midpoint of the list
• Defining the median as the midpoint of a distribution means that the scores are being divided into two equal-sized groups
• We are not locating the midpoint between the highest and lowest X values
Calculating the Median:
1. With an odd number of scores, list the values in order and the
median is the middle score in the list
X values: 5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8
2,2,3,4,4,5,5,6,8,8,8
5 would be the median
2. With an even number of scores, list the values in order, and the median is half-way between the middle two scores
X values: 61, 98, 75, 77, 66, 75, 70, 83, 52, 53
52,53,61,66,70,75,75,77,83,98
Find the middle of 70 & 75
70+75
-----
2
= 72.50
The Mode
The score or category that has the greatest frequency
MOST OCCURING SCORE
• The only measure of central tendency that will always correspond to an actual score in the data
• The mean and median are both calculated values and often produce an answer that does not equal any score in the distribution
Although a distribution will have only one mean, and only one
median, it is possible to have more than one mode
• Bimodal: A distribution with two modes
• Multimodal: A distribution with more than two modes
The Mean, the Median and the Mode
Mean: A “balance point” – the distances above the mean have the
same total as the distances below the mean
Median: The middle of the distribution (in terms of scores)
Mode: The score/value that occurs most often
Selecting a Measure of Central Tendency
Extreme Scores or Skewed Distributions
• When a distribution has a few extreme scores, scores that are very different in value from most of the others, then the mean may not be a good representative of the majority of the distribution
• Because it is relatively unaffected by extreme scores, the median commonly is used when reporting the average value for a skewed distribution
Median - skewed distribution
Selecting a Measure of Central Tendency
Undetermined Values
• Occasionally, you will encounter a situation in which an individual has an unknown or undetermined score
• This often occurs when you are measuring the number of errors (or amount of time) required for an individual to complete a task
• It is impossible to compute the mean for these data because of the undetermined value
However, it is possible to determine the median
Open-ended distributions
• When there is no upper limit (or lower limit) for one of the categories
• It is impossible to compute a mean for these data because you cannot find Σ𝑋
You can find the median
Ordinal Data
• Many researchers believe that it is not appropriate to use the mean to describe central tendency for ordinal data
• When scores are measured on an ordinal scale, the median is always appropriate and is usually the preferred measure of central tendency
When to use the Mode:
• Nominal Scale
• Always identifies an actual score and is thus useful in describing discrete variables
• The mode gives an indication of the shape of the distribution as well as a measure of central tendency
Graphs can also be used to report and compare measures of central tendency
• The means (or medians) are displayed using a line graph, histogram, or bar graph, depending on the scale of measurement used for the independent variable
• The height of a graph should be approximately two-thirds to three-quarters of its length
• Normally, the zero point for both the x- and y-axis is at the point where the two axes intersect
• However, when a value of zero is part of the data, it is common to move the zero point so that the graph does not overlap the axes
Central Tendency and the Shape of the Distribution
Symmetrical Distribution: The right-hand side is a mirror image of the left-hand side
• The median is exactly at the centre because exactly half of the area in the graph will be on either side of the centre
• The mean is exactly at the centre because each score on the left side of the distribution is balanced by a corresponding score on the right
• If a symmetrical distribution has only one mode, it will also be in the center of the distribution
Measures of Central Tendency for Skewed Distributions
Skewed Distributions: There is a strong tendency for the mean,
median, and mode to be located in predictably different positions
(especially for continuous variables)
• Positively Skewed: The most likely order of the 3 measures of central tendency from smallest to largest (left to right) is the mode, median, and mean
• Negatively Skewed: The most probably order is mean, median, and modeFrequency Distributions
After collecting data, the first task for a researcher is to organize & simplify the data so that it is possible to get a general overview of the results
This is the goal of descriptive statistical techniques
One method for simplifying and organizing data is to construct a frequency distribution
Frequency Distribution: an organized tabulation showing exactly how many individuals are located in each category on the scale of measurement(Table shows each possible x value/scores and how frequently they occur)
Can be structured either as a table or as a graph, and presents the same two elements
The set of categories that make up the of measurement scale
A record of the frequency or number of individuals in each category
Frequency distribution tables
A frequency distribution table consists of at least 2 columns
Lists the categories on the scale of measurement (X)
Values are listed from highest to lowest without skipping any
Frequency (f)
Tallies are determined for each value (how often each x value occurs in the data set)
The sum of the frequencies should equal N
Frequency Distribution Tables
3. A third column can be used for the proportion (p) for each category:
p = f/N
• Because the proportions describe the frequency (f) in relation to the total
number (N), they often are called relative frequencies
4. A fourth column can display the percentage of the distribution corresponding to each X value
• The percentage is found by multiplying p by 100
• The sum of the percentage column is 100%
𝑝 = 𝑓 ÷ 𝑁
% = 𝑝×100
Grouped Frequency Distribution
Sometimes a set of scores covers a wide range of values
A list of all the X values would be too long to be a “simple” presentation of the data
In a grouped table, the X column lists groups of scores, called class intervals,rather than individual values
The grouped frequency distribution table should have about 10 class intervals
The width of each interval should be a relatively simple number such as 2, 5, 10, or 20
The bottom score is each class interval should be a multiple of the width
All intervals should be the same width
Frequency Distribution Graphs
X axis: Contains the score categories (X)
Y axis: Contains the frequencies
• When the score categories consist of numerical scores from an interval or ratio scale, the graph should be either a histogram or a polygon
Histograms
• In a histogram, a bar is centered above each score (or class interval)
The height of the bar corresponds to the frequency
The width extends to the real limits, so that adjacent bars touch
Polygons
In a polygon, a dot is centered above each score
The height of the dot corresponds to the frequency
A continuous line is drawn from dot to dot to connect the series of dots
The graph is completed by drawing a line down the x-axis (zero frequency) at each end of the range of scores
Bar Graphs
Used when the score categories (X values) are measurements from a nominal or an ordinal scale
A bar graph is just like a histogram, except that gaps or spaces are left between adjacent bars
Nominal Scale: The space emphasizes that the scale consists of separate, distinct categories
Ordinal Scale: Separate bars are used because you cannot assume that the categories are all the same size
Categories - bar graph
Numerical - histogram/polygon
Graphs for Population Distributions
• When you can obtain an exact frequency for each score in a population, you can construct frequency distribution graphs that are exactly the same as the histograms, polygons, and bar graphs that are typically used for samples
Many population are so large that it is impossible to know the exact number of individuals (frequency) for any specific category
Relative Frequencies
When the exact number of individuals is not known, population distributions can be shown using relative frequency instead of the absolute number of individuals for each category
Smooth Curve
• If the scores in a population are measured on an interval or ratio scale, it is customary to present the distribution as a smooth curve rather than a jagged histogram or polygon
• The smooth curve emphasizes the fact that the distribution is not showing the exact frequency for each category
Normal Curve: one commonly occurring population distribution
The word normal refers to a specific shape that can be precisely defined by an equation
Describing Frequency Distributions
• Researchers often simply describe a distribution by listing its characteristics
Characteristics:
1. Central Tendency: measures where the center of the distribution is located
2. Variability: measures the degree to which the scores are spread over a wide range or are clustered together
3. Shape
Shape
A graph shows the shape of the distribution
• Symmetrical: the left side of the graph is (roughly) a mirror image of the right side
• Skewed: the scores tend to pile up toward one end of the scale and taper off gradually at the other end
• Tail: the section where the scores taper off toward one end of a distribution
Positively and Negatively Skewed Distributions
Positively Skewed: the scores tend to pile up on the left side of the
distribution with the tail tapering off to the right
• Example: Wealth among citizens
Negatively Skewed: the scores tend to pile up on the right side and the
tails points to the left
• Example: Life Expectancy
Stem-and-Leaf Displays
Stem-and-Leaf: provides an efficient method for obtaining and
displaying a frequency distribution
Each score is divided into a stem consisting of the first digit/digits, and a leaf consisting of the final digit
Then, go through the list of scores, one at a time, and write the leaf for each score beside its stem
The resulting display provides an organized picture of the entire distribution
The number of leaves beside each stem corresponds to the frequency, and the individual leaves identify the individual scores
Identify the unique stems - 2,3,4,5,6,7
The leaves are all of final values go in second column
Central Tendency
A statistical measure to determine a single score that defines the centre of a distribution
Goal: to find the single score that is most typical or most representative of the entire group
“Average” or “Typical” score
This average value can be used to provide a simple description of an entire population or a sample
Measures of central tendency are also useful for making comparisons between groups of individuals or between sets of data
There is no single, standard procedure for determining central tendency
The problem is that no single measure produces a central, representative value in every situation
There can be problems defining the “centre” of a distribution
To deal with these problems, statisticians have developed 3 different methods for measuring central tendency
• Mean
• Median
• Mode
Funky letters = population
Normal letters = sample
Alternative Definitions of the Mean
Dividing the total equally:
• Think of the mean as the amount each individual received when the total (Σ𝑋) is divided equally among all the individuals (N) in the distribution
The mean is a balance point:
• Think of the mean as a balance point for the distribution
• The total distance below the mean is the same as the total distance above the mean
1. The overall sum of the scores for the combined group (Σ𝑋), and
2. The total number of scores in the combined group (n)
The Weighted Mean
Example 1: Two Samples with the same n
Professor L wants to examine the weighted mean of exam scores across sections 01 and 02 of Psyc*1010 at U of G. She collects a sample of 10 students from each section and has them report their exam grade. Calculate the weighted mean for the two samples.
Section 01:87, 49, 78, 59, 66, 42, 59, 52, 69, 44
Σ𝑋! = 605
𝑀! = 60.50
Section 02:61, 54, 43, 48, 67, 84, 48, 70, 89, 65
Σ𝑋" = 629
𝑀" = 62.90
Weighted Mean (M W ) = 61.70
When the two samples are the same size, the weighted mean will be halfway between the original two sample means
Unless there are the same number of scores for each group, the
overall mean will not be halfway between the original two sample
means
When the samples are not the same size, one makes a larger contribution to the total group and therefore carries more weight in determining the overall mean
Characteristics of the Mean
In general, the characteristics of the mean result from the fact that every score in the distribution contributes to the value of the mean
Specifically, every score adds to the total (Σ𝑋) and every score contributes one point to the number of scores (n)
1. Changing the value of any score will change the mean
2. Adding a new score to a distribution, or removing an existing score, will usually change the mean
The exception is when the new score (or the removed score) is exactly equal to the mean
Original X values: 5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10
𝑀 = Σ𝑋 ÷ 𝑛
= 5.42
Adding X values:5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10, 15
𝑀 = 6.15
Removing X values:5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10,
𝑀 = 5
3. If a constant value is added to every score in a distribution, the same constant will be added to the mean
• Similarly, if you subtract a constant from every score, the same constant will be subtracted from the mean
4. If every score in a distribution is multiplied by (or divided by) a constant value, the mean will change in the same way
The Median
Goal: To locate the midpoint of the distribution
• If the scores in a distribution are listed in order from smallest to largest, the median is the midpoint of the list
• Defining the median as the midpoint of a distribution means that the scores are being divided into two equal-sized groups
• We are not locating the midpoint between the highest and lowest X values
Calculating the Median:
1. With an odd number of scores, list the values in order and the
median is the middle score in the list
X values: 5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8
2,2,3,4,4,5,5,6,8,8,8
5 would be the median
2. With an even number of scores, list the values in order, and the median is half-way between the middle two scores
X values: 61, 98, 75, 77, 66, 75, 70, 83, 52, 53
52,53,61,66,70,75,75,77,83,98
Find the middle of 70 & 75
70+75
-----
2
= 72.50
The Mode
The score or category that has the greatest frequency
MOST OCCURING SCORE
• The only measure of central tendency that will always correspond to an actual score in the data
• The mean and median are both calculated values and often produce an answer that does not equal any score in the distribution
Although a distribution will have only one mean, and only one
median, it is possible to have more than one mode
• Bimodal: A distribution with two modes
• Multimodal: A distribution with more than two modes
The Mean, the Median and the Mode
Mean: A “balance point” – the distances above the mean have the
same total as the distances below the mean
Median: The middle of the distribution (in terms of scores)
Mode: The score/value that occurs most often
Selecting a Measure of Central Tendency
Extreme Scores or Skewed Distributions
• When a distribution has a few extreme scores, scores that are very different in value from most of the others, then the mean may not be a good representative of the majority of the distribution
• Because it is relatively unaffected by extreme scores, the median commonly is used when reporting the average value for a skewed distribution
Median - skewed distribution
Selecting a Measure of Central Tendency
Undetermined Values
• Occasionally, you will encounter a situation in which an individual has an unknown or undetermined score
• This often occurs when you are measuring the number of errors (or amount of time) required for an individual to complete a task
• It is impossible to compute the mean for these data because of the undetermined value
However, it is possible to determine the median
Open-ended distributions
• When there is no upper limit (or lower limit) for one of the categories
• It is impossible to compute a mean for these data because you cannot find Σ𝑋
You can find the median
Ordinal Data
• Many researchers believe that it is not appropriate to use the mean to describe central tendency for ordinal data
• When scores are measured on an ordinal scale, the median is always appropriate and is usually the preferred measure of central tendency
When to use the Mode:
• Nominal Scale
• Always identifies an actual score and is thus useful in describing discrete variables
• The mode gives an indication of the shape of the distribution as well as a measure of central tendency
Graphs can also be used to report and compare measures of central tendency
• The means (or medians) are displayed using a line graph, histogram, or bar graph, depending on the scale of measurement used for the independent variable
• The height of a graph should be approximately two-thirds to three-quarters of its length
• Normally, the zero point for both the x- and y-axis is at the point where the two axes intersect
• However, when a value of zero is part of the data, it is common to move the zero point so that the graph does not overlap the axes
Central Tendency and the Shape of the Distribution
Symmetrical Distribution: The right-hand side is a mirror image of the left-hand side
• The median is exactly at the centre because exactly half of the area in the graph will be on either side of the centre
• The mean is exactly at the centre because each score on the left side of the distribution is balanced by a corresponding score on the right
• If a symmetrical distribution has only one mode, it will also be in the center of the distribution
Measures of Central Tendency for Skewed Distributions
Skewed Distributions: There is a strong tendency for the mean,
median, and mode to be located in predictably different positions
(especially for continuous variables)
• Positively Skewed: The most likely order of the 3 measures of central tendency from smallest to largest (left to right) is the mode, median, and mean
• Negatively Skewed: The most probably order is mean, median, and mode
