Chapter 2 data set

Frequency Distributions

 

  • After collecting data, the first task for a researcher is to organize & simplify the data so that it is possible to get a general overview of the results

  • This is the goal of descriptive statistical techniques

  • One method for simplifying and organizing data is to construct a frequency distribution

 

 

Frequency Distribution: an organized tabulation showing exactly how many individuals are located in each category on the scale of measurement(Table shows each possible x value/scores and how frequently they occur)

  • Can be structured either as a table or as a graph, and presents the same two elements

 

  1. The set of categories that make up the of measurement scale

  2. A record of the frequency or number of individuals in each category

 

 

Frequency distribution tables

 

A frequency distribution table consists of at least 2 columns

 

  1. Lists the categories on the scale of measurement (X)

  • Values are listed from highest to lowest without skipping any

  1. Frequency (f)

  • Tallies are determined for each value (how often each x value occurs in the data set)

  • The sum of the frequencies should equal N

 

Frequency Distribution Tables

 

3. A third column can be used for the proportion (p) for each category:

 

p = f/N

• Because the proportions describe the frequency (f) in relation to the total

number (N), they often are called relative frequencies

 

4. A fourth column can display the percentage of the distribution corresponding to each X value

• The percentage is found by multiplying p by 100

• The sum of the percentage column is 100%

 

 

𝑝 = 𝑓 ÷ 𝑁

% = 𝑝×100

 

 

Grouped Frequency Distribution

 

Sometimes a set of scores covers a wide range of values

  •  A list of all the X values would be too long to be a “simple” presentation of the data

  • In a grouped table, the X column lists groups of scores, called class intervals,rather than individual values

  •  The grouped frequency distribution table should have about 10 class intervals

  •  The width of each interval should be a relatively simple number such as 2, 5, 10, or 20

  •  The bottom score is each class interval should be a multiple of the width

  •  All intervals should be the same width

 

 


Frequency Distribution Graphs

 

X axis: Contains the score categories (X)

Y axis: Contains the frequencies

 

• When the score categories consist of numerical scores from an interval or ratio scale, the graph should be either a histogram or a polygon

 

Histograms

 

• In a histogram, a bar is centered above each score (or class interval)

  •  The height of the bar corresponds to the frequency

  •  The width extends to the real limits, so that adjacent bars touch

 


 

 

Polygons

In a polygon, a dot is centered above each score

  • The height of the dot corresponds to the frequency

  •  A continuous line is drawn from dot to dot to connect the series of dots

  •  The graph is completed by drawing a line down the x-axis (zero frequency) at each end of the range of scores

 

 


 

 

Bar Graphs

  •  Used when the score categories (X values) are measurements from a nominal or an ordinal scale

 

  • A bar graph is just like a histogram, except that gaps or spaces are left between adjacent bars

    • Nominal Scale: The space emphasizes that the scale consists of separate, distinct categories

    • Ordinal Scale: Separate bars are used because you cannot assume that the categories are all the same size


 

 

 

Categories - bar graph

Numerical - histogram/polygon

 

Graphs for Population Distributions

 

• When you can obtain an exact frequency for each score in a population, you can construct frequency distribution graphs that are exactly the same as the histograms, polygons, and bar graphs that are typically used for samples

  • Many population are so large that it is impossible to know the exact number of individuals (frequency) for any specific category

 

 

Relative Frequencies

 

  • When the exact number of individuals is not known, population distributions can be shown using relative frequency instead of the absolute number of individuals for each category

 

 


Smooth Curve

 

• If the scores in a population are measured on an interval or ratio scale, it is customary to present the distribution as a smooth curve rather than a jagged histogram or polygon

 

• The smooth curve emphasizes the fact that the distribution is not showing the exact frequency for each category

 

Normal Curve: one commonly occurring population distribution

  •  The word normal refers to a specific shape that can be precisely defined by an equation

 


 

 

Describing Frequency Distributions

 

• Researchers often simply describe a distribution by listing its characteristics

 

Characteristics:

 

1. Central Tendency: measures where the center of the distribution is located

2. Variability: measures the degree to which the scores are spread over a wide range or are clustered together

3. Shape

 

 

Shape

 

A graph shows the shape of the distribution

 

• Symmetrical: the left side of the graph is (roughly) a mirror image of the right side

 

• Skewed: the scores tend to pile up toward one end of the scale and taper off gradually at the other end

 

• Tail: the section where the scores taper off toward one end of a distribution

 

Positively and Negatively Skewed Distributions

 

Positively Skewed: the scores tend to pile up on the left side of the

distribution with the tail tapering off to the right

• Example: Wealth among citizens

 

Negatively Skewed: the scores tend to pile up on the right side and the

tails points to the left

• Example: Life Expectancy

 

 

 


 

Stem-and-Leaf Displays

 

Stem-and-Leaf: provides an efficient method for obtaining and

displaying a frequency distribution

  •  Each score is divided into a stem consisting of the first digit/digits, and a leaf consisting of the final digit

  •  Then, go through the list of scores, one at a time, and write the leaf for each score beside its stem

 

  • The resulting display provides an organized picture of the entire distribution

  •  The number of leaves beside each stem corresponds to the frequency, and the individual leaves identify the individual scores

 


Identify the unique stems - 2,3,4,5,6,7

 

The leaves are all of final values go in second column

 

 

Central Tendency

 

A statistical measure to determine a single score that defines the centre of a distribution

  •  Goal: to find the single score that is most typical or most representative of the entire group

 

“Average” or “Typical” score

 

  •  This average value can be used to provide a simple description of an entire population or a sample

 

  •  Measures of central tendency are also useful for making comparisons between groups of individuals or between sets of data

 

There is no single, standard procedure for determining central tendency

  • The problem is that no single measure produces a central, representative value in every situation

 

 There can be problems defining the “centre” of a distribution

  •  To deal with these problems, statisticians have developed 3 different methods for measuring central tendency

• Mean

• Median

• Mode


 

Funky letters = population

Normal letters = sample

 

Alternative Definitions of the Mean

 

Dividing the total equally:

 

• Think of the mean as the amount each individual received when the total (Σ𝑋) is divided equally among all the individuals (N) in the distribution

 

The mean is a balance point:

 

• Think of the mean as a balance point for the distribution

• The total distance below the mean is the same as the total distance above the mean

 


 

 


 

 

1. The overall sum of the scores for the combined group (Σ𝑋), and

2. The total number of scores in the combined group (n)

 

The Weighted Mean

 

Example 1: Two Samples with the same n

 

Professor L wants to examine the weighted mean of exam scores across sections 01 and 02 of Psyc*1010 at U of G. She collects a sample of 10 students from each section and has them report their exam grade. Calculate the weighted mean for the two samples.

 

 

Section 01:87, 49, 78, 59, 66, 42, 59, 52, 69, 44

Σ𝑋! = 605

𝑀! = 60.50
 

Section 02:61, 54, 43, 48, 67, 84, 48, 70, 89, 65

Σ𝑋" = 629

𝑀" = 62.90

 

Weighted Mean (M W ) = 61.70

 

  •  When the two samples are the same size, the weighted mean will be halfway between the original two sample means

 

 Unless there are the same number of scores for each group, the

overall mean will not be halfway between the original two sample

means

 

  •  When the samples are not the same size, one makes a larger contribution to the total group and therefore carries more weight in determining the overall mean

 

 

 

 


 

 


 

Characteristics of the Mean

 

 In general, the characteristics of the mean result from the fact that every score in the distribution contributes to the value of the mean

 

Specifically, every score adds to the total (Σ𝑋) and every score contributes one point to the number of scores (n)

 

1. Changing the value of any score will change the mean

 

2. Adding a new score to a distribution, or removing an existing score, will usually change the mean

 

 The exception is when the new score (or the removed score) is exactly equal to the mean

 

Original X values: 5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10

𝑀 = Σ𝑋 ÷ 𝑛

= 5.42

 

Adding X values:5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10, 15

𝑀 = 6.15

Removing X values:5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10,

𝑀 = 5

 

3. If a constant value is added to every score in a distribution, the same constant will be added to the mean

• Similarly, if you subtract a constant from every score, the same constant will be subtracted from the mean

 

4. If every score in a distribution is multiplied by (or divided by) a constant value, the mean will change in the same way

 

 

The Median

 

Goal: To locate the midpoint of the distribution

• If the scores in a distribution are listed in order from smallest to largest, the median is the midpoint of the list

• Defining the median as the midpoint of a distribution means that the scores are being divided into two equal-sized groups

• We are not locating the midpoint between the highest and lowest X values

 

Calculating the Median:

 

1. With an odd number of scores, list the values in order and the

median is the middle score in the list

 

X values: 5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8

 

2,2,3,4,4,5,5,6,8,8,8

 

5 would be the median

 

2. With an even number of scores, list the values in order, and the median is half-way between the middle two scores

 

X values: 61, 98, 75, 77, 66, 75, 70, 83, 52, 53

 

52,53,61,66,70,75,75,77,83,98

 

Find the middle of 70 & 75

 

70+75

-----

    2

= 72.50

 

 

The Mode

 

The score or category that has the greatest frequency

 

MOST OCCURING SCORE

 

• The only measure of central tendency that will always correspond to an actual score in the data

 

• The mean and median are both calculated values and often produce an answer that does not equal any score in the distribution

 

Although a distribution will have only one mean, and only one

median, it is possible to have more than one mode

 

• Bimodal: A distribution with two modes

Multimodal: A distribution with more than two modes


 

 

 

The Mean, the Median and the Mode

 

Mean: A “balance point” – the distances above the mean have the

same total as the distances below the mean

 

Median: The middle of the distribution (in terms of scores)

 

Mode: The score/value that occurs most often

 

 

Selecting a Measure of Central Tendency

 

Extreme Scores or Skewed Distributions

 

• When a distribution has a few extreme scores, scores that are very different in value from most of the others, then the mean may not be a good representative of the majority of the distribution

 

• Because it is relatively unaffected by extreme scores, the median commonly is used when reporting the average value for a skewed distribution

 

Median - skewed distribution

 


 

Selecting a Measure of Central Tendency

 

Undetermined Values

 

• Occasionally, you will encounter a situation in which an individual has an unknown or undetermined score

 

• This often occurs when you are measuring the number of errors (or amount of time) required for an individual to complete a task

 

• It is impossible to compute the mean for these data because of the undetermined value

  •  However, it is possible to determine the median

 

Open-ended distributions

 

• When there is no upper limit (or lower limit) for one of the categories

• It is impossible to compute a mean for these data because you cannot find Σ𝑋

  •  You can find the median

 

 

Ordinal Data

 

• Many researchers believe that it is not appropriate to use the mean to describe central tendency for ordinal data

• When scores are measured on an ordinal scale, the median is always appropriate and is usually the preferred measure of central tendency

 

 

When to use the Mode:

• Nominal Scale

• Always identifies an actual score and is thus useful in describing discrete variables

• The mode gives an indication of the shape of the distribution as well as a measure of central tendency


Graphs can also be used to report and compare measures of central tendency

 

• The means (or medians) are displayed using a line graph, histogram, or bar graph, depending on the scale of measurement used for the independent variable

 

 

• The height of a graph should be approximately two-thirds to three-quarters of its length

 

• Normally, the zero point for both the x- and y-axis is at the point where the two axes intersect

 

• However, when a value of zero is part of the data, it is common to move the zero point so that the graph does not overlap the axes


Central Tendency and the Shape of the Distribution

 

Symmetrical Distribution: The right-hand side is a mirror image of the left-hand side

 

• The median is exactly at the centre because exactly half of the area in the graph will be on either side of the centre

 

• The mean is exactly at the centre because each score on the left side of the distribution is balanced by a corresponding score on the right

 

• If a symmetrical distribution has only one mode, it will also be in the center of the distribution

 

Measures of Central Tendency for Skewed Distributions

 

Skewed Distributions: There is a strong tendency for the mean,

median, and mode to be located in predictably different positions

(especially for continuous variables)

 

• Positively Skewed: The most likely order of the 3 measures of central tendency from smallest to largest (left to right) is the mode, median, and mean

 

• Negatively Skewed: The most probably order is mean, median, and modeFrequency Distributions

 

  • After collecting data, the first task for a researcher is to organize & simplify the data so that it is possible to get a general overview of the results

  • This is the goal of descriptive statistical techniques

  • One method for simplifying and organizing data is to construct a frequency distribution

 

 

Frequency Distribution: an organized tabulation showing exactly how many individuals are located in each category on the scale of measurement(Table shows each possible x value/scores and how frequently they occur)

  • Can be structured either as a table or as a graph, and presents the same two elements

 

  1. The set of categories that make up the of measurement scale

  2. A record of the frequency or number of individuals in each category

 

 

Frequency distribution tables

 

A frequency distribution table consists of at least 2 columns

 

  1. Lists the categories on the scale of measurement (X)

  • Values are listed from highest to lowest without skipping any

  1. Frequency (f)

  • Tallies are determined for each value (how often each x value occurs in the data set)

  • The sum of the frequencies should equal N

 

Frequency Distribution Tables

 

3. A third column can be used for the proportion (p) for each category:

 

p = f/N

• Because the proportions describe the frequency (f) in relation to the total

number (N), they often are called relative frequencies

 

4. A fourth column can display the percentage of the distribution corresponding to each X value

• The percentage is found by multiplying p by 100

• The sum of the percentage column is 100%

 

 

𝑝 = 𝑓 ÷ 𝑁

% = 𝑝×100

 

 

Grouped Frequency Distribution

 

Sometimes a set of scores covers a wide range of values

  •  A list of all the X values would be too long to be a “simple” presentation of the data

  • In a grouped table, the X column lists groups of scores, called class intervals,rather than individual values

  •  The grouped frequency distribution table should have about 10 class intervals

  •  The width of each interval should be a relatively simple number such as 2, 5, 10, or 20

  •  The bottom score is each class interval should be a multiple of the width

  •  All intervals should be the same width

 

 


Frequency Distribution Graphs

 

X axis: Contains the score categories (X)

Y axis: Contains the frequencies

 

• When the score categories consist of numerical scores from an interval or ratio scale, the graph should be either a histogram or a polygon

 

Histograms

 

• In a histogram, a bar is centered above each score (or class interval)

  •  The height of the bar corresponds to the frequency

  •  The width extends to the real limits, so that adjacent bars touch

 


 

 

Polygons

In a polygon, a dot is centered above each score

  • The height of the dot corresponds to the frequency

  •  A continuous line is drawn from dot to dot to connect the series of dots

  •  The graph is completed by drawing a line down the x-axis (zero frequency) at each end of the range of scores

 

 


 

 

Bar Graphs

  •  Used when the score categories (X values) are measurements from a nominal or an ordinal scale

 

  • A bar graph is just like a histogram, except that gaps or spaces are left between adjacent bars

    • Nominal Scale: The space emphasizes that the scale consists of separate, distinct categories

    • Ordinal Scale: Separate bars are used because you cannot assume that the categories are all the same size


 

 

 

Categories - bar graph

Numerical - histogram/polygon

 

Graphs for Population Distributions

 

• When you can obtain an exact frequency for each score in a population, you can construct frequency distribution graphs that are exactly the same as the histograms, polygons, and bar graphs that are typically used for samples

  • Many population are so large that it is impossible to know the exact number of individuals (frequency) for any specific category

 

 

Relative Frequencies

 

  • When the exact number of individuals is not known, population distributions can be shown using relative frequency instead of the absolute number of individuals for each category

 

 


Smooth Curve

 

• If the scores in a population are measured on an interval or ratio scale, it is customary to present the distribution as a smooth curve rather than a jagged histogram or polygon

 

• The smooth curve emphasizes the fact that the distribution is not showing the exact frequency for each category

 

Normal Curve: one commonly occurring population distribution

  •  The word normal refers to a specific shape that can be precisely defined by an equation

 


 

 

Describing Frequency Distributions

 

• Researchers often simply describe a distribution by listing its characteristics

 

Characteristics:

 

1. Central Tendency: measures where the center of the distribution is located

2. Variability: measures the degree to which the scores are spread over a wide range or are clustered together

3. Shape

 

 

Shape

 

A graph shows the shape of the distribution

 

• Symmetrical: the left side of the graph is (roughly) a mirror image of the right side

 

• Skewed: the scores tend to pile up toward one end of the scale and taper off gradually at the other end

 

• Tail: the section where the scores taper off toward one end of a distribution

 

Positively and Negatively Skewed Distributions

 

Positively Skewed: the scores tend to pile up on the left side of the

distribution with the tail tapering off to the right

• Example: Wealth among citizens

 

Negatively Skewed: the scores tend to pile up on the right side and the

tails points to the left

• Example: Life Expectancy

 

 

 


 

Stem-and-Leaf Displays

 

Stem-and-Leaf: provides an efficient method for obtaining and

displaying a frequency distribution

  •  Each score is divided into a stem consisting of the first digit/digits, and a leaf consisting of the final digit

  •  Then, go through the list of scores, one at a time, and write the leaf for each score beside its stem

 

  • The resulting display provides an organized picture of the entire distribution

  •  The number of leaves beside each stem corresponds to the frequency, and the individual leaves identify the individual scores

 


Identify the unique stems - 2,3,4,5,6,7

 

The leaves are all of final values go in second column

 

 

Central Tendency

 

A statistical measure to determine a single score that defines the centre of a distribution

  •  Goal: to find the single score that is most typical or most representative of the entire group

 

“Average” or “Typical” score

 

  •  This average value can be used to provide a simple description of an entire population or a sample

 

  •  Measures of central tendency are also useful for making comparisons between groups of individuals or between sets of data

 

There is no single, standard procedure for determining central tendency

  • The problem is that no single measure produces a central, representative value in every situation

 

 There can be problems defining the “centre” of a distribution

  •  To deal with these problems, statisticians have developed 3 different methods for measuring central tendency

• Mean

• Median

• Mode


 

Funky letters = population

Normal letters = sample

 

Alternative Definitions of the Mean

 

Dividing the total equally:

 

• Think of the mean as the amount each individual received when the total (Σ𝑋) is divided equally among all the individuals (N) in the distribution

 

The mean is a balance point:

 

• Think of the mean as a balance point for the distribution

• The total distance below the mean is the same as the total distance above the mean

 


 

 


 

 

1. The overall sum of the scores for the combined group (Σ𝑋), and

2. The total number of scores in the combined group (n)

 

The Weighted Mean

 

Example 1: Two Samples with the same n

 

Professor L wants to examine the weighted mean of exam scores across sections 01 and 02 of Psyc*1010 at U of G. She collects a sample of 10 students from each section and has them report their exam grade. Calculate the weighted mean for the two samples.

 

 

Section 01:87, 49, 78, 59, 66, 42, 59, 52, 69, 44

Σ𝑋! = 605

𝑀! = 60.50
 

Section 02:61, 54, 43, 48, 67, 84, 48, 70, 89, 65

Σ𝑋" = 629

𝑀" = 62.90

 

Weighted Mean (M W ) = 61.70

 

  •  When the two samples are the same size, the weighted mean will be halfway between the original two sample means

 

 Unless there are the same number of scores for each group, the

overall mean will not be halfway between the original two sample

means

 

  •  When the samples are not the same size, one makes a larger contribution to the total group and therefore carries more weight in determining the overall mean

 

 

 

 


 

 


 

Characteristics of the Mean

 

 In general, the characteristics of the mean result from the fact that every score in the distribution contributes to the value of the mean

 

Specifically, every score adds to the total (Σ𝑋) and every score contributes one point to the number of scores (n)

 

1. Changing the value of any score will change the mean

 

2. Adding a new score to a distribution, or removing an existing score, will usually change the mean

 

 The exception is when the new score (or the removed score) is exactly equal to the mean

 

Original X values: 5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10

𝑀 = Σ𝑋 ÷ 𝑛

= 5.42

 

Adding X values:5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10, 15

𝑀 = 6.15

Removing X values:5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8, 10,

𝑀 = 5

 

3. If a constant value is added to every score in a distribution, the same constant will be added to the mean

• Similarly, if you subtract a constant from every score, the same constant will be subtracted from the mean

 

4. If every score in a distribution is multiplied by (or divided by) a constant value, the mean will change in the same way

 

 

The Median

 

Goal: To locate the midpoint of the distribution

• If the scores in a distribution are listed in order from smallest to largest, the median is the midpoint of the list

• Defining the median as the midpoint of a distribution means that the scores are being divided into two equal-sized groups

• We are not locating the midpoint between the highest and lowest X values

 

Calculating the Median:

 

1. With an odd number of scores, list the values in order and the

median is the middle score in the list

 

X values: 5, 3, 4, 8, 5, 2, 8, 4, 2, 6, 8

 

2,2,3,4,4,5,5,6,8,8,8

 

5 would be the median

 

2. With an even number of scores, list the values in order, and the median is half-way between the middle two scores

 

X values: 61, 98, 75, 77, 66, 75, 70, 83, 52, 53

 

52,53,61,66,70,75,75,77,83,98

 

Find the middle of 70 & 75

 

70+75

-----

    2

= 72.50

 

 

The Mode

 

The score or category that has the greatest frequency

 

MOST OCCURING SCORE

 

• The only measure of central tendency that will always correspond to an actual score in the data

 

• The mean and median are both calculated values and often produce an answer that does not equal any score in the distribution

 

Although a distribution will have only one mean, and only one

median, it is possible to have more than one mode

 

• Bimodal: A distribution with two modes

Multimodal: A distribution with more than two modes


 

 

 

The Mean, the Median and the Mode

 

Mean: A “balance point” – the distances above the mean have the

same total as the distances below the mean

 

Median: The middle of the distribution (in terms of scores)

 

Mode: The score/value that occurs most often

 

 

Selecting a Measure of Central Tendency

 

Extreme Scores or Skewed Distributions

 

• When a distribution has a few extreme scores, scores that are very different in value from most of the others, then the mean may not be a good representative of the majority of the distribution

 

• Because it is relatively unaffected by extreme scores, the median commonly is used when reporting the average value for a skewed distribution

 

Median - skewed distribution

 


 

Selecting a Measure of Central Tendency

 

Undetermined Values

 

• Occasionally, you will encounter a situation in which an individual has an unknown or undetermined score

 

• This often occurs when you are measuring the number of errors (or amount of time) required for an individual to complete a task

 

• It is impossible to compute the mean for these data because of the undetermined value

  •  However, it is possible to determine the median

 

Open-ended distributions

 

• When there is no upper limit (or lower limit) for one of the categories

• It is impossible to compute a mean for these data because you cannot find Σ𝑋

  •  You can find the median

 

 

Ordinal Data

 

• Many researchers believe that it is not appropriate to use the mean to describe central tendency for ordinal data

• When scores are measured on an ordinal scale, the median is always appropriate and is usually the preferred measure of central tendency

 

 

When to use the Mode:

• Nominal Scale

• Always identifies an actual score and is thus useful in describing discrete variables

• The mode gives an indication of the shape of the distribution as well as a measure of central tendency


Graphs can also be used to report and compare measures of central tendency

 

• The means (or medians) are displayed using a line graph, histogram, or bar graph, depending on the scale of measurement used for the independent variable

 

 

• The height of a graph should be approximately two-thirds to three-quarters of its length

 

• Normally, the zero point for both the x- and y-axis is at the point where the two axes intersect

 

• However, when a value of zero is part of the data, it is common to move the zero point so that the graph does not overlap the axes


Central Tendency and the Shape of the Distribution

 

Symmetrical Distribution: The right-hand side is a mirror image of the left-hand side

 

• The median is exactly at the centre because exactly half of the area in the graph will be on either side of the centre

 

• The mean is exactly at the centre because each score on the left side of the distribution is balanced by a corresponding score on the right

 

• If a symmetrical distribution has only one mode, it will also be in the center of the distribution

 

Measures of Central Tendency for Skewed Distributions

 

Skewed Distributions: There is a strong tendency for the mean,

median, and mode to be located in predictably different positions

(especially for continuous variables)

 

• Positively Skewed: The most likely order of the 3 measures of central tendency from smallest to largest (left to right) is the mode, median, and mean

 

• Negatively Skewed: The most probably order is mean, median, and mode