1/60
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Descriptive Statistics
involves methods of organizing, picturing, illustrating, and summarizing data and information derived from samples or populations
Unlike inferential statistics, the goal of performing this is simply to create a profile of the dataset without drawing further conclusions.
Distribution
set of test scores arrayed for recording or study
Raw Score
straightforward, unmodified (often numerical) account of performance
Simple Frequency Distribution
scores are listed alongside the number of times that each score has occurred
Grouped Frequency Distribution
test score (class) intervals replace actual scores
Class Interval
(Max Value - Min Value) / Number of Levels Set
Histograms
graph with vertical lines drawn at the true limits of each test score / class interval, thus forming a series of contiguous rectangles
Ideal for continuous data
Frequency Polygon
a continuous line connects various points where test scores / class intervals (x-axis) meet frequencies (y-axis)
Ideal for continuous data
Bar Graph
numbers indicative of frequency appear on the y-axis while reference to some categorization or label runs on the x-axis
These graphs tend to be more useful for visualizing nominal or ordinal data.
Graphs
may also be used to summarize data.
Measures of Central Tendency
indicate average / midmost score between extremes in a distribution
Mean (̄x, M)
sum of the frequencies divided by the number of cases
the most common measure of central tendency is the arithmetic _____ , colloquially referred to as the “average.”
This accounts the actual numerical value of every score within a dataset.
Often the most apt measure of central tendency for continuous data when the distributions are believed to be approximately normal.
x̄
represents the sample mean
μ
represents the population mean.
x
raw score in a set of score
Mean Solution
x̄ = ∑ x / N
Median
The middle score in a distribution
This is determined by ordering scores by magnitude in a list according to ascending or descending order.
x͂, Md, Mdn, and Med are symbols used to represent this.
The median may be particularly useful when:
A distribution is skewed or has a few extreme scores
An individual has an unknown or undetermined score
A distribution is said to be open-ended
Scores are measured on an ordinal scale
Position of Median
N + 1 / 2
Mode
Most frequently occurring score in a distribution
Scores tied for “most frequently recurring” can create more than one _____ and are called multimodal distributions (e.g. bimodal = two ____) as they co-occur with the same frequency.
There is no universal symbol for mode, so we can use Mo or Mod
The mode may be particularly useful when:
Data are measured on a nominal scale
Variables produce discrete data
We indicate the shape of the distribution (peak)
Measures of Variability / Dispersion
indicate how scores in a distribution are scattered or dispersed
Range
difference between minimum and maximum scores
distance covered by scores in a distribution
simplest measure of variability to calculate, but potential use is limited
Range Formula
Xmax - Xmin
Interquartile Range (IQR)
difference between Q1 and Q3 of a distribution
range of scores comprising the middle 50% of a distribution
based on quartiles, data points that divide the distribution into four equal sections or quarters (25% of the distribution).
used when measuring central tendency with the median
Standard Deviation and Variance
The ________ provides a measure of the standard or average distance from the mean, and it describes whether the scores are clustered closely or widely scattered around the mean
While this as a concept is straightforward, actual equations are often more complex and lead to related concepts of variance before the standard deviation.
The process of squaring deviation scores gets rid of plus and minus signs, but it also creates a measure of variability based on squared distances which are not very useful or easy to appreciate for descriptives
Deviation Scores
The first step in finding the standard distance from the mean is to determine the ______ or distance from the mean, for each individual score.
The _____ is the difference between a data point (score) and the mean.
There are two parts to this: the sign (telling us if the score is above or below the mean) and the number (giving us the actual distance).
deviation for sample
x – x̄
deviation for population
x – μ
Variance
equals the mean of the squared deviations and is the average squared distance from the mean
Standard Deviation
we can take the square root of the variance and produce a more meaningful number
is the square root of the variance. Conversely, the variance is the square of the _______________
Skewness and Kurtosis
Using frequency polygons can help visualize the various shapes and forms taken by every distribution.
Some distributions are symmetrical while others contain extreme scores in either direction.
Some distributions are said to be ____ (extreme cases in one direction that the other) while variations in symmetrical distributions may also differ in terms of peakedness (_____)
Skewness
Nature and extent to which symmetry is absent
Normal or no skew
if equal on both sides
Positive skew
when relatively few of the scores fall at the high end of the distribution (i.e. most scores are low)
Negative skew
when relatively few of the scores fall at the low end of the distribution (i.e. most scores are high).
Perfectly Symmetrical or Normally Distributed
Ssk = 0
Moderately Symmetrical
–0.5 < Ssk < +0.5
Moderately Skewed
±0.5 < Ssk < ±1.0
Highly Skewed
Ssk < –1.0 ; Ssk > +1.0
Kurtosis
Refers to distribution steepness at its center
The root suffix kurtic (curve / bulge)
Leptokurtic
(from ___-, i.e. thin) has a tall peak or slender curve as well as fat tails
κ = positive values
Platykurtic
(from ____, i.e. flat) has a low peak or broad curve as well as thin tails
κ = negative values
Mesokurtic
(from _____, i.e. middle) has an intermediate peak as well as moderate tails
“tails neither thin nor fat”
κ = value of zero
Measures of Position / Location
determines the position of a single value in relation to other values within a given sample or population dataset
Percentiles
split sorted data into one hundred equal parts
divides dataset into a hundred equal parts (per cent = per a hundred; -ile = point)
Percentiles, quartiles, and deciles share some overlapping equivalences
Q1 = P25 ; Q2 = P50 ; Q3 = P75
D1 = P10 ; D2 = P20 ; D3 = P30 ; D4 = P40 ; and so on
The ____ rank (PR) is the percentage of scores equal to or below the data point in question
Quartiles
split sorted data into four equal parts
divides dataset into four equal parts (per quart = per four ; -ile = point); not to be confused with quarters which are the sections between quartiles
Percentiles, quartiles, and deciles share some overlapping equivalences
Q1 = P25 ; Q2 = P50 ; Q3 = P75 ; Q2 = D5
The median is always conveniently located at the middle of a dataset, making it the 2nd quartile, the 5th decile, and the 50th percentile in distributions
Deciles
split sorted data into ten equal parts
Percentiles, quartiles, and deciles share some overlapping equivalences
Q2 = D5
D1 = P10 ; D2 = P20 ; D3 = P40 ; D5 = P50 ; and so on
z-Scores
A raw dataset on its own can be difficult to work with, so we need a reference point or standard score to make it more meaningful.
Finding distance from the mean yields a value called a ________ or standard score, indicates the direction and degree that any given raw score deviates from the mean of a distribution on a scale of units.
z
represents a coordinate on the standard normal distribution (also known as the Z-distribution).
The Normal Curve
sometimes called the Laplace-Gaussian curve or Gaussian curve
bell-shaped, smooth, symmetrically defined curve
tapers on both sides toward the x-axis asymptotically
Karl Pearson
first to use the term “normal curve”
The Area Under The Normal Curve
divided into areas defined in standard deviation units
tail : area between 2 - 3 SD units
simplifies score interpretation
work well with standard scores
Basic Standard Scores
raw scores converted from one scale to another with an arbitrarily set M and SD
z Score
zero plus or minus one
M = 0 ; SD = 1
The _____ is the common standard score used as it matches the unit distance of the standard deviation. In a normal distribution, this score ranges from -4 to +4
T Score
fifty plus or minus ten
M = 50 ; SD = 10
Used in psychological statistics and assessment. In a normal distribution, its range goes from 10 to 90.
Stanine
standard nine
M = 5 ; SD = ~2
Deviation IQ
scores on IQ tests
M = 100 ; SD = 15 *
Standard Score
Finding distance from the mean yields a value called a ________ that indicates the direction and degree that any given raw score deviates from the mean of a distribution on a scale of units.
z score equation
(x - μ) / σ
t score equation
50 + ( z * 10)