Statistics for Kinesiology
Why Stats
Statistics → collection of information (data) and the methods used to analyze them. A discipline. “The science of making an educated guess”
Having good data skills (gathering, analyzing, summarizing, displaying & interpreting) makes you an intelligent decision maker
Statistics are formal methods/approaches to:
collect/summarize data
Objectively analyze data
Make unbiased interpretations leading to a decision
Key terms
Data - result of measurement
Evaluation - philosophical process of determining the worth of data
Reliability - consistency
Objective - no bias
Subjective - with bias
Variable - characteristic that can assume more than one value
Constant - characteristic that can only assume one value
Observational - research that involves passive collection of data or self-reported data
Experimental - research process that involves active intervention and control
Validity - Appropriateness of the test in measuring what it is designed to measure
Internal validity - a measure of the control within the experiment to be sure that the results are due to the intervention
Affected by confounding variables, instrument error, investigator error, bias
External validity - The ability to generalize the sample results to the population interest
Affected by appropriate inclusion and exclusion criteria
Population Vs Sample
Why not the whole population
Numbers
Time
Manpower
Cost
Willingness of all members of the population to participate
Destructive nature of experiment
Fundamental assumption of sampling
Sample is representative of the population (PARAMETRIC)
Otherwise we will apply non-parametric text
Random Sample
Random sample is scientifically drawn group that actually possesses the same characteristics as the population
Every person has an equal selection opportunity for you sample
Selection of one person is independent of the selection of another person
If repeatedly draw sub-samples from the population, the means of multiple sample will approximate the population mean
Sample error
E.g. suppose we survey the weight of all UOGH students, N = 5000, Avg = 134.21lbs
Measured N = 200 students, sample avg = 135
The difference between a population parameter and a sample statistic is known as Sampling error
When n and repeated sampling goes up, sampling error goes down
Data levels of measurement
Nominal → Attributes are named, weakest
Ordinal → attributes can be ordered
Interval → distance is meaningful
Ration → Absolute zero
Important Because data type defines which statistical test can be performed
Nominal
Allows to distinguish differences between items qualitatively
There is no quantitative ordering or value, no hierarchy
Assign responses to to different categories = discrete data
Examples → city of birth, postal code, biological sex, favourite movies, id number
Ordinal
The categories have logical, meaningful order or rank
Starts at lowest ends at highest
BUT don’t know the numerical distance between categories, so still discrete data
Still cannot do mathematical calculation but can compare
Examples → likert scales, starred reviews, perceived exertion, letter Grades
Interval
Measurements are numerical values
Intervals of equal length represent equal differences in the characteristic. This is the key difference between interval and ordinal
Data is usually continuous (can hold any value including fractions/decimals
No true zero – zero is not the absence of what is being measured
Starting point is arbitrary
Data can be added or subtracted, but not multiplied or divided
Examples – standardized scores like IQ, temperature
Ratio
Allows for identification of absolute differences
absolute/true zero
Absence of a characteristics
Usually continuous
Can be added subtracted multiplied or divided
Most measured data fits into this category
Examples – time to completion, age, weight, distance
Organizing data
Frequency distribution
Rank order distribution
Simplest distribution
Only feasible with very small data sets
Ex) 13, 2, 9, 17, 5
Rank ordered = 2, 5, 9, 13, 17
Determine number of data points, N = 5
Can determined high = 17, low = 2
Can determine range (R) = H - L = 17 -2 = 15
Sample frequency distribution
For larger data sets
Tabulates the occurrence of each number
N = the sum of all frequencies
Ex) 5, 4, 9, 6, 4, 5, 2, 9, 9, 3, 9
N = 11
H = 9
L = 2
R = 7
Grouped frequency distribution
For very large data sets, particularly with very large ranges
Tabulates the occurrence of each number within a range
Lose some information - do not know exact values of number
Types of frequency distribution graphs
Histogram
Frequency, probability density (for continuous data)
Bar graph
Frequency, but bars don't touch each other (for discrete data)
Types of frequency distribution
Frequency polygon → connecting the dots - continuous
Cumulative frequency → each new data points is the sum of previous data points
Interpreting graphs
4 important characteristics:
Shape, centre, spread, outlines
Nomral data has avery characteristic shape, above describe how the shape varies from normal
Shape
First indicator of the distribution of date
If data is not symmetrical it is described as skewed
Positive skew (right) has long tail towards the positive direction of the graph
Negative skew (left) has long tail towards the negative direction of graph
Centre and spread
Centre
Equal amount of data on both side
The mean (average)
For symmetrical data, this will also be the median and mode
Spread
Measures how close individual data points are to the centre
Normal data is described as mesokurtic
Data that is widely spread from the centre is platykurtic
Data that is tightly grouped is leptokurtic
Outliers
Data that do not fit the expected pattern
Different ways to identify
Gap greater then 2 in a histogram
Off the line of best fit
Cannot just ignore these pieces of data