Statistics for Kinesiology

Why Stats

Statistics → collection of information (data) and the methods used to analyze them. A discipline. “The science of making an educated guess”


Having good data skills (gathering, analyzing, summarizing, displaying & interpreting) makes you an intelligent decision maker


Statistics are formal methods/approaches to:

  1. collect/summarize data

  2. Objectively analyze data

  3. Make unbiased interpretations leading to a decision

Key terms

  • Data - result of measurement

  • Evaluation - philosophical process of determining the worth of data

  • Reliability - consistency

  • Objective - no bias

  • Subjective - with bias

  • Variable - characteristic that can assume more than one value

  • Constant - characteristic that can only assume one value

  • Observational - research that involves passive collection of data or self-reported data

  • Experimental - research process that involves active intervention and control

  • Validity - Appropriateness of the test in measuring what it is designed to measure

    • Internal validity - a measure of the control within the experiment to be sure that the results are due to the intervention

      • Affected by confounding variables, instrument error, investigator error, bias

    • External validity - The ability to generalize the sample results to the population interest

      • Affected by appropriate inclusion and exclusion criteria

Population Vs Sample


  • Population

    • Collection of ALL subjects/things of interest that have more then 1 thing in common

  • Population parameter

    • Any measure/characteristic computed/determined for a population

  • Population size - Denoted as N

  • An observation of interest within a population - denoted as X

  • Sample

    • Subset of population of interest

  • Sample statistics

    • Any measure/characteristic computed/determined for a sample

  • Sample size - Denoted as n

  • Observation of interest within a sample - Denoted as x


Why not the whole population

  1. Numbers

  2. Time

  3. Manpower

  4. Cost

  5. Willingness of all members of the population to participate

  6. Destructive nature of experiment

Fundamental assumption of sampling

  • Sample is representative of the population (PARAMETRIC)

    • Otherwise we will apply non-parametric text

Random Sample

  • Random sample is scientifically drawn group that actually possesses the same characteristics as the population

  • Every person has an equal selection opportunity for you sample

  • Selection of one person is independent of the selection of another person

    • If repeatedly draw sub-samples from the population, the means of multiple sample will approximate the population mean

Sample error

  • E.g. suppose we survey the weight of all UOGH students, N = 5000, Avg = 134.21lbs

  • Measured N = 200 students, sample avg = 135

  • The difference between a population parameter and a sample statistic is known as Sampling error

When n and repeated sampling goes up, sampling error goes down

Data levels of measurement

  • Nominal → Attributes are named, weakest

  • Ordinal → attributes can be ordered

  • Interval → distance is meaningful

  • Ration → Absolute zero

Important Because data type defines which statistical test can be performed

Nominal

  • Allows to distinguish differences between items qualitatively 

  • There is no quantitative ordering or value, no hierarchy

  • Assign responses to to different categories = discrete data

    • Examples → city of birth, postal code, biological sex, favourite movies, id number

Ordinal

  • The categories have logical, meaningful order or rank

  • Starts at lowest ends at highest

  • BUT don’t know the numerical distance between categories, so still discrete data

  • Still cannot do mathematical calculation but can compare

    • Examples → likert scales, starred reviews, perceived exertion, letter Grades

Interval

  • Measurements are numerical values

  • Intervals of equal length represent equal differences in the characteristic. This is the key difference between interval and ordinal

  • Data is usually continuous (can hold any value including fractions/decimals

  • No true zero – zero is not the absence of what is being measured

  • Starting point is arbitrary

  • Data can be added or subtracted, but not multiplied or divided

    • Examples – standardized scores like IQ, temperature

Ratio

  • Allows for identification of absolute differences

  • absolute/true zero

    • Absence of a characteristics

  • Usually continuous

  • Can be added subtracted multiplied or divided

  • Most measured data fits into this category

    • Examples – time to completion, age, weight, distance

Organizing data

Frequency distribution

Rank order distribution

  • Simplest distribution

  • Only feasible with very small data sets

    • Ex) 13, 2, 9, 17, 5

    • Rank ordered = 2, 5, 9, 13, 17

    • Determine number of data points, N = 5

    • Can determined high = 17, low = 2

    • Can determine range (R) = H - L = 17 -2 = 15

Sample frequency distribution

  • For larger data sets

  • Tabulates the occurrence of each number

  • N = the sum of all frequencies

Ex) 5, 4, 9, 6, 4, 5, 2, 9, 9, 3, 9


X

F

2

1

3

1

4

2

5

2

6

1

9

4

N = 11

H = 9

L = 2

R = 7

Grouped frequency distribution

  • For very large data sets, particularly with very large ranges

  • Tabulates the occurrence of each number within a range

  • Lose some information - do not know exact values of number








Types of frequency distribution graphs

Histogram

  • Frequency, probability density (for continuous data)

Bar graph 

  • Frequency, but bars don't touch each other (for discrete data)

Types of frequency distribution

  • Frequency polygon → connecting the dots - continuous

  • Cumulative frequency → each new data points is the sum of previous data points

Interpreting graphs

4 important characteristics:

  • Shape, centre, spread, outlines

Nomral data has avery characteristic shape, above describe how the shape varies from normal

Shape

  • First indicator of the distribution of date

  • If data is not symmetrical it is described as skewed

  • Positive skew (right) has long tail towards the positive direction of the graph

  • Negative skew (left) has long tail towards the negative direction of graph

Centre and spread

Centre

  • Equal amount of data on both side

  • The mean (average)

  • For symmetrical data, this will also be the median and mode

Spread

  • Measures how close individual data points are to the centre

  • Normal data is described as mesokurtic

  • Data that is widely spread from the centre is platykurtic

  • Data that is tightly grouped is leptokurtic

Outliers

  • Data that do not fit the expected pattern

  • Different ways to identify

    • Gap greater then 2 in a histogram

    • Off the line of best fit

  • Cannot just ignore these pieces of data