Statistics
Chapter 1: Introduction to Statistics
Definition: statistics is the science of planning studies and experiments and collecting data, organizing, summarizing, presenting, analyzing and interpreting data, and then drawing conclusions based on them.
Definition: Data are observations that have been collected.
types of data:
Quantitative: a number (age,weight,cost)
Qualitative: a category/words (gender, car type, color of your shoes)
Definition: a population is a complete collection of all elements under study
Notation: N= total population size
Definition: a parameter is a numerical measurement calculated from the population data
Definition: a sample is a subcollection of our population
Definition: a statistic is a numerical measurement calculated from out sample data
types of quantitative data:
discrete: a number that dosen’t involve a decimal or fraction
ex. the number of…..
continues: a number that can involve a decimal/ fraction
ex. the amount of or how much
Data can be:
descriptive: presented in words pictures, or graphs
inferential: presented using standard numeric symbols
levels of measurement for data
nominal: categories only. no inherent order (qualitative)
ex. favorite color
ordinal: leveled/order data. order exists (qualitative)
ex. low risk or high risk
interval: ordered with no true zero.
ex. temperature, time
ratio: ordered with true zero
ex. weight, amount of gas in your car
data collection method:
two methods of collecting data:
observational studies: observe facts and draw conclusions
experiment: apply a treatment in a controlled environment to determine a response/outcome to the treatment
methods of sampling: sampling methods are different ways to collect a sample from a population
simple random sampling: a sample of N is chosen so that every element in the population has an equal chance of being selected in the sample
ex. picking names from a hat, picking cards from a box
systematic sampling: select a starting point then pick every k element in the population
ex. choosing every 10th customer that enters a store, selecting every 23rd order made for a given product
convenance sampling: select elements from a population which are easiest to get
ex. asking people in the math learning center if they love math and making a conclusion about how many people love math at MNSU
stratified sampling: breaking population up into meaningful subgroups called strata, then drawing a simple random sample from each stratum
ex………………….
cluster sampling: dividing population into diverse clusters and randomly selecting some clusters (include all elements of those clusters)
ex. a university randomly selects 5 dorms and surveys every student living in the dorm
chapter 2: descriptive statistics
frequency tables: organize raw data into more manageable groups
allow you to organize raw data into classes i (intervals) and ii counts (frequency’s)
class = intervals
counts = frequency
How to construct a frequency table:
Determine the number of classes ( c = In(n) ) round up
Calculate the range (range = maximum - minimum)
Determine the class width (width = range/ number of classes)
determine the classes by introducing an extra decimal position
make your table
class types:
lower class limits: lower end of each class (classes)
upper class limits: upper end of each class (counts)
class boundaries: numbers between each class
class midpoints: the number in the middle of each class
class width: difference between lower- and upper-class limit
cumulative frequency: the sum of the frequencies for a class and all previous classes
graphical displays:
stem-leaf plot: a way to summarize data when n (sample size) is small
bar graph: categorical data displayed by area
pie charts: categorical data displayed by area (we hate theses!)
histograms: display variables with ordinal measurements (type of bar graph)
relative frequency histograms: use relative frequencies (type of bar graph)
histogram shapes: uniform, normal (hill shaped), right-skewed, left-skewed
box plot: ……………………………
times series graph: we don’t care about this one ether!
Numerical measures:
measures of central tendency (i.e. mean, medium, etc..)
mean- ……………………...
median- how to find the median
sort data from smallest to largest
find the middle value:
if n is odd, the median is in the exact middle
is n is even, the median is the average of the two middle values
mode- data that appears the most frequently in a data set
if two values appear with the same maximum frequency, it is bimodal
if more than two values appear with the same maximum frequency, it is multimodal
if no values are repeated, it has no mode
midrange- minimum value + maximum value over 2
measures of variation/spread (i.e. variance, mean, absolute, error, etc..)
measure of relative position (i.e. z-score, empirical rule)
properties of measures of central tendency
mean: highly affected by outliers since all data values used for calculation
median: less affected by outliers since we just find the middle value
mode: applicable to categorical data
midrange: highly affected by outliers because we use two extreme values (min+max)
relation between mean, median, and mode
normal: mean = median = mode
right-skewed: mode< median< mean (positive)
left-skewed: mode> median> mean > mode (negative)
measures of variation
range = maximum value - minimum value
mean absolute divination:
population MAD: