Statistics

  1. Chapter 1: Introduction to Statistics

    1. Definition: statistics is the science of planning studies and experiments and collecting data, organizing, summarizing, presenting, analyzing and interpreting data, and then drawing conclusions based on them.

    2. Definition: Data are observations that have been collected.

      1. types of data:

        1. Quantitative: a number (age,weight,cost)

        2. Qualitative: a category/words (gender, car type, color of your shoes)

    3. Definition: a population is a complete collection of all elements under study

      1. Notation: N= total population size

    4. Definition: a parameter is a numerical measurement calculated from the population data

    5. Definition: a sample is a subcollection of our population

    6. Definition: a statistic is a numerical measurement calculated from out sample data

      1. types of quantitative data:

        1. discrete: a number that dosen’t involve a decimal or fraction

          1. ex. the number of…..

        2. continues: a number that can involve a decimal/ fraction

          1.  ex. the amount of or how much

    7. Data can be:

      1. descriptive: presented in words pictures, or graphs

      2. inferential: presented using standard numeric symbols

    8. levels of measurement for data

      1. nominal: categories only. no inherent order (qualitative)

        1. ex. favorite color

      2. ordinal: leveled/order data. order exists (qualitative)

        1. ex. low risk or high risk

      3. interval: ordered with no true zero.

        1. ex. temperature, time

      4. ratio: ordered with true zero

        1. ex. weight, amount of gas in your car

    9. data collection method:

      1. two methods of collecting data:

        1. observational studies: observe facts and draw conclusions

        2. experiment: apply a treatment in a controlled environment to determine a response/outcome to the treatment

    10. methods of sampling: sampling methods are different ways to collect a sample from a population

      1. simple random sampling: a sample of N is chosen so that every element in the population has an equal chance of being selected in the sample

        1. ex. picking names from a hat, picking cards from a box

      2. systematic sampling: select a starting point then pick every k element in the population

        1. ex. choosing every 10th customer that enters a store, selecting every 23rd order made for a given product

      3. convenance sampling: select elements from a population which are easiest to get

        1. ex. asking people in the math learning center if they love math and making a conclusion about how many people love math at MNSU

      4. stratified sampling: breaking population up into meaningful subgroups called strata, then drawing a simple random sample from each stratum

        1. ex………………….

      5. cluster sampling: dividing population into diverse clusters and randomly selecting some clusters (include all elements of those clusters)

        1. ex. a university randomly selects 5 dorms and surveys every student living in the dorm

  2. chapter 2: descriptive statistics

    1. frequency tables: organize raw data into more manageable groups

      1. allow you to organize raw data into classes i (intervals) and ii counts (frequency’s)

        1. class = intervals

        2. counts = frequency

    2. How to construct a frequency table:

      1. Determine the number of classes ( c = In(n) ) round up

      2. Calculate the range (range = maximum - minimum)

      3. Determine the class width (width = range/ number of classes)

      4. determine the classes by introducing an extra decimal position

      5. make your table

    3. class types:

      1. lower class limits: lower end of each class (classes)

      2. upper class limits: upper end of each class (counts)

      3. class boundaries: numbers between each class

      4. class midpoints: the number in the middle of each class

      5. class width: difference between lower- and upper-class limit

    4. cumulative frequency: the sum of the frequencies for a class and all previous classes

    5. graphical displays:

      1. stem-leaf plot: a way to summarize data when n (sample size) is small

      2. bar graph: categorical data displayed by area

      3. pie charts: categorical data displayed by area (we hate theses!)

      4. histograms: display variables with ordinal measurements (type of bar graph)

      5. relative frequency histograms: use relative frequencies (type of bar graph)

        1. histogram shapes: uniform, normal (hill shaped), right-skewed, left-skewed

      6. box plot: ……………………………

      7. times series graph: we don’t care about this one ether!

    6. Numerical measures:

      1. measures of central tendency (i.e. mean, medium, etc..)

        1. mean- ……………………...

        2. median- how to find the median

          1. sort data from smallest to largest

          2. find the middle value:

            1. if n is odd, the median is in the exact middle

            2. is n is even, the median is the average of the two middle values

        3. mode- data that appears the most frequently in a data set

          1. if two values appear with the same maximum frequency, it is bimodal

          2. if more than two values appear with the same maximum frequency, it is multimodal

          3. if no values are repeated, it has no mode

        4. midrange- minimum value + maximum value over 2

      2. measures of variation/spread (i.e. variance, mean, absolute, error, etc..)

      3. measure of relative position (i.e. z-score, empirical rule)

    7. properties of measures of central tendency

      1. mean: highly affected by outliers since all data values used for calculation

      2. median: less affected by outliers since we just find the middle value

      3. mode: applicable to categorical data

      4. midrange: highly affected by outliers because we use two extreme values (min+max)

    8. relation between mean, median, and mode

      1. normal: mean = median = mode

      2. right-skewed: mode< median< mean (positive)

      3. left-skewed: mode> median> mean > mode (negative)

    9. measures of variation

      1. range = maximum value - minimum value

      2. mean absolute divination:

        1. population MAD: