Sampling and Sampling Distribution

Sampling and External Validity

External validity

is the extent to which results can be generalized

  • can be across individuals, or across methods and settings

Population

  • the accessible population is a smaller part of the sample

Sampling

  • can have an impact on how valid the data is, is it representative, can it be generalized, how and why?

  • You want to get a result that can be inferred to the population from the smaller number of respondents

  • the larger the sampling size, the more accurate the population estimate

    • the sample does not need to be a certain percentage of the population

      • there are mathematical formulas to determine adequate sample size (e.g. power analysis)

Sampling Error, Sampling Bias, and Sampling Representativeness

sampling error

is the difference between the statistic used to estimate the population parameter, and the parameter’s actual value

  • larger samples are more likely to yield data that accurately reflects the true population value

Sampling Bias

  • occurs when some members of a population have a greater chance of being selected than other members of the population

    • can dramatically increase sample error in a systematic way

    • causes misleading results by over-representing one groups; opinion and under-representing another group’s opinion

  • Haphazard Sampling

    • There are more of one characteristic, e.g. more higher income, more people from a demographic, etc. within the population, so the samples end sup being much more of one demographic.

Types of Probability Sampling

  • is actually pretty rare in HDFS research

  • population info is available BEFORE participants are picked

    • Simple Random sampling

      • Done with a random table or a computer randomization (e.g. using excel)

    • systematic random

      • picking up every nth subject (every 4th and 9th)

      • sensitive to the way the list is order

    • stratified random

      • Guarantees that the proportion of the population is the same as the sample

        • divide the population into stratum (can make it tow here it matches the

          • strata should be homogenous within and heterogenous between one another

            • e.g. dividing into

    • cluster

      • sample a ready-made group within the population, assuming it has a similar composition to the populations

      • clusters must be heterogeneous within (has a diff exact variation within it) and homogenous between one another (all have 3 blue marbles, and 2 green marbles)

Types of Non-Probability Sampling

  • population info is not available

    • chance of people included is unknown and unequal

  • convenience

    • get any available people in the population (

      • can have low representativeness/generalizability (i.e., external validity)

purposive (judgemental)

  • Obtain subjects who are believed to meet a predetermined criterion

    • e.g. only young adults entering a theatre to watch popular movies

  • Also has low representativeness and generalizability

  • Quality of sampling depends on researchers ability to make judgements on the qualifications of participants

  • quota

    • Predetermine the proportion or number of groups in the sample (50% male, 50% female or 30 males and 30 females)

      • used when you want to ensure the sample reflects the proportion of the group in the populations (e.g. TXST is 11% AA, 70% White/European American, etc.)

      • Also used to secure enough number of group members for analysis (e.g. if the amt of children w/ gay parents is very small, or amount of fraternal twins) (can sometimes be called over-sampling to make sure you get enough)

        • in this case, wouldn’t be representative of the whole population, since you purposely picking out and targeting one demographic

  • snowball

    • used to locate hidden, hard to reach populations

      • drug addicted individuals, married straight individuals that actually consider themselves gay, gang-affiliated members, sex-workers, etc.

    • You obtain your subjects through a chain of personal networks

Inferential Statistics

  • take data from a sample and make inferences about the population parameters

    • are the statistics of the population parameters (e.g. mean, median, ANOVA score, etc.

    • We can repeat the study several times in order to determine this, and represent it like this:

This is a sampling distribution (a combo of several samples, not just one sample, which is just a sample distribution)

margin of error is significantly smaller as the sample grows larger

  • as we take

    • we give them a significance levels, which give them a margin of error. we then put that value (rep. by alpha) into this equation to make it a confidence level:

      • (1-a)*100)


Central Limit Theorem

  • if the population distribution is normal, the sampling distribution of mean will also be normal

    • If its not normal, the sample will still become more and more normal as the sample size increases

When a sample has more than 30 participants, the more normal it will be. This is better, and is represented well in the sampling distribution

SD of the sampling distrib, is called the standard error or standard sampling error

We can calculate where a value is on the normal curve based on the z-scores

for 95%, your z-score must be plus or minus 1.96

Class Notes

  • Variance is the exact sum of the squared values of the deviations of all other values

    • it is the arithmetic average of the deviation

    • Standard deviation is a standardized value of how much is left over when the avg deviation is taken away

  • Coefficient of variation (CV) = SD/M is sometimes useful to make more use of variance

    • when the value is less than 10% low relative variability; when it is greater than 30% is high relative variability

      • however, this does not always work

Visualization Lab

  • each band in a histogram is usually 1.

    • however, you can customize it through the chart editor

Box Plot Commands

  • analyze

    • explore

      • Add desired variables into the Dependent List

        • Click plots (make sure nothing but factor levels together in box plots is clicked, everything else should not be clicked)

  • IQR tells us the range of where people are clustered the most

    • It is a better measure of range than the actual range, actually, a lot of times

  • The values you see outside of the minimum shows the real minimum

    • the line for minimum shows us the minimum excluding those extreme outliers

  • “5% Trimmed mean”

    • this shows us the mean with the (top and bottom 5% excluded to try to show what it is without the outliers)