Data Collection & Analysis


In statistics we can generally group data into two broad types: quantitative and qualitative.

  • Qualitative data is non-numerical data, such as species, colour or shape.

  • Quantitative data is numerical data, such as height, age or mass.


We can further divide quantitative data into discrete and continuous data.

  • Discrete data can only take particular values and there are clear gaps between its possible values.

  • Continuous data can take any value in an uninterrupted range.


A study that observes or measures all members of a population is called a census.

Rather than measure or observe every member of a population, we can instead measure or observe only some members of a population, which we call a sample. We use the term sampling units to refer to the individual members of a population. In order to collect a sample, we often make use of a sampling frame, which is a named or numbered list of all the sampling units within a population.

In order to make samples more representative of the population they are taken from and reduce bias, we make use of random sampling.

There are several different methods of random sampling:

Simple random sampling

Simple random sampling describes procedures where sampling units are chosen at random such that every sample of a particular size has the same probability of being selected. One way to collect a simple random sample is to assign a number to every sampling unit and then generate random numbers to select sampling units (discarding and replacing duplicates) until the sample reaches the required size.

Advantages

  • Free of bias

  • Easy and cheap to implement for small populations or samples

  • Each sampling unit has a known and equal chance of selection

Disadvantages

  • Requires a sampling frame

  • Not suitable for large samples or populations due to cost or lack of access to each selected sampling unit

Systematic sampling

A systematic sample is selected at regular intervals from an ordered list. One way to collect a systematic sample is to select a random starting point and then select each sampling unit after a regular interval. For a population of size N and a systematic sample of size n, the size of the interval will be k=nN.

For example, suppose we need a sample of size 15 from a population of 300. We select the first sampling unit randomly , for example by generating a random number between 1 and 15300​=20, and then select every 20th sampling unit after that (eg. 3,23,43,63,...).

Advantages

  • Simple and quick to use

  • Suitable for large samples or populations

Disadvantages

  • Requires a sampling frame

  • Can introduce bias if the sample coincides with a pattern in the population

Stratified sampling

In stratified sampling, the population is divided into mutually exclusive strata and a random sample is taken from each. The proportion of the sample taken from each stratum should be equal. The strata are based on the characteristics of the population and every sampling unit belongs to exactly one stratum. This can be calculated as:

number to sample from stratum=population sizesample size​×size of stratum

For example, suppose we wish to study pollinator preferences by selecting a sample of 30 tulips from a field of 150 tulips, which is comprised of 50 red tulips, 65 yellow tulips and 35 purple tulips. We would then randomly select 15030​×50=10 red tulips, 15030​×65=13 yellow tulips and 15030​×35=7 red tulips.

Advantages

  • Accurately reflects the population's structure

  • Guarantees proportional representation of groups

  • Allows for comparison between different groups

Disadvantages

  • Requires a sampling frame

  • Population must be classified into distinct strata

  • Selection within the strata suffers from the same disadvantages as other random samples

Cluster sampling

In cluster sampling, the population is divided into similar clusters and then a simple random sample of the clusters is selected. All members of the selected clusters are included in the sample. The clusters are not based on the characteristics of the population (unlike strata), but every sampling unit still belongs to exactly one cluster.

For example, suppose we wished to select a sample of the residents of a particular street. Rather than select from the individual residents, we could group them into clusters by which house they live in. We would then randomly select houses and the residents of the selected houses would be included in the sample.

Advantages

  • Can be cheaper than other random samples

  • Suitable for large populations

Disadvantages

  • Data may not reflect the population as a whole if the clusters vary significantly from each other

  • It can be difficult to analyse and interpret the data


Opportunity sampling

In opportunity sampling, the sample is collected based on availability. This method is also called convenience sampling.

For example, a researcher could interview the first 20 people they meet outside a supermarket about their shopping habits. This method introduces bias as only people who visit that supermarket are selected for the sample.

Advantages

  • Easy to implement

  • Inexpensive

  • No sampling frame required

Disadvantages

  • Introduces bias

  • Unlikely to be representative

  • Highly dependent upon the individual researcher

Quota sampling

In quota sampling, the population is divided into mutually exclusive strata and a sample is collected until the required number for each stratum has been selected. Since each quota is fulfilled based on availability, an opportunity sample is collected for each quota.

For example, a researcher may wish to speak to 15 people aged 21−30, 20 people aged 31−40 and 10 people aged 41−50. They would interview people, allocate them to the appropriate age group and continue until they had the correct number for each age group. If anyone declined to be interviewed, or was in an age group for which the quota had already been fulfilled, the researcher would ignore them. Since people are allocated to each age group as they are interviewed, an opportunity sample is being collected for each age group.

Advantages

  • Allows a small sample to still be representative

  • No sampling frame required

  • Fast, simple and inexpensive to implement

  • Allows for comparison between different groups

Disadvantages

  • Non-random sampling can introduce bias

  • The population must be divided into groups, which might not accurately represent the sizes of the groups in the population

  • Increasing the scope of the study requires more groups, which increases cost

  • Non-responses are not recorded

Self-selected sampling

In self-selected sampling, sampling units volunteer to participate in the sample. They are often recruited through advertisements. This is sometimes used to recruit participants who have a rare condition or experience a rare phenomenon.

Advantages

  • Easy and cheap to implement

  • Can reach a large population

  • No sampling frame required

Disadvantages

  • Likely to introduce bias

  • Unlikely to be representative of the population

  • Difficult to generalise findings to the population







Histograms

  • A histogram is for displaying grouped continuous data whereas a bar chart is for discrete or qualitative data

  • On a histogram frequency density is plotted on the y – axis


class width x frequency density = frequency