Data Collection & Analysis
In statistics we can generally group data into two broad types: quantitative and qualitative.
Qualitative data is non-numerical data, such as species, colour or shape.
Quantitative data is numerical data, such as height, age or mass.
We can further divide quantitative data into discrete and continuous data.
Discrete data can only take particular values and there are clear gaps between its possible values.
Continuous data can take any value in an uninterrupted range.
A study that observes or measures all members of a population is called a census.
Rather than measure or observe every member of a population, we can instead measure or observe only some members of a population, which we call a sample. We use the term sampling units to refer to the individual members of a population. In order to collect a sample, we often make use of a sampling frame, which is a named or numbered list of all the sampling units within a population.
In order to make samples more representative of the population they are taken from and reduce bias, we make use of random sampling.
There are several different methods of random sampling:
Simple random sampling
Simple random sampling describes procedures where sampling units are chosen at random such that every sample of a particular size has the same probability of being selected. One way to collect a simple random sample is to assign a number to every sampling unit and then generate random numbers to select sampling units (discarding and replacing duplicates) until the sample reaches the required size.
Advantages
Free of bias
Easy and cheap to implement for small populations or samples
Each sampling unit has a known and equal chance of selection
Disadvantages
Requires a sampling frame
Not suitable for large samples or populations due to cost or lack of access to each selected sampling unit
Systematic sampling
A systematic sample is selected at regular intervals from an ordered list. One way to collect a systematic sample is to select a random starting point and then select each sampling unit after a regular interval. For a population of size N and a systematic sample of size n, the size of the interval will be k=nN.
For example, suppose we need a sample of size 15 from a population of 300. We select the first sampling unit randomly , for example by generating a random number between 1 and 15300=20, and then select every 20th sampling unit after that (eg. 3,23,43,63,...).
Advantages
Simple and quick to use
Suitable for large samples or populations
Disadvantages
Requires a sampling frame
Can introduce bias if the sample coincides with a pattern in the population
Stratified sampling
In stratified sampling, the population is divided into mutually exclusive strata and a random sample is taken from each. The proportion of the sample taken from each stratum should be equal. The strata are based on the characteristics of the population and every sampling unit belongs to exactly one stratum. This can be calculated as:
number to sample from stratum=population sizesample size×size of stratum
For example, suppose we wish to study pollinator preferences by selecting a sample of 30 tulips from a field of 150 tulips, which is comprised of 50 red tulips, 65 yellow tulips and 35 purple tulips. We would then randomly select 15030×50=10 red tulips, 15030×65=13 yellow tulips and 15030×35=7 red tulips.
Advantages
Accurately reflects the population's structure
Guarantees proportional representation of groups
Allows for comparison between different groups
Disadvantages
Requires a sampling frame
Population must be classified into distinct strata
Selection within the strata suffers from the same disadvantages as other random samples
Cluster sampling
In cluster sampling, the population is divided into similar clusters and then a simple random sample of the clusters is selected. All members of the selected clusters are included in the sample. The clusters are not based on the characteristics of the population (unlike strata), but every sampling unit still belongs to exactly one cluster.
For example, suppose we wished to select a sample of the residents of a particular street. Rather than select from the individual residents, we could group them into clusters by which house they live in. We would then randomly select houses and the residents of the selected houses would be included in the sample.
Advantages
Can be cheaper than other random samples
Suitable for large populations
Disadvantages
Data may not reflect the population as a whole if the clusters vary significantly from each other
It can be difficult to analyse and interpret the data
Opportunity sampling
In opportunity sampling, the sample is collected based on availability. This method is also called convenience sampling.
For example, a researcher could interview the first 20 people they meet outside a supermarket about their shopping habits. This method introduces bias as only people who visit that supermarket are selected for the sample.
Advantages
Easy to implement
Inexpensive
No sampling frame required
Disadvantages
Introduces bias
Unlikely to be representative
Highly dependent upon the individual researcher
Quota sampling
In quota sampling, the population is divided into mutually exclusive strata and a sample is collected until the required number for each stratum has been selected. Since each quota is fulfilled based on availability, an opportunity sample is collected for each quota.
For example, a researcher may wish to speak to 15 people aged 21−30, 20 people aged 31−40 and 10 people aged 41−50. They would interview people, allocate them to the appropriate age group and continue until they had the correct number for each age group. If anyone declined to be interviewed, or was in an age group for which the quota had already been fulfilled, the researcher would ignore them. Since people are allocated to each age group as they are interviewed, an opportunity sample is being collected for each age group.
Advantages
Allows a small sample to still be representative
No sampling frame required
Fast, simple and inexpensive to implement
Allows for comparison between different groups
Disadvantages
Non-random sampling can introduce bias
The population must be divided into groups, which might not accurately represent the sizes of the groups in the population
Increasing the scope of the study requires more groups, which increases cost
Non-responses are not recorded
Self-selected sampling
In self-selected sampling, sampling units volunteer to participate in the sample. They are often recruited through advertisements. This is sometimes used to recruit participants who have a rare condition or experience a rare phenomenon.
Advantages
Easy and cheap to implement
Can reach a large population
No sampling frame required
Disadvantages
Likely to introduce bias
Unlikely to be representative of the population
Difficult to generalise findings to the population
Histograms
A histogram is for displaying grouped continuous data whereas a bar chart is for discrete or qualitative data
On a histogram frequency density is plotted on the y – axis
class width x frequency density = frequency