BIOSCI220 MODULE 1

0.0(0)
Studied by 5 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/158

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 4:03 AM on 9/18/25
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

159 Terms

1
New cards

Analysis of Variance (ANOVA)

The one-way analysis of variance is used to determine whether there are any differences between the means of three or more independent groups.

2
New cards

3 Assumptions of ANOVA

  • Each of the samples is a random sample from its population

  • The variable is normally distributed in each population

  • The standard deviation (and variance) of the variable is the same in all populations.


3
New cards

What are 4 common assumptions of parametric tests?

  • Data are random samples from the population

  • Data are normally distributed

  • Data are independent

  • If comparing groups, they have equal variances


4
New cards

What happens if parametric test assumptions are violated?

The conclusions of the research and interpretation of the results may change.

5
New cards

Average/mean

The sample mean is the sum of all the observations in a sample divided by n, the number of observations.

6
New cards

Blinding

The process of concealing information from participants (sometimes also researchers) about which experimental units receive which treatment.

7
New cards

Blocking

The grouping of experimental units that have similar properties. Within each block, treatments are randomly assigned to experimental units.

8
New cards

Categorical variable/data

Qualitative characteristics of individuals that do not have magnitude on a numerical scale (e.g. survival (dead or alive), size class (small, medium, large), life stage (egg, larva, juvenile, adult). Also called factor variables.

9
New cards

Centroid

The multivariate equivalent to an average/mean. The centroid is where the sum of the squared distances from each data point to the centroid is minimised in multivariate space.

10
New cards

Cluster analysis

Grouping data in such a way that observations in the same group (called a cluster) are more similar (in some sense) to each other than to those in other groups (clusters). K-means clustering minimises total within-cluster variation, based on k number of clusters, specified by the researcher.

11
New cards

Coefficient

Coefficients are numbers used to multiply a variable, it is located next to and in front of that variable. In regression with a single explanatory variable, the coefficient tells you how much the response variable is expected to increase (if the coefficient is positive) or decrease (if the coefficient is negative) when that explanatory variable increases by one.

12
New cards

Confidence Interval

A range of values around a population estimate that are believed to contain, with a certain probability (e.g. 95%), the true value of that statistic (i.e. the population value).

13
New cards

Confounding variable

An unmeasured explanatory variable that can mask or distort the causal relationship between measured variables in a study.

14
New cards

Control

A group or groups of subjects that do not receive an experimental treatment, but otherwise experience similar conditions as the treated subjects. Used to help determine if there are any treatment effects in chosen response variable(s).

15
New cards

Degrees of freedom (df)

The number of independent pieces of information that have gone into calculating your estimate.

16
New cards

Distribution

A distribution is a mathematical function that describes a collection of data, or scores, on a variable. For example, a sampling distribution is the probability distribution of all values for an estimate that we might obtain when we sample a population.

17
New cards

Dummy variable

Used with categorical factor variables in linear regression modelling to code levels/groups for estimating coefficients. Usually takes the value of 0 or 1.

18
New cards

Experimental unit

The independent physical entity which can be assigned, at random, to a treatment.

19
New cards

Factor

A type of explanatory variable used to explain/predict the response variable. Each factor will have two or more levels/groups. Both numeric and character variables can be made into factors, but a factor's levels will always be character values and there will be a limited/set number of these chosen by the scientist. Combination of factor levels are called treatments.

20
New cards

Frequency distribution

A frequency distribution describes the number of times each value of a variable occurs in a sample.

21
New cards

Hypothesis testing

Compares data to what we would expect to see if a specific null hypothesis were true. If the data are too unusual, compared to what we would expect to see if the null were true, we state we have evidence against the null hypothesis.

22
New cards

Null hypothesis

A specific statement about a population parameter. It is made for argument and often embodies the sceptical point of view. Often it is that the population parameter of interest is zero (no effect, no preference, no correlation, no difference).

23
New cards

Alternative hypothesis:

Includes all other feasible values for the population parameter besides the value stated in the null hypothesis.

24
New cards

Inference

Using probability to deal with uncertainty in drawing conclusions about a population parameter we wish to estimate or measure.

25
New cards

Interaction effect

An interaction effect happens when one explanatory variable has a relationship with another explanatory variable on a response variable. In other words, the explanatory variables do not act independently on the response variable.

26
New cards

Level/s

Levels are associated with Factor variables (see Factor). It is the splitting of the factor into two or more groups.

27
New cards

Linear regression

A method that draws a straight line through data to predict/explain the response variable (Y) from the explanatory variable (X). Assumptions are: the relationship between X and Y must be linear; at each value of X, the distribution of possible Y-values is normal (normally distributed residuals); the variance of Y-values is the same at all values of X (the residuals have constant variance); at each value of X, the Y-measurements represent a random sample from the population of possible Y-values.

28
New cards

Matrix (or data matrix)

Is easiest to explain by comparing to a data frame. Both are two-dimensional, where columns usually represent variables and rows represent samples. However, in a data frame the columns contain different types of data compared to a matrix where all elements are the same type of data.

29
New cards

Mean/average

The sample mean is the sum of all the observations in a sample divided by n, the number of observations.

30
New cards

Model

In simple terms, statistical modelling is a simplified, mathematically-formalized way to approximate reality (i.e. what generates your data) and optionally to make predictions from this approximation. The statistical model is the mathematical equation that is used.

31
New cards

n

The number of observations in a sample.

32
New cards

Non-metric multidimensional scaling (nMDS)

Is a commonly used multivariate data tool that works to collapse information from multiple dimensions into just a few dimensions to make it easy to visualise and interpret the information. Objects that are placed closer together are more similar to each other than objects found further apart. However, we need to check the stress value to ensure what is being viewed is interpretable (see “stress”).

33
New cards

Normal distribution

A continuous probability distribution describing a bell-shaped curve. It is a good approximation to the frequency distribution of many biological variables.

34
New cards

Numerical variable/data

Quantitative measurements that have magnitude on a numerical scale. These variables are numbers and can be discrete (e.g. number of amino acids in a protein) or continuous (e.g. height).

35
New cards

Null distribution

The sampling distribution of outcomes for a test statistic under the assumption that the null hypothesis is true.

36
New cards

Observation

An observation is the sampling unit from which we collect measurements and/counts from (related to the different variables of interest). Scientists can also use the terms subject, replicate, individual, unit.

37
New cards

One-sample t-test

Compares the mean of a random sample from a normal population with the population mean proposed in a null hypothesis. Assumptions are: the data are a random sample from the population; the variable is normally distributed in the population.

38
New cards

Outlier

An observation/data point well outside the range of values of other observations in a data set.

39
New cards

Parameter

A parameter is a quantity describing a population. A parameter never changes because it relates to the underlying population. We can use statistics to try and estimate parameters (Module 1 & 3 in BIOSCI 220), or we can use parameters to simulate or predict data (Module 2 in BIOSCI 220). Parameters are usually Greek letters (e.g. σ) or capital letters (e.g. P).

40
New cards

Parametric tests

Parametric statistical tests assume the underlying population from which you have sampled from are normally distributed. Non-parametric tests do not require this assumption to be met (they often use ranks of data points rather than the actual values of the data).

41
New cards

Population

A population is all the individual units of interest, whereas a sample is a subset of units taken from the population.

42
New cards

Principal Components Analysis (PCA)

The process of computing new uncorrelated variables, called principal components, from a dataset that explain the maximum amount of variation in the data. The idea of PCA is to reduce the number of variables of a data set, while preserving as much information as possible.

43
New cards

p-value

Each member of a population has an equal and independent chance of being selected into the sample being taken.

44
New cards

Random sample

Each member of a population has an equal and independent chance of being selected into the sample being taken.

45
New cards

Randomisation in experimental design

The random assignment of treatments to units in an experimental study.This process helps eliminate bias and ensures that treatment groups are comparable.

46
New cards

Randomisation test/permutation test

Generates a null distribution for the association between two variables by repeatedly and randomly rearranging the values of one of the two variables in the data. Assumptions are: the data are a random sample from the population; for tests that are comparing means or medians between groups, the distribution of the variable must have the same shape in every population (robust to this assumption when n is large).

47
New cards

Replication in experimental design

The application of every treatment to multiple, independent experimental units to allow us to generalise our results to our population of interest.This process helps to improve the reliability and validity of the results by reducing the effects of variability during the experiment.

48
New cards

Reproducibility

Research is considered to be reproducible when the exact results can be reproduced if given access to the original data, software, or code.

49
New cards

Residual/Error

Residuals (or errors) are the distances between data points and the model. They indicate the extent to which a model accounts for the variation in the observed data. Useful in linear regression and ANOVA models.

50
New cards

Response variable

Aka Dependent variable, outcome variable, Y variable. A response variable is what changes as a result of changes in one or more explanatory variables. It is the expected effect, and it responds to explanatory variables.

51
New cards

Sample

A sample is a subset of units taken from the population.

52
New cards

Scaling/Standardising

A technique used to allow comparison between variables that have different measurement units (e.g., grams and millimetres) or are on different scales. We centre the data and make the spread of the data equal by subtracting the variable mean then dividing by the variable standard deviation. This makes changes between and within variables relative to each other.

53
New cards

Screeplot

A diagnostic tool to check whether Principal Components Analysis works well on your data or not. A screeplot shows how much variation each PC captures from the data. An ideal curve should be steep, then bends at an “elbow” — this is your cutting-off point — and after that flattens out.

54
New cards

Skew

Refers to asymmetry in the shape of a frequency distribution for a numerical variable.

55
New cards

Standard deviation

A common measure of spread of a distribution. It indicates just how different measurements typically are from the mean.

56
New cards

Standard error

The standard deviation divided by the square root of n.

57
New cards

Statistical power

The probability that a random sample will lead to strong evidence against a null hypothesis, which is indeed false.

58
New cards

Stress

A term used in non-metric multidimensional scaling (nMDS) to describe how well the nMDS represents the original data. If stress is too high (>0.2), the nMDS should not be interpreted.

59
New cards

Test statistic

A one-number summary calculated from the data that is used to evaluate how compatible that data are with the result expected under the null hypothesis.

60
New cards

Two-sample t-test (independent)

Compares the means of two random and independent samples from two normal populations with the difference in population means proposed in a null hypothesis. Assumptions are: each of the two samples is a random sample from its population; the variable is normally distributed in each population; the standard deviation (and variance) of the variable is the same in both populations.

61
New cards

Type I error

Rejecting a true null hypothesis. The significance level α sets the probability of committing a type I error.It occurs when a test indicates a significant effect or difference when there is none.

62
New cards

Type II error

Failing to find evidence against a false null hypothesis.

63
New cards

Variable

A variable is any characteristic or measurement that differs from individual to individual. Data/observations are the measurements of one or more variables made on a sample of individuals. Explanatory variables predict or affect another variable, called a response variable.

64
New cards

Welch’s t-test

Compares the means of two random and independent samples from two normal populations with the difference in population means proposed in a null hypothesis. Differs from the two-sample independent t-test in that it can be used even when the variances of the samples are not equal.

65
New cards

Variance

A measure of spread of a distribution. It is the average of the squared differences from the mean. The standard deviation is the square root of the variance.

66
New cards

Variation

Variability in statistics refers to the difference being exhibited by data points within a data set, as related to each other or as related to the mean. This can be expressed through the range, variance or standard deviation of a data set. Variation is also a biology term and is used in a similar way, where it refers to the differences or deviations from the recognized norm or standard.

67
New cards

?function

If you have questions about a particular R function, or what arguments can be given to that function, you can access its documentation by writing the code ?insert_function_name_here into the console in R. It will return a help file in the bottom right-hand pane in RStudio. For operators like %>%, enclose in backticks like this: ?%>%``.

68
New cards

%>% (piping operator) – tidyverse


The pipe operator allows us to combine multiple operations in R into a single sequential chain of actions. %>% takes the output of one function and then “pipes” it to be the input of the next function. Read %>% as “then” or “and then.”

69
New cards

<- (assignment operator) – base

The assignment operator is used to assign a value to a variable.

70
New cards

+ – base

Sums two variables together.

71
New cards

- – base

Subtracts one variable from another.

72
New cards

* – base

Multiplies two variables together.

73
New cards

/ – base

Divides one variable from another.

74
New cards

$ – base

Allows you to access objects stored within an object (e.g. df$variable).

75
New cards

apply() – base

Apply a function across an array, matrix or data frame. MARGIN = 1 (rows), 2 (columns), or c(1, 2).

76
New cards

args() – base

Tells you what arguments a function will take.

77
New cards

data.frame – base

The most common way of storing data in R. Columns = variables, rows = samples.

78
New cards

drop_na() – tidyr

Drop rows containing missing values.

79
New cards

factor() – base

Convert variable(s) into factor variables.

80
New cards

filter() – dplyr

Subset a data frame, retaining or dropping rows based on conditions.

81
New cards

getwd() – base

Returns filepath of the current working directory.

82
New cards

ggplot() – ggplot2

Opening command to create a data visualisation.

83
New cards

glimpse() – dplyr

Transposed version of print: columns run down, data across.

84
New cards

group_by() – dplyr

Group dataframe by variables (columns).

85
New cards

install.packages("package") – base

Installs R package from repository.

86
New cards

is.na() – base

Returns TRUE for each data point that is NA.

87
New cards

library() – base

Loads installed R packages into current session.

88
New cards

list.files() – base

Returns all files in working directory.

89
New cards

mutate() – dplyr

Create/modify/delete columns in a data frame.

90
New cards

pivot_longer() – tidyr

“Lengthens” data (more rows, fewer columns).

91
New cards

pivot_wider() – tidyr

“Widens” data (more columns, fewer rows).

92
New cards

read_csv() – readr

Imports a .csv file into R as a tibble.

93
New cards

tibble – tidyverse

A special type of data frame (cleaner printing, easier to work with).

94
New cards

select() – dplyr

Subset a data frame by columns.

95
New cards

setwd() – base

Sets working directory (use at beginning of every R session).

96
New cards

summarise() / summarize() – dplyr

Reduces multiple values to a single value, often with group_by().

97
New cards

view() – dplyr

Opens dataset in spreadsheet style format in RStudio.

98
New cards
99
New cards
100
New cards