Chapter 2 Notes – Hypothesis, Theory, Variables, Error, Sampling, and Data Tools

Hypothesis and Theory in Science

  • Hypothesis: a tentative or possible answer to a research question; more than a simple educated guess because it is based on prior research and knowledge.
  • Phenomenon observation: identify something in nature you want to study; propose a possible explanation (the hypothesis).
  • Good hypothesis criterion: testable. If it cannot be tested via an experiment, survey, or investigation, it’s not a good hypothesis.
  • If testing yields results in agreement with the hypothesis, the work is published in a scientific journal and, if repeatable by others, the idea becomes a theory.
  • Theory definition: the best explanation we have for a natural process, based on data, experimentation, observations, and consensus among scientists. It is not merely a guess or a single opinion.
  • Examples of theories: cell theory, atomic theory, theory of evolution, Big Bang theory. The Big Bang theory can be updated over time as new evidence emerges.
  • Important nuance: theories can be revised with new data; scientific knowledge evolves.

The Scientific Method: Testing, Publishing, and Theory Acceptance

  • After forming a hypothesis, you test it through experiments, investigations, or surveys.
  • If results support the hypothesis, publish in a scientific journal.
  • Reproducibility is critical: other researchers should be able to repeat the method and obtain similar results.
  • When replication is successful, the hypothesis gains status as part of a broader theoretical framework.

Variables in Experiments

  • Dependent variable: what you measure or observe (the outcome).
  • Independent variable: what you deliberately change (the cause or input).
  • Example from a plant lab: dependent variable could be the change in pH; independent variable is the type of water or fertilizer used.
  • Controlled variables: all other conditions kept the same (e.g., sunlight, water amount, temperature, pot size).
  • Control group: a baseline group that is not given the experimental treatment or is left untreated; used for comparison with experimental groups.
  • Experimental design goal: isolate the effect of the independent variable on the dependent variable while minimizing confounding factors.

Data Quality: Error and Bias in Data Collection

  • Systematic error: an error that consistently skews measurements in the same direction due to a faulty instrument or procedure (e.g., a miscalibrated pH meter).
  • Random error: fluctuations caused by unpredictable, uncontrollable factors or human variability; affects precision.
  • Bias: systematic favoritism that leads to distorted data collection or interpretation; can be intentional (e.g., funding bias or disinformation) or unintentional (design or technique flaws).
  • Examples of bias in science: funding bias (e.g., pharmaceutical sponsors), selective reporting, selective data inclusion, or data “cherry-picking.”
  • Objective data collection aims to minimize both systematic and random error and to avoid bias, ensuring data meaningfully informs decisions.

Distinguishing Random vs Systematic Data Collection Errors

  • When collecting data, some investigations require random sampling to avoid bias and representative results.
  • Random sampling example: selecting random locations in a study area to sample organisms, ensuring each unit has an equal chance of selection.
  • Systematic sampling example: selecting samples at regular intervals (e.g., every 1 meter along a transect) to study how a variable changes across space.
  • The same words (random vs systematic) apply differently depending on context (data collection vs data analysis); keep straight the distinction between how data are collected and how data are interpreted.

Quadrat Sampling: Random Sampling Approach

  • Quadrat: a square frame (often 1 m^2) used to sample stationary or slow-moving organisms (e.g., plants).
  • Random sampling steps (example workflow):
    • Define the study area and divide it into a grid of quadrats (each 1 m^2 in the common example).
    • Assign coordinates to each square (e.g., top-right corner coordinates).
    • Use a random number generator to select which quadrats to sample (random sampling to remove bias).
    • Place the quadrat at each selected coordinate and count the target organisms within it.
    • Record results and compute the mean count per quadrat.
    • Estimate total population by scaling up from sampled area to entire study area.
  • Edge rule: when an organism lies on the boundary, count it if at least half of the organism is inside the quadrat; if less than half is inside, do not count.
  • Example workflow in the transcript (quadrats for daisies):
    • Area: 620 m^2 field; use a 1 m^2 quadrat grid.
    • Sampling goal: sample ~10% of the total area (e.g., 30 squares; here 10 would be used).
    • For each sampled square, count daisies within the square; determine if edge plants count by the 50% rule.
    • Compute the mean number of daisies per square (per m^2) across all sampled squares.
    • Estimate total population: multiply the average density by the total area:
      • The example yields: average density = 7 daisies per m^2; total area = 620 m^2; estimated population = xˉimesA=7imes620=4340.\bar{x} imes A = 7 imes 620 = 4340.
  • Quadrat sampling advantages: simple, inexpensive, good for non-moving organisms, and flexible for different habitats.
  • Limitations: may require many quadrats for precision; edge effects; assumes random distribution or representative sampling; potential observer bias in counting.

Systematic Sampling with Transects

  • Systematic sampling collects data along a defined path (transect) using a regular interval.
  • Common setup: lay a tape measure along the study area (e.g., 10 meters) and sample at fixed distances (e.g., every 1 meter).
  • Procedure with quadrats along transect:
    • Place a quadrat at the start of the transect and count organisms within.
    • Move the quadrat 1 meter along the transect and count again; repeat along the entire transect.
  • Data representation: line graphs or kite graphs can show how population density changes with distance along the transect.
  • Example interpretation: density of daisies increases or decreases with distance from a gate; a biotic factor (e.g., trampling by people) likely affects abundance near high-traffic areas.
  • Transect sampling helps detect spatial trends and the influence of environmental gradients or human activity on populations.

Comparison: Random Quadrat Sampling vs Systematic Transect Sampling

  • Random quadrat sampling provides an unbiased estimate of mean density across the study area, good for heterogenous habitats.
  • Systematic transect sampling helps identify spatial patterns and gradients (e.g., edge effects, pollution gradients).
  • Both methods are valuable; the choice depends on the research question, habitat, and desired information (density vs distribution pattern).

Low-Cost, Low-Tech Sampling Tools and Methods

  • Quadrat-based sampling (already covered).
  • Pitfall trap: a container flush with the ground; insects fall in and cannot escape; used to sample ground-dwelling arthropods.
  • Sweeping nets: large nets for catching flying insects or those in vegetation.
  • Beating tray: a tray placed beneath a branch; gently beat the branch to dislodge insects onto the tray for counting.
  • Beating tray example: estimate insect populations by observing dropped insects; count and scale up.
  • Kick sampling: in aquatic environments; disturb sediment to flush organisms into a net downstream; used for aquatic invertebrates and small fish.
  • Light traps: a sheet or blanket placed with a light source at night to attract nocturnal insects for counting.
  • Quadrat vs capture/recapture methods:
    • Quadrat sampling is best for stationary organisms (plants, corals, barnacles).
    • Capture-recapture (mark-recapture) is used for mobile animals (e.g., Florida panther), involving capturing, tagging, releasing, recapturing, and counting.
  • Capture-recapture method (Mark-Recapture) basics:
    • First capture: n1 individuals captured and marked.
    • Second capture: n2 individuals captured; among these, m2 are marked from the first capture.
    • Population estimate (assuming a closed population and marks retained):
      • N^=n<em>1n</em>2m2\, \hat{N} = \frac{n<em>1 \cdot n</em>2}{m_2}
  • Assumptions and limitations of capture-recapture:
    • Marks are not lost or overlooked; marked individuals mix back into the population; no immigration or emigration during the study; closed population during sampling period.
    • Marks affecting recapture rates, or animals learning to avoid capture, can bias results.

Population Density and Population Size: Key Formulas

  • Random quadrat estimate (density-based extrapolation):
    • If each quadrat area is a=1 m2a = 1 \text{ m}^2 and total study area is AA square meters, the estimated population size is:
    • N^=xˉA\hat{N} = \bar{x} \cdot A, where \bar{x} is the mean number of individuals per quadrat.
  • Example: density per m^2 = 7; total area = 620 m^2;
    • N^=7×620=4340\hat{N} = 7 \times 620 = 4340.
  • Capture-recapture notation and formula:
    • First capture: n<em>1n<em>1, second capture: n</em>2n</em>2, recaptured marked individuals: m2m_2
    • Estimated population: N^=n<em>1n</em>2m2\hat{N} = \dfrac{n<em>1 \cdot n</em>2}{m_2}

Biodiversity and Diversity Indices

  • Simpson's Index (for comparing biodiversity between areas):
    • Let n<em>in<em>i be the number of individuals in species i, and N=</em>iniN = \sum</em>i n_i be the total individuals across all species.
    • The standard formula for the probability that two randomly selected individuals belong to the same species is:
    • D=<em>in</em>i(ni1)N(N1)D = \dfrac{\sum<em>i n</em>i(n_i - 1)}{N(N - 1)}
    • The diversity measure can be reported as 1 − D (often called Simpson's Diversity Index) to reflect higher values for more diverse communities.
  • Note: Simpson’s index is not a direct estimate of population size; it measures biodiversity and species evenness across areas.

High-Tech Data Collection and Analysis: Geospatial and Remote Sensing

  • Geospatial systems (GIS): computer-based maps and analyses; data input by specialists and analysts; output can be static (maps) or interactive (GIS dashboards).
    • Examples: maps used to assess drought conditions, wind farm siting, or crime distributions.
    • Privacy and ethics: interactive GIS can reveal sensitive information (e.g., residential locations of offenders); access and privacy controls are important.
  • Satellite data: numerous parameters collected from space (e.g., ozone, hurricanes, phytoplankton productivity, methane, atmospheric pollution, vegetation).
  • Tracking and wildlife monitoring with satellites and radio collars: e.g., Florida panther monitoring to study range, reproduction, and mortality; tagging and re-release to study movement and population dynamics.
  • Computer models: simulate weather and climate; two main timescales:
    • Weather models: days to weeks
    • Climate models: decades to centuries
  • Key limitation of computer models: predictive reliability decreases the further into the future you forecast due to increasing variables and uncertainties.
  • Hurricane forecasting example:
    • Cone forecasts illustrate uncertainty grows with time; tracks become broader as projected horizon increases.
    • The European (ECMWF) model is often more accurate for longer-range forecasts than some American models, though all models have strengths and weaknesses.
  • Crowdsourcing and big data:
    • Crowdsourced data sources include traffic data from smartphones (e.g., Google Maps using Bluetooth/Wi-Fi signals to infer crowd density) and social media.
    • Crowdsourcing can greatly increase data volume but requires substantial processing power and sophisticated analytics.
    • Privacy concerns and data ethics are central: data collection can reveal location, behavior, and other sensitive information.
  • Big data: benefits include rich information and patterns that may not be visible in small datasets; limitations include processing time, storage, and potential for spurious correlations if not properly analyzed.

Practical Takeaways for Exams and Lab Work

  • Distinguish clearly between hypothesis, theory, and laws (the transcript emphasizes that a theory is well-supported and not just a guess).
  • Be able to identify independent vs dependent variables and control variables in a described experiment.
  • Understand when to use random sampling (to avoid bias) vs systematic sampling (to examine spatial patterns or gradients).
  • Know how to implement a quadrat-based population estimate, including edge rules and the steps to scale from a sample to an entire area.
  • Be able to describe the capture-recapture method, its formula, and the assumptions that underlie it.
  • Recognize different sampling tools (quadrats, pitfall traps, beating trays, kick sampling, light traps) and what they estimate.
  • Understand the basics of biodiversity indices (e.g., Simpson’s index) and what they measure.
  • Appreciate the role of GIS, satellites, computer models, and crowdsourcing in modern data collection and how these methods complement traditional field techniques.
  • Be mindful of ethical considerations, privacy concerns, and avoiding bias in data collection and interpretation.

Quick Reference Equations

  • Quadrat density-based estimate:
    • N^=xˉA\hat{N} = \bar{x} \cdot A
    • where xˉ\bar{x} is mean count per quadrat and AA is total study area in square units (e.g., m^2).
  • Edge counting rule for quadrats: count an organism if at least one-half of its body is inside the quadrat; otherwise do not count.
  • Capture-recapture (Lincoln-Petersen) estimator:
    • N^=n<em>1n</em>2m2\hat{N} = \dfrac{n<em>1 \cdot n</em>2}{m_2}
    • Definitions: n<em>1n<em>1 = number captured and marked in first sample; n</em>2n</em>2 = number captured in second sample; m2m_2 = number of marked individuals recaptured in second sample.
  • Simpson’s Diversity Index (probability-based):
    • D=<em>in</em>i(ni1)N(N1)D = \dfrac{\sum<em>i n</em>i(n_i - 1)}{N(N - 1)}
    • Biodiversity measure commonly presented as 1D1 - D for intuitive interpretation (higher = more diverse).

Closing Notes

  • Homework vs studying: homework is a task to be completed, often in-class or at home; studying involves independent review, questions, and practice to prepare for exams.
  • The PowerPoints and videos are available for revision; use a mix of reading, watching, and practice problems to solidify understanding.
  • Expect to encounter the above equations and concepts on exams; the lecturer will provide the exact definitions and variable meanings on tests.