Chapter 2 Notes – Hypothesis, Theory, Variables, Error, Sampling, and Data Tools
Hypothesis and Theory in Science
- Hypothesis: a tentative or possible answer to a research question; more than a simple educated guess because it is based on prior research and knowledge.
- Phenomenon observation: identify something in nature you want to study; propose a possible explanation (the hypothesis).
- Good hypothesis criterion: testable. If it cannot be tested via an experiment, survey, or investigation, it’s not a good hypothesis.
- If testing yields results in agreement with the hypothesis, the work is published in a scientific journal and, if repeatable by others, the idea becomes a theory.
- Theory definition: the best explanation we have for a natural process, based on data, experimentation, observations, and consensus among scientists. It is not merely a guess or a single opinion.
- Examples of theories: cell theory, atomic theory, theory of evolution, Big Bang theory. The Big Bang theory can be updated over time as new evidence emerges.
- Important nuance: theories can be revised with new data; scientific knowledge evolves.
The Scientific Method: Testing, Publishing, and Theory Acceptance
- After forming a hypothesis, you test it through experiments, investigations, or surveys.
- If results support the hypothesis, publish in a scientific journal.
- Reproducibility is critical: other researchers should be able to repeat the method and obtain similar results.
- When replication is successful, the hypothesis gains status as part of a broader theoretical framework.
Variables in Experiments
- Dependent variable: what you measure or observe (the outcome).
- Independent variable: what you deliberately change (the cause or input).
- Example from a plant lab: dependent variable could be the change in pH; independent variable is the type of water or fertilizer used.
- Controlled variables: all other conditions kept the same (e.g., sunlight, water amount, temperature, pot size).
- Control group: a baseline group that is not given the experimental treatment or is left untreated; used for comparison with experimental groups.
- Experimental design goal: isolate the effect of the independent variable on the dependent variable while minimizing confounding factors.
Data Quality: Error and Bias in Data Collection
- Systematic error: an error that consistently skews measurements in the same direction due to a faulty instrument or procedure (e.g., a miscalibrated pH meter).
- Random error: fluctuations caused by unpredictable, uncontrollable factors or human variability; affects precision.
- Bias: systematic favoritism that leads to distorted data collection or interpretation; can be intentional (e.g., funding bias or disinformation) or unintentional (design or technique flaws).
- Examples of bias in science: funding bias (e.g., pharmaceutical sponsors), selective reporting, selective data inclusion, or data “cherry-picking.”
- Objective data collection aims to minimize both systematic and random error and to avoid bias, ensuring data meaningfully informs decisions.
Distinguishing Random vs Systematic Data Collection Errors
- When collecting data, some investigations require random sampling to avoid bias and representative results.
- Random sampling example: selecting random locations in a study area to sample organisms, ensuring each unit has an equal chance of selection.
- Systematic sampling example: selecting samples at regular intervals (e.g., every 1 meter along a transect) to study how a variable changes across space.
- The same words (random vs systematic) apply differently depending on context (data collection vs data analysis); keep straight the distinction between how data are collected and how data are interpreted.
Quadrat Sampling: Random Sampling Approach
- Quadrat: a square frame (often 1 m^2) used to sample stationary or slow-moving organisms (e.g., plants).
- Random sampling steps (example workflow):
- Define the study area and divide it into a grid of quadrats (each 1 m^2 in the common example).
- Assign coordinates to each square (e.g., top-right corner coordinates).
- Use a random number generator to select which quadrats to sample (random sampling to remove bias).
- Place the quadrat at each selected coordinate and count the target organisms within it.
- Record results and compute the mean count per quadrat.
- Estimate total population by scaling up from sampled area to entire study area.
- Edge rule: when an organism lies on the boundary, count it if at least half of the organism is inside the quadrat; if less than half is inside, do not count.
- Example workflow in the transcript (quadrats for daisies):
- Area: 620 m^2 field; use a 1 m^2 quadrat grid.
- Sampling goal: sample ~10% of the total area (e.g., 30 squares; here 10 would be used).
- For each sampled square, count daisies within the square; determine if edge plants count by the 50% rule.
- Compute the mean number of daisies per square (per m^2) across all sampled squares.
- Estimate total population: multiply the average density by the total area:
- The example yields: average density = 7 daisies per m^2; total area = 620 m^2; estimated population = xˉimesA=7imes620=4340.
- Quadrat sampling advantages: simple, inexpensive, good for non-moving organisms, and flexible for different habitats.
- Limitations: may require many quadrats for precision; edge effects; assumes random distribution or representative sampling; potential observer bias in counting.
Systematic Sampling with Transects
- Systematic sampling collects data along a defined path (transect) using a regular interval.
- Common setup: lay a tape measure along the study area (e.g., 10 meters) and sample at fixed distances (e.g., every 1 meter).
- Procedure with quadrats along transect:
- Place a quadrat at the start of the transect and count organisms within.
- Move the quadrat 1 meter along the transect and count again; repeat along the entire transect.
- Data representation: line graphs or kite graphs can show how population density changes with distance along the transect.
- Example interpretation: density of daisies increases or decreases with distance from a gate; a biotic factor (e.g., trampling by people) likely affects abundance near high-traffic areas.
- Transect sampling helps detect spatial trends and the influence of environmental gradients or human activity on populations.
Comparison: Random Quadrat Sampling vs Systematic Transect Sampling
- Random quadrat sampling provides an unbiased estimate of mean density across the study area, good for heterogenous habitats.
- Systematic transect sampling helps identify spatial patterns and gradients (e.g., edge effects, pollution gradients).
- Both methods are valuable; the choice depends on the research question, habitat, and desired information (density vs distribution pattern).
- Quadrat-based sampling (already covered).
- Pitfall trap: a container flush with the ground; insects fall in and cannot escape; used to sample ground-dwelling arthropods.
- Sweeping nets: large nets for catching flying insects or those in vegetation.
- Beating tray: a tray placed beneath a branch; gently beat the branch to dislodge insects onto the tray for counting.
- Beating tray example: estimate insect populations by observing dropped insects; count and scale up.
- Kick sampling: in aquatic environments; disturb sediment to flush organisms into a net downstream; used for aquatic invertebrates and small fish.
- Light traps: a sheet or blanket placed with a light source at night to attract nocturnal insects for counting.
- Quadrat vs capture/recapture methods:
- Quadrat sampling is best for stationary organisms (plants, corals, barnacles).
- Capture-recapture (mark-recapture) is used for mobile animals (e.g., Florida panther), involving capturing, tagging, releasing, recapturing, and counting.
- Capture-recapture method (Mark-Recapture) basics:
- First capture: n1 individuals captured and marked.
- Second capture: n2 individuals captured; among these, m2 are marked from the first capture.
- Population estimate (assuming a closed population and marks retained):
- N^=m2n<em>1⋅n</em>2
- Assumptions and limitations of capture-recapture:
- Marks are not lost or overlooked; marked individuals mix back into the population; no immigration or emigration during the study; closed population during sampling period.
- Marks affecting recapture rates, or animals learning to avoid capture, can bias results.
- Random quadrat estimate (density-based extrapolation):
- If each quadrat area is a=1 m2 and total study area is A square meters, the estimated population size is:
- N^=xˉ⋅A, where \bar{x} is the mean number of individuals per quadrat.
- Example: density per m^2 = 7; total area = 620 m^2;
- N^=7×620=4340.
- Capture-recapture notation and formula:
- First capture: n<em>1, second capture: n</em>2, recaptured marked individuals: m2
- Estimated population: N^=m2n<em>1⋅n</em>2
Biodiversity and Diversity Indices
- Simpson's Index (for comparing biodiversity between areas):
- Let n<em>i be the number of individuals in species i, and N=∑</em>ini be the total individuals across all species.
- The standard formula for the probability that two randomly selected individuals belong to the same species is:
- D=N(N−1)∑<em>in</em>i(ni−1)
- The diversity measure can be reported as 1 − D (often called Simpson's Diversity Index) to reflect higher values for more diverse communities.
- Note: Simpson’s index is not a direct estimate of population size; it measures biodiversity and species evenness across areas.
High-Tech Data Collection and Analysis: Geospatial and Remote Sensing
- Geospatial systems (GIS): computer-based maps and analyses; data input by specialists and analysts; output can be static (maps) or interactive (GIS dashboards).
- Examples: maps used to assess drought conditions, wind farm siting, or crime distributions.
- Privacy and ethics: interactive GIS can reveal sensitive information (e.g., residential locations of offenders); access and privacy controls are important.
- Satellite data: numerous parameters collected from space (e.g., ozone, hurricanes, phytoplankton productivity, methane, atmospheric pollution, vegetation).
- Tracking and wildlife monitoring with satellites and radio collars: e.g., Florida panther monitoring to study range, reproduction, and mortality; tagging and re-release to study movement and population dynamics.
- Computer models: simulate weather and climate; two main timescales:
- Weather models: days to weeks
- Climate models: decades to centuries
- Key limitation of computer models: predictive reliability decreases the further into the future you forecast due to increasing variables and uncertainties.
- Hurricane forecasting example:
- Cone forecasts illustrate uncertainty grows with time; tracks become broader as projected horizon increases.
- The European (ECMWF) model is often more accurate for longer-range forecasts than some American models, though all models have strengths and weaknesses.
- Crowdsourcing and big data:
- Crowdsourced data sources include traffic data from smartphones (e.g., Google Maps using Bluetooth/Wi-Fi signals to infer crowd density) and social media.
- Crowdsourcing can greatly increase data volume but requires substantial processing power and sophisticated analytics.
- Privacy concerns and data ethics are central: data collection can reveal location, behavior, and other sensitive information.
- Big data: benefits include rich information and patterns that may not be visible in small datasets; limitations include processing time, storage, and potential for spurious correlations if not properly analyzed.
Practical Takeaways for Exams and Lab Work
- Distinguish clearly between hypothesis, theory, and laws (the transcript emphasizes that a theory is well-supported and not just a guess).
- Be able to identify independent vs dependent variables and control variables in a described experiment.
- Understand when to use random sampling (to avoid bias) vs systematic sampling (to examine spatial patterns or gradients).
- Know how to implement a quadrat-based population estimate, including edge rules and the steps to scale from a sample to an entire area.
- Be able to describe the capture-recapture method, its formula, and the assumptions that underlie it.
- Recognize different sampling tools (quadrats, pitfall traps, beating trays, kick sampling, light traps) and what they estimate.
- Understand the basics of biodiversity indices (e.g., Simpson’s index) and what they measure.
- Appreciate the role of GIS, satellites, computer models, and crowdsourcing in modern data collection and how these methods complement traditional field techniques.
- Be mindful of ethical considerations, privacy concerns, and avoiding bias in data collection and interpretation.
Quick Reference Equations
- Quadrat density-based estimate:
- N^=xˉ⋅A
- where xˉ is mean count per quadrat and A is total study area in square units (e.g., m^2).
- Edge counting rule for quadrats: count an organism if at least one-half of its body is inside the quadrat; otherwise do not count.
- Capture-recapture (Lincoln-Petersen) estimator:
- N^=m2n<em>1⋅n</em>2
- Definitions: n<em>1 = number captured and marked in first sample; n</em>2 = number captured in second sample; m2 = number of marked individuals recaptured in second sample.
- Simpson’s Diversity Index (probability-based):
- D=N(N−1)∑<em>in</em>i(ni−1)
- Biodiversity measure commonly presented as 1−D for intuitive interpretation (higher = more diverse).
Closing Notes
- Homework vs studying: homework is a task to be completed, often in-class or at home; studying involves independent review, questions, and practice to prepare for exams.
- The PowerPoints and videos are available for revision; use a mix of reading, watching, and practice problems to solidify understanding.
- Expect to encounter the above equations and concepts on exams; the lecturer will provide the exact definitions and variable meanings on tests.