1/53
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Population
the entire collection of all individuals or items under consideration in a statistical study.
• The target ____ should be clearly defined since there are ____ within _____.
Population size
total number of individuals or items in the population under study.
Census
collecting data from the entire population.
Often too expensive or even impossible to undertake; therefore, a sample is taken.
Sample
a subset of individuals from the population. Data are only recorded on these individuals.
a relatively small number of observations from the population being investigated
Can be used to mean one observation or a collection of measurements from a population.
Sample size
number of individuals or observations in a single sample
Inferential statistics
uses information from a sample to make decisions, conclusions, and predictions about the entire population
Parameter
descriptive measure of a population (symbolized by Greek letters),
population mean (u)
population standard deviation (omega)
slope of the populationregression line 1 (beta).
Statistic
a descriptive measure of a sample used to estimate a parameter,
sample mean y (with line on top)
sample standard deviation S
and slope of the sample regression line (beta1 with ^)
Components of research design
• Study units (within the context of the target population and study area/sites)
• Variables
• Spatial Aspects of Design
• Temporal Aspect of Design
• Techniques and Methods of Data Collection
• Sampling strategy & Randomness: More about Where and When
• Overall Type of Research Design: Observational versus Experimental
Study units
individuals, or subjects (people, animals, objects, or things) about which information is required and on which measurements are recorded.
AKA units of analysis or cases.
In an experimental study = experimental units
observation study = units of observation.
In agricultural research, these are the pre-determined plots where different treatments are applied.
Social sciences = “who”
Variable
characteristic that varies from one study unit (individual, subject, person, or thing) to another.
social sciences, these may be opinions, behavior, attitudes, perception, etc.
In physics – weight, force, energy, light, etc.
In biology – growth rate, chlorophyll content, height, density, color, etc.
Distribution of a variable
all the values that a variable takes on.
Data
the values of a variable, i.e., actual measurements/observations recorded for each variable.
Datum
individual piece of data (an observation) or a single measurement.
Categorical/Qualitative variable
categorical variable is a nonnumerically valued variable and does not follow an ordered sequence.
On nominal scale
Values of the variable are classified by some quality or attribute, i.e., the values are put into categories.
cannot be measured, but rather, the frequencies of individuals in the categories are counted to obtain numbers to analyze the data.
Examples: color, regions of a country (north, south, east, west); marital status; like or dislike a certain hobby or activity; yes, no, or indifferent to something; types of animals; types of items to purchase; languages.
Ordinal scale
Data or observations which can be put in order from lowest to highest, but which do not have a constant interval between successive units, i.e., the data can be ranked.
Relative magnitudes are known, so many types of statistical analysis can be applied.
• scale of 1 – 5 can be used, where 1 = very poor, 2 = poor, 3 = moderate, 4 = good, 5 =very good.
Qualitative scale
numerically valued variable.
Constant interval size between successive units
Discreate/continuous quantitative variables
Discrete
quantitative variable that can only take on specific values, usually whole numbers.
a countable variable.
Continuous quantitative variable
quantitative variable that can have an infinite number of values between any observed range.
a measurable variable.
Indicator variable
dummy variable
Categorical variables that are coded in order to obtain quantitative variables that can be analyzed using hypotheses tests like ANOVA and regression.
Ex. Responses to a question, Yes or No, can be coded as 1 = Yes, 0 = No.
Derived variable and indices
This is when a combination of categorical, ordinal, and/or quantitative variables are recorded and the values are multiplied, divided, added, or otherwise combined to obtain one derived quantitative variable with units (e.g., basal area of mangroves calculated as cm2/25-m2 plot).
Or they can be combined to obtain an index (without units).
Explanatory/predictor/independent variable
variables of interest that are hypothesized to explain or affect other variables in the study, but which are not likely to be affected by those other variables
application of variables must either precede or occur during the same time period as the expected reaction of the response variable.
Response/dependent variable
variable that is hypothesized to be affected by the explanatory or independent variables.
Extraneous variables
possible explanatory variables that are NOT of interest or are NOT related to the purpose of the study, though they could be of interest in a different study.
May potentially affect the response variable, interfere with the study and lead to “experimental error“
Lurking, confounding, hidden variables
Factors
explanatory variables, which are categorical, are applied as treatments in an experiment or considered as levels in an observational study
Spatial aspects of design
• The “Where” component of research design.
• Linked to the study units – where are they sampled and measured?
• Involves the way the observations or replicates are arranged in space (distance, area, or volume).
• Study sites and geographical location of target populations.
Temporal aspects of design
• The “When” component of research design – the way observations or replicates are arranged in time.
• Time period (year, month, time of day) and frequency of observations.
• Start, end, frequency of recording the variables
Techniques and methods of data collection
• The “How” component of research design.
• Specific methods and techniques used to take measurements of the variables or to record data.
• The specific techniques to be applied will differ from one field of natural or social science to the other
Random sampling
the selection of individuals or study units from a population without bias, such that:
1. All individuals have an equal chance of selection (or each possible sample of a given size is equally
likely to be the one obtained).
2. The selection of individuals is independent, i.e., the selection of one does not affect the selection of others.
ensures that the sample is as representative as possible of the entire population.
simple random sampling (SRS), systematic random sampling,
stratified random sampling, and cluster random sampling.
Sampling strategies
Study units or individuals on which data are recorded must be randomly sampling from the target population so that they truly represent the population
in a given study must specify the way observations are recorded in space and time.
must eliminate bias as much as possible because bias over-emphasizes or under-emphasizes some characteristics of the population.
Observational studies
when conducted in order to get opinions from people it is often called a sample survey or social survey
opinions from people it is often called a sample survey or social survey.
Aims at estimating population parameters.
Researcher collects data about a particular phenomenon as it occurs in nature or in society.
Randomness
Variables of interest are measured or recorded for the study units.
No imposing of treatments on the subjects or individuals.
No manipulation or control of any variables or conditions.
Extraneous variables cannot be controlled.
Population inferences
can be made if there is random selection from the target population in observational studies.
Can be made if there is random selection from the target population in experiments
Causal inferences
can NOT be made in observational studies, that is, causation or cause-and-effect relationships among variables can NOT be established because there are so many unmeasured factors (or extraneous variables) that may affect the variable being measured.
can be made if there is random assignment to treatment and control groups In experiments
Experimental studies sampling randomness
First, study units (experimental units) are randomly selected from the target population.
Secondly, the experimental units are randomly assigned to treatment and control groups
Treatment groups
exposed to new conditions, that is, one or more levels of the predictor variable or factor being manipulated; the treatments are imposed upon these groups.
Control groups
exposed to the usual level of the manipulated variable or not exposed to it at all
case of human subjects, the group receives a placebo, so the subjects don’t know whether they are receiving a treatment or not
Replication is required to
Check or confirm the results,
Apply statistical analysis – analysis is based on replicates, and
Increase the power of the test.
No. of replicates = no. of observations in a sample or sample size (n).
Describing quantitative data
1. Shape – symmetric, left skewed, or right skewed.
2. Center – the middle of the distribution.
3. Spread – variation or dispersion of the distribution.
Mean
center of gravity of the distribution (e.g., histogram).
not a resistant measure of center, because it is seriously influenced by skewness (pulled in the direction of a few extreme observations).
Best for symmetric distributions
Median
divides the area under the curve into two equal halves.
resistant measure because it is more robust to extreme values or skewness than the mean and therefore is a better measure of center for a very skewed distribution.
Mode
only measure of center that can also be used for qualitative data. Distributions may be unimodal, bimodal, or multimodal.
Five-number summary
Min, Q1, Q2, Q3, Max
Outliers
observations that lie outside the overall pattern of the data.
• May be due to recording error, may belong to a different population, or may just be unusually extreme observations.
Graphs used for quantitative data
1. Histograms
2. Dotplots
3. Stemplots
4. Boxplots
5. Normal probability plots
6. Scatter diagrams (xy graphs)