1/48
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Statistics
the science and art of collecting analyzing and drawing conclusions from data
Individual
An individual is a person, animal, or thing described in a set of data.
Variable
is any attribute that can take different values for different individuals
Categorical vs Quantitative Variables
A categorical variable takes values that are labels, which place each individual into a particular group, called a category.
A quantitative variable takes number values that are quantities—counts or measurements.
Population
The population in a statistical study is the entire group of individuals we want information about.
Census
A census collects data from every individual in the population.
Sample
A sample is a subset of individuals in the population from which we collect data.
Observational Study vs. Experimental Study
An observational study observes individuals and measures variables of interest, but does not attempt to influence the responses.
An experiment deliberately imposes treatments (conditions) on individuals to measure their responses.
Bad Sampling
A voluntary response sample - consists of people who choose to be the sample by responding to a general invitation.
A convenience sample consists of individuals from the population who are easy to reach
The design of a statistic study shows BIAS if It is very likely to underestimate or very likely to overestimate the value you want to know
Good Sampling
a random sample consists of individuals from the population who are selected for the sample using a chance process
Under Coverage Bias
occurs when some members of the population are less likely to be chosen or cannot be chose for the sample.
Non Response Bias
occurs when an individual chosen for the sample can’t be contacted or refuses to participate.
Response Bias
Occurs when there is a consistent patter of inaccurate responses to a survey question
Simple Random Sampling SRS
A simple random sample (SRS) of size n is a sample chosen in such a way that every group of n individuals in the population has an equal chance of being selected as the sample.
Sampling Variability
The fact that different random samples of the same size from the same population produce different estimates is called sampling variability.
Systematic Random Sampling
Systematic random sampling selects a sample from an ordered arrangement of the population by randomly selecting one of the first k individuals and choosing every kth individual thereafter.
Stratified Random Sampling
Strata are groups of individuals in a population that share characteristics thought to be associated with the variables being measured in a study.
Stratified random sampling selects a sample by choosing an SRS from each stratum and combining the SRSs into one overall sample.
Cluster Random Sampling
A cluster is a group of individuals in the population that are located near each other.
Cluster random sampling selects a sample by randomly choosing clusters and including each member of the selected clusters in the sample.
Confounding
• Confounding occurs when two variables are associated in such a way that their effects on a response variable cannot be distinguished from each other.
Conducting a well designed experiment is the best way to avoid confounding.
Response Variable
A response variable measures an outcome of a study.
Explanatory Variable
An explanatory variable may help explain or predict changes in a response variable.
Treatment
A treatment is a specific condition applied to the individuals in an experiment
Experimental Unit
An experimental unit is the object to which a treatment is randomly assigned. When the experimental units are human beings, they are often called subjects
Placebo
A placebo is a treatment that has no active ingredient but is otherwise like other treatments
Control Group
A control group is a group used to provide a baseline for comparing the effects of other treatments
Random Assignment
In an experiment, random assignment means that treatments are assigned to experimental units (or experimental units are assigned to treatments) using a chance process.
Replication
Replication is the idea that we should use enough subjects to create roughly equivalent groups.
Placebo Effect
The placebo effect describes the fact that some subjects in an experiment will respond favorably to any treatment, even an inactive treatment (placebo).
Double Blind Experiment
In a double-blind experiment, neither the subjects nor those who interact with them and measure the response variable know which treatment a subject is receiving
Single Blind Experiment
In a single-blind experiment, either the subjects or the people who interact with them and measure the response variable don't know which treatment a subject is receiving.
Benefits of randomized designs
One reason to keep other variables the same for each subject is to prevent confounding.
Another reason is to reduce the variability in the response variable, making it easier to determine if one treatment is more effective than another.
experimental units must be assigned to the treatments completely at random.
Statistically Significant
When an observed difference in responses between the groups in an experiment is so large that it is unlikely to be explained by chance variation in the random assignment, we say that the result is Statistically Significant
Distribution
The distribution of a variable tells us what values the variable takes and how often it takes each value.
Frequency Table
A frequency table shows the number of individuals having each data value.
Relative Frequency
A relative frequency table shows the proportion or percentage of individuals having each data value. typically in a percentage!
Bar Chart
shows each category as a bar, the heights of the bars show the categories frequencies or relative frequencies.
Pie Chart
shows each category as a sector of a circle the areas of the sectors are proportional to the category frequencies or relative frequencies. uses categorical variables.
Two Way Table
A two-way table is a table of frequencies (or relative frequencies) that summarizes the relationship between two categorical variables for some group of individuals.
side by side bar chart
A side-by-side bar chart displays the distribution of a categorical variable for each value of another categorical variable. The bars are grouped together based on the values of one of the categorical variables and placed next to each other.
segmented bar chart
A segmented bar chart displays the distribution of a categorical variable as portions (segments) of a rectangle, with the area of each segment proportional to the percentage of individuals in the corresponding category.
Association
There is an association between two variables if knowing the value of one variable helps us predict the value of the other. If knowing the value of one variable does not help us predict the value of the other, then there is no association between the variables.
In a positive association, values of one variable tend to increase as the values of the other variable increase.
In a negative association, values of one variable tend to decrease as the values of the other variable increase.
Shapes of Quantitative Data
If the frequencies are about the same for all values, we say the distribution is ( approximately uniform )
one major peak (unimodal) … bimodal, trimodal, and multimodal when there are three or more major peaks.
Look for clusters of values and obvious gaps. Decide if the distribution is ( roughly symmetric or clearly skewed )
Direction of skewness is toward the long tail, not the direction where most observations are clustered.
Center
Center: Value where half the dots are at that point or above (and half the
dots are at that point or below).
Variability
Variability: For starters, we will give the range. As the semester progresses,
we will quantify "how spread out the data ponts are"
Outlier
As for clear departures from the pattern - perhaps the most important is an outlier - a value that falls outside the overall pattern.
Stemplots
A stemplot shows each data value separated into two parts: a stem, which consists of the leftmost digit(s), and a leaf, the final digit. The stems are ordered from least to greatest and arranged in a vertical column. The leaves are arranged in increasing order out from the stems.
Histograms
A histogram shows each interval as a bar. The heights of the bars show the frequencies or relative frequencies of values in each interval.
Scatterplot
the relationship between two quantitative variables measured on the same individuals.The values of one variable appear on the horizontal axis and the values of the other variable appear on the vertical axis. Each individual in the data set appears as a point in the graph.
scatterplot can show a weak, moderate, or strong association. An association is strong if the points don’t deviate much from the form identified.