1/47
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Data
Collections of observations, such as measurements, or survey
responses.
Statistics
The science of planning studies and experiments; obtaining data; and organizing, summarizing, presenting, analyzing, and interpreting those data and then drawing conclusions based on them
Population
Complete collecting of all measurements or data that are being considered. It is the complete collection of data that we would Like to make inferences about.
Sample
A sub-collecyion of members selected from a population
Census
Collection of data from every member of the population
Process involved in Statistical Study
Prepare
Analyze
Conclude
Prepare
Components:
Context
Source of the data
Sampling method
Analyze
Components:
Gather the data
Explore the data
Apply statistical method
Conclude
Significance
Misleading Conclusions
Avoid making statements not justified by statistical analysis. Make it clear
Correlation
When two things move together. Variable A goes up, variable B goes up too
Causation
One things causes the other. A creates change in variable B
The “third variable” problem
A third factor (C) is driving both A and B
Reverse Casuality
B is actually causing A, not the other way around.
Pure Coincidence (Spurious Correlation)
With enough data points, random noise will naturally form matching patterns purely by chance.
Loaded Questions
Survey questions can be “loaded” or intentionally worded to elicit a desired response
Order of Questions
Sometimes survey questions are unintentionally loaded by such factors as the order of the items being considered.
Non-response
occurs when someone either refuses to respond to a survey question or is unavailable. When people are asked survey questions, some firmly refuse to answer
Percentage
Some studies cite misleading or unclear percentages. Note that 100% of some quantity is all of it, but if there are references made to percentages that exceed 100%, such references are often not justified
Parameter
A numerical measurement describing some characteristic of a population.
Statistic
A numerical measurement describing some characteristic of a sample
Qualitative (categorical data) Data
consist of names or labels (not numbers that represent counts or measurements).
Quantitative (numerical) data
consist of numbers representing counts or measurements
Discrete Data
result when the data values are quantitative and the number of values is finite or “countable.”
Continuous Data
result from infinitely many possible quantitative values, where the collection of values is not
countable
Levels of Measurement
Nominal level
Ordinal level
Interval level
Ratio level
Nominal level
simplest form of measurement. At this level, data is purely categorical. Numbers or labels are used only to identify or classify objects.
No mathematical value, no logical order, and no distance between categories count frequencies (mode) and calculate percentages. You cannot add, subtract, or average them.
Ex: sex, marital status, types of software
Ordinal level
the data can be placed in a logical order or rank. However, while you know which value is "greater" than another, you do not know the exact distance between them.
There is a clear rank or position, but the intervals between the ranks are not necessarily equal
Median and Mode. You still cannot perform meaningful addition or subtraction.
Scales, education level, race results
Interval level
possesses all the characteristics of the ordinal level, with the added benefit that the space (interval) between values is equal and measurable.
Equal distances between points. However, there is no "True Zero." A value of zero does not mean the total absence of the variable
You can calculate the Mean, Median, and Mode. You can add and subtract, but you cannot multiply or divide
Temperature in Celsius or Fahrenheit, IQ scores, Dates Years
Ratio level
most complex and informative level of measurement. It has all the properties of the interval level, plus a "True Zero" point.
A value of zero means the variable is completely absent. This allows for the comparison of magnitudes (ratios)
All operation possible
Mean, median, mode and geometric mean
Height, weight, age, income, number of students in a classroom
Big Data
refers to data sets so large and so complex that their analysis is beyond the capabilities of traditional software tools.
Analysis of big data may require software simultaneously running in parallel on many different computers.
Missing Data
missing completely at random if the likelihood of its being missing is independent of its value or any of the other values in the data set. That is, any data value is just as likely to be missing as any other data value.
A data value is missing not at random if the missing value is related to the reason that it is missing
Delete cases
One very common method for dealing with missing data is to delete all subjects having any missing values.
Impute missing values
We impute missing data values when we substitute values for them. There are different methods of determining the replacement values, such as using the mean of the other values, or using a randomly selected value from other similar cases, or using a method based on regression analysis
Gold Standard
Randomization with placebo-treatment groups is sometimes called the because it is so effective. (A placebo such as a sugar pill has no medicinal effect.
Experiment
we apply some treatment and then proceed to observe its effects on the individuals. (The individuals in experiments are called experimental units, and they are often called subjects when they are people.)
Observational
we observe and measure specific characteristics, but we don’t attempt to modify the individuals being studied
Replication
repetition of an experiment on more than one individual. Good use of replication requires sample sizes That are large enough so that we can see effects of treatments.
Blinding
used when the subject doesn’t know whether he or she is receiving a treatment or a placebo. It is a way to get around the placebo effect, which occurs when an untreated subject reports an improvement in symptoms. (The reported improvement in the placebo group may be real or imagined.
Randomization
used when individuals are assigned to different groups through a process of random selection,
Simple Random Sampling
A sample of n subjects is selected in such a way that every possible sample of the same size n has the same chance of being chosen
Systematic Sampling
we select some starting point and then select every kth (such as every 50th) element in the population
Convenience Sampling
we simply use data that are very easy to get
Stratified Sampling
we subdivide the population into at least two different subgroups (or strata) so that subjects within the same subgroup share the same characteristics (such as gender). Then we draw a sample from each subgroup (or stratum)
Cluster Sampling
we first divide the population area into sections (or clusters). Then we randomly select some of those clusters and choose all the members from those selected clusters.