1/39
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Statistics
A way to transform data into information. Creating new understanding from a set of numbers.
Descriptive Statistics
Summarizes and describes the variables in a dataset using numeric and visual techniques.
Inferential Statistics
Uses sample data to make predictions or decisions about a larger population.
ex. estimation, hypothesis testing, regression. Make generalizations or predictions.
Population
The entire set of entities that we wish to study.
Sample
A subset of the population
Parameter
A descriptive measure of a population. A fixed measure that will describe the entire population.
So given the variable “has blonde hair” which can either be true or false. The parameter would be the total number of people in the population that have blonde hair.
Statistic
A descriptive measure of a sample. It’s valuye depends on which sample is selected. For example we grab a sample size n from the larger population of N people.
Then a statistic would be that x people out of a sample size n have blonde hair.
Confidence level
The proportion of times taht an estimating procedure will be correct.
ex. 95% confidence level means estimates based on this form of inference will be correct 95% of the time
Significance level
Measures how frequently the conclusion will be wrong in the long run.
1 - confidence level
Sampling plan
Is a method or procedure for specifying how a sample will be taken from a population.
1. define population
the sampling method
Sample size
Simple random sampling
Given a population of N people, each individual in the population has an equal weighted chance of being selected into our sample. The issue is that we would need a complete list of the entire population.
Replacement
Replacement: when doing simple random sampling we add the subject back to the population so they could be selected again
Non-replacement: we remove them from the pool so they could not be selected again
Stratified random sampling
Here the population is divided up into stratum, s.t. the population is grouped. Then we go into each one of these groups and we conduct random sampling within each of the stratum
Systematic random sampling
Identify the size of the population ex. N
Identiify sample size n, you wish to get to
k = N/n, which will tell you your sampling interval\
Then choose a random starting point i between 1 and k, 1<=i<=k
Include the units at positions i, i + k, i + 2k…
Cluster sampling
The population is divided into naturally occurring groups, called clusters. A random sample of these clusters is chosen, and every (or some) members of each selected cluster is studied.
For example, if a population has 20 clusters, a researcher might randomly select 10 of them and collect data from all members of those 10 clusters.
Convenience sampling
Selecting people from the total population to be in the sample where the data is easy to collect. However this is dangerous and can lead toward conveince bias.
Sampling error
The difference between the parameter and the statistic
Non-sampling error
Non-sampling error is any error that isn't caused by random chance in choosing the sample. It comes from how the data is collected, recorded, or processed.
Observation / Variable
Observations (also called cases or individuals) are the entities being studied, such as people, animals, or companies. Each observation is one row in the data table.
Variables are the characteristics recorded about each observation, such as age, height, or favorite color. Their values can differ from one observation to the next. Each variable is one column.
Cross-sectional data
Cross-sectional data is collected on many entities at one point in time, with each entity measured once.
Example: the scores from one class exam. Every student is measured once on the same test, so the data shows how students compare at that moment.
Time-series data
Time series data is collected over time, with each row representing a time period (day, month, year, etc.) and the variables measured at each one. It's used to track changes, spot trends, and forecast future values.
It often follows a single entity, but it can include more than one.
Example: the daily closing prices of Apple and Microsoft over the past year, with one row per day and one column per stock.
Time series data vs. cross-sectional data
Time series vs. cross-sectional: In time series data, each row is a time period, and the goal is to see how variables change over time. In cross-sectional data, each row is an entity (a person, company, etc.), and all entities are measured at roughly the same moment.
It doesn't have to be the exact same second, but time isn't a factor. The goal is to compare entities, not track change.
Qualitative Data
Two categories
Nominal data: Catagories with no inherent order ex. (is X, vs. is Y) (true false type)
Ordinal data: Catagories with an order ex. a course rating where it is a scale
Quantitative Data
Interval data: equal spacing between values, but zero is arbitrary and doesn't mean "none." Differences make sense, but ratios don't. Ex: temperature in °F or °C, calendar years.
Ratio data: equal spacing and a true zero that means a complete absence. Differences and ratios both make sense. Ex: weight, height, income, age.
Continous vs. Discrete Data
Continous: Can take any value within a range
Discrete: Can onyl take specific seperate valyes, usually whole numbers ex. a count
Relative frequency
frequency of the category / n
Lets say that you have 3 categories, then the sum of all of the relative frequencies should add up to one
Bar chart
Used for nonimal data, where the order does not matter. It descirbes the count, or the frequency for which some variable appears in the data set for each category.P
Pie chart
Also used for nominal data where the order does not matter. Describes the relative frequency for each category and shows them adding up to 1 whole “pie”
Grouped frequency distribution
Displays the number of observations in each of the distributions distinct classes. Used for continous data or when there are many discrete values.
Use a histogram
How to make a histogram
determine the number of classes using k >= log2n
bin width = (max - min) / k
start at the min value in the dataset and count all of the values that belong to that bin
Histogram
A type of bar graph that shows the distribution of a set of numerical data. Counts the frequency for data that belongs in differnt bins / classes
A histogram is symmetric when a vertical line down the middle splits the two sides into identical shape and size. (No skew)
Skewness
When there is a “tail” on one side of the histogram.
Postive / right - A bulk of the observations are on the left side
Negative / left - A bulk of the observations are on the right side
Modality
A mode is the class / value which is observeed the most times in the dataset.
In a histogram modality describes the peaks. A histogram can have multiple peaks and be bimodal (a special case of multimodal)
Joint freq distribution / relative freq
Shows the relationship between two or more variables by counting how many observations fall into each combination of categories or intervals
summarizes how many observations falls into a catergory
summarizes how many observations fall into a class / interval
We can also represent these observations as a percentage of the grand total which would be a joint relative frequency table
measures of center
mean: population and sample mean
population = μ
sample = xˉ
The median has two cases, either odd values, in which case you do (n + 1) / 2
Or even values where you find average between n / 2 and (n / 2) + 1
mode
the value or category that occurs most frequently
right vs. left skew
median < mean (right skew)
median > mean (left skew)
The extreme values that are included in the calculation of the mean will bring the tal toward the mean
range
largest observation - smallest observation
Variance
sigma ^ 2 = sum (xi - u) / N (population)
sigma ^ 2 = sum (xi - x bar) / n -1 (sample)
Standard deviation
How far on average a value deviates from the mean it is sqrt( sigma ^ 2)