1/71
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Numerical variable
Quantitative, represents meaningful numbers
Categorical variables
Qualitative, represents categories
Discrete variable
Assumes a countable number of values
Continuous variable
Uncountable (Unlimited) values within an interval
Nominal scale
Uses names, labels or categories to sort items into distinct groups. Least sophisticated level of measurement
Ordinal scale
Uses numbers to categorize and rank items with respect to some characteristic or trait (Ex: Customer service ratings)
Interval scale
Uses numbers to categorize and rank items while finding meaningful differences between them (Ex: Temperature)
Ratio scale
Same as interval scales, but also has a true zero point (Ex: Sales, profit, inventory levels, …)
Omission strategy
Also called complete-case analysis, excludes observations with missing values from analysis
When to use omission strategy
When the amount of missing values is small and are expected to be distributed randomly across observations
Mean imputation strategy
Replacing missing values with the mean values across relevant observations
When not to use mean imputation strategy
When there is a large number of missing values, or if the missing data is not random (Ex: Sensitive data)
How to address missing values for a categorical variable
Use to most frequent category
Create an “Unknown” category
Subsetting
Extracting portions of a data set (variables) that are relevant to an analysis
Benefits of subsetting
Eliminate unwanted data
Can reveal insights in data (descriptive analytics)
Binning
Converts numerical variables into categorical variables by grouping numerical values into a small number of bins (groups)
Benefits of binning
Reduces noise in data caused by minor observation errors
Create equal intervals
Create bins of equal counts
Rescaling
Rescale numerical data so that variables in a data set are measured in a same scale
Category reduction
Collapse some categories to create fewer nonoverlapping categories
When do we use category reduction
When variables have too many categories
Variables have some categories that rarely occur
When a small sample doesn’t have any observations in certain categories
How to do category reduction
Create a “Other” category for categories with very few observations
Combine categories with similar impacts
Dummy variables
An indicator/binary variable, takes on values of 1 or 0 to describe 2 categories of a categorical variable
Amount of dummy variables to create
One less than the number of categories of the variable
Assigning bin number 1 to categories
The category you are testing or measuring
Assigning bin number 0 to categories
When the row does not match that category
Category scores
Use when the data is ordinal or have natural ordered categories
Appending data
Consolidating multiple data sources that share the same format and variables
Merging data
Using a common variable between 2 data sets to merge them together
=AVERAGE
Calculates the mean or average between a number of observations
=MEDIAN
Calculates the median as a measure of central location. Is the middle value of the ordered observations of a variable
=MODE
Calculates the value that occurs most frequently in a variable
Unimodal
If a variable has 1 mode
Bimodal
If a variable has 2 modes
Multimodal
If a variable has more than 2 modes
Range
Simplest measure of dispersion, is the difference between the maximum and minimum observations of a variable
Interquartile Range (IQR)
The difference between the third quartile and the first quartile. The range of the middle 50% of observations of the variable
Mean absolute deviation
An average of the absolute differences between the observations and the mean
2 most widely used measures of dispersion
Variance and standard deviation
Variance
An average of the squared differences between the observations and the mean
Standard deviation
The positive square root of the variance
Coefficient of variation (CV)
Standard deviation divided by the mean. Is a relative measure of dispersion and adjusts for differences in the magnitudes of the means
The Sharpe Ratio
Measures an investments excess return per unit of total risk or volatility
High Sharpe Ratio
The better the investment compensates its investors for risk
Skewness coefficient
The degree to which a distribution is not symmetric about its mean
Kurtosis Coefficient
Tells us whether the tails of the distribution are more or less extreme than the normal distribution
Covarience
The linear relationship between 2 variables
Correlation coefficient
The direction and the strength of the linear relationship between x and y (-1 to 1)
Boxplot
A summary that shows the min, Q1, Q2, Q3 and max value of the variable
Empirical rule 1
Approximately 68% of all observations fall in interval x + s
Empirical rule 2
Approximately 95% of all observations fall in the interval x + 2s
Empirical rule 3
Almost all (99.7%) of observations fall in the interval x + 3s
z score
Find the relative position of an observation by dividing the difference of the observation from the mean by the standard deviation
Treat observation as an outlier if z-score is
less than -3 or more than 3
In a boxplot, outliers are present when
They are farther than 1.5 x IQR from the IQR box
Lower fence of box plot
Q1 - (1.5 x IQR)
Upper fence
Q3 + (1.5 x IQR)
z-score formula
(xi - mean)/std dev
=COUNT
Counts the number of cells in range that contains values
=COUNTA
Counts the number of cells in a range that are not empty
=COUNTIF
Counts the number of cells in range that meet a certain criteria
=COUNTIFS
Counts the number of cells in range that meet multiple criterias
=COUNTBLANK
Counts the number of cells in a range that are blank
=LN
Returns the natural log of a certain cell
=SUM
Total value of all cells
=YEARFRAC
Returns the year fraction representing the number of whole days between 2 dates
=VLOOKUP
Searches for a specific value in the first column of a table and returns a value from another column in the same row
=MONTH
Returns the month from 1 (January) to 12 (December)
=MAX
Finds the maximum value of a range of cells
=MIN
Finds the minimum value of a range of cells
=STDEV.S
Finds the standard deviation based on a sample
=CORREL
Returns the correlation coefficient between 2 datasets
=VAR.S
Finds the variance based on the sample