BUS 230 Quiz 1

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/71

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 4:04 AM on 9/30/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

72 Terms

1
New cards

Numerical variable

Quantitative, represents meaningful numbers

2
New cards

Categorical variables

Qualitative, represents categories

3
New cards

Discrete variable

Assumes a countable number of values

4
New cards

Continuous variable

Uncountable (Unlimited) values within an interval

5
New cards

Nominal scale

Uses names, labels or categories to sort items into distinct groups. Least sophisticated level of measurement

6
New cards

Ordinal scale

Uses numbers to categorize and rank items with respect to some characteristic or trait (Ex: Customer service ratings)

7
New cards

Interval scale

Uses numbers to categorize and rank items while finding meaningful differences between them (Ex: Temperature)

8
New cards

Ratio scale

Same as interval scales, but also has a true zero point (Ex: Sales, profit, inventory levels, …)

9
New cards

Omission strategy

Also called complete-case analysis, excludes observations with missing values from analysis

10
New cards

When to use omission strategy

When the amount of missing values is small and are expected to be distributed randomly across observations

11
New cards

Mean imputation strategy

Replacing missing values with the mean values across relevant observations

12
New cards

When not to use mean imputation strategy

When there is a large number of missing values, or if the missing data is not random (Ex: Sensitive data)

13
New cards

How to address missing values for a categorical variable

  1. Use to most frequent category

  2. Create an “Unknown” category


14
New cards

Subsetting

Extracting portions of a data set (variables) that are relevant to an analysis

15
New cards

Benefits of subsetting

  1. Eliminate unwanted data

  2. Can reveal insights in data (descriptive analytics)


16
New cards

Binning

Converts numerical variables into categorical variables by grouping numerical values into a small number of bins (groups)

17
New cards

Benefits of binning

  1. Reduces noise in data caused by minor observation errors

  2. Create equal intervals

  3. Create bins of equal counts


18
New cards

Rescaling

Rescale numerical data so that variables in a data set are measured in a same scale

19
New cards

Category reduction

Collapse some categories to create fewer nonoverlapping categories

20
New cards

When do we use category reduction

  1. When variables have too many categories

  2. Variables have some categories that rarely occur

  3. When a small sample doesn’t have any observations in certain categories


21
New cards

How to do category reduction

  1. Create a “Other” category for categories with very few observations

  2. Combine categories with similar impacts


22
New cards

Dummy variables

An indicator/binary variable, takes on values of 1 or 0 to describe 2 categories of a categorical variable

23
New cards

Amount of dummy variables to create

One less than the number of categories of the variable

24
New cards

Assigning bin number 1 to categories

The category you are testing or measuring

25
New cards

Assigning bin number 0 to categories

When the row does not match that category

26
New cards

Category scores

Use when the data is ordinal or have natural ordered categories

27
New cards

Appending data

Consolidating multiple data sources that share the same format and variables

28
New cards

Merging data

Using a common variable between 2 data sets to merge them together

29
New cards

=AVERAGE

Calculates the mean or average between a number of observations

30
New cards

=MEDIAN

Calculates the median as a measure of central location. Is the middle value of the ordered observations of a variable

31
New cards

=MODE

Calculates the value that occurs most frequently in a variable

32
New cards

Unimodal

If a variable has 1 mode

33
New cards

Bimodal

If a variable has 2 modes

34
New cards

Multimodal

If a variable has more than 2 modes

35
New cards

Range

Simplest measure of dispersion, is the difference between the maximum and minimum observations of a variable

36
New cards

Interquartile Range (IQR)

The difference between the third quartile and the first quartile. The range of the middle 50% of observations of the variable

37
New cards

Mean absolute deviation

An average of the absolute differences between the observations and the mean

38
New cards

2 most widely used measures of dispersion

Variance and standard deviation

39
New cards

Variance

An average of the squared differences between the observations and the mean

40
New cards

Standard deviation

The positive square root of the variance

41
New cards

Coefficient of variation (CV)

Standard deviation divided by the mean. Is a relative measure of dispersion and adjusts for differences in the magnitudes of the means

42
New cards

The Sharpe Ratio

Measures an investments excess return per unit of total risk or volatility

43
New cards

High Sharpe Ratio

The better the investment compensates its investors for risk

44
New cards

Skewness coefficient

The degree to which a distribution is not symmetric about its mean

45
New cards

Kurtosis Coefficient

Tells us whether the tails of the distribution are more or less extreme than the normal distribution

46
New cards

Covarience

The linear relationship between 2 variables

47
New cards

Correlation coefficient

The direction and the strength of the linear relationship between x and y (-1 to 1)

48
New cards

Boxplot

A summary that shows the min, Q1, Q2, Q3 and max value of the variable

49
New cards

Empirical rule 1

Approximately 68% of all observations fall in interval x + s

50
New cards

Empirical rule 2

Approximately 95% of all observations fall in the interval x + 2s

51
New cards

Empirical rule 3

Almost all (99.7%) of observations fall in the interval x + 3s

52
New cards

z score

Find the relative position of an observation by dividing the difference of the observation from the mean by the standard deviation

53
New cards


Treat observation as an outlier if z-score is

less than -3 or more than 3

54
New cards

In a boxplot, outliers are present when

They are farther than 1.5 x IQR from the IQR box

55
New cards

Lower fence of box plot

Q1 - (1.5 x IQR)

56
New cards

Upper fence

Q3 + (1.5 x IQR)

57
New cards

z-score formula

(xi - mean)/std dev

58
New cards

=COUNT

Counts the number of cells in range that contains values

59
New cards

=COUNTA

Counts the number of cells in a range that are not empty

60
New cards

=COUNTIF

Counts the number of cells in range that meet a certain criteria

61
New cards

=COUNTIFS

Counts the number of cells in range that meet multiple criterias

62
New cards

=COUNTBLANK

Counts the number of cells in a range that are blank

63
New cards

=LN

Returns the natural log of a certain cell

64
New cards

=SUM

Total value of all cells

65
New cards

=YEARFRAC

Returns the year fraction representing the number of whole days between 2 dates

66
New cards

=VLOOKUP

Searches for a specific value in the first column of a table and returns a value from another column in the same row

67
New cards

=MONTH

Returns the month from 1 (January) to 12 (December)

68
New cards

=MAX

Finds the maximum value of a range of cells

69
New cards

=MIN

Finds the minimum value of a range of cells

70
New cards

=STDEV.S

Finds the standard deviation based on a sample

71
New cards

=CORREL

Returns the correlation coefficient between 2 datasets

72
New cards

=VAR.S

Finds the variance based on the sample