Data Mining & Predictive Analytics Test #1

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/31

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 12:36 AM on 10/6/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

32 Terms

1
New cards

Produce summary of items in R (5 number summary)

summary(DN)

2
New cards

How to import(?) the data into R

dataname=read.table(“Data.Name.txt”)

3
New cards

Box Plot

Visual representation of our 5 number summary: min, Q1, median, Q3, max. Particulary useful in providing a visual comparison of two or more groups/variables when produced side-by-side and show evidence by outliers.

R Code: boxplot(DN$VN,…)

4
New cards

Histogram

Histogram shows the distribution of the data modality (bimodal, mimodal, multimodal) outliers, etc. Note: Some general drawbacks to the histogram, require a large number of observations to be beneficial.

The specification of the bins used to construct the histograms can greatly impact the visualization of the plot. Generally more difficult to compare multiple histograms to eachother.

R Code: hist(DN$VN,…)

5
New cards

Scatterplot

Drawbacks: when the plot contains a large number of observations all of the points can overlap and it may be hard to get a clear idea of the results.

Every individual needs to have measurements for each of the variables included in the plots. Therefore, the scatterplot can be highly susceptible to missing values.

Benefits: The scatterplot is particulary useful when examining the potential relationship between two variables.

** Can also be used with a grouping variable (categorical variable) to provide additional insight.

R Code: plot(DNVN1,DNVN1,DNVN2,…)

6
New cards

Confidence Interval

An interval that provides a certain amount of confidence that the unknown population paramter value is found within the bounds.

7
New cards

Interpretation of Confidence Interval

We are (1-a)100% confident that the true population parameter value is between the lower and upper bound.

Common Misconception: A (1-a)100% confidence interval does not indicate that there is a (1-a)100% probability that the interval contains the unknown population parameter.

8
New cards

Statistical Hypothesis Steps

  1. State the statistical hypotheses: The null hypothesis, assumed to be true, and the alternative hypothesis, needs evidence from the data to be claimed as being supported as true.

  2. Test Statistic: A numeric value calculated based on the data, used as evidence to evaluate the claims made in the hypothesis.

  3. Critical Value/p-value approach: Critical value approach sets up a region of test statistic values (known as the rejection region) that provides evidence of the alternative hypothesis being true based on the significance level alpha. P value approach computes a probability. If the p value is less than the significance level (alpha), the null hypothesis is rejected.

  4. Conclusion: Comprised of two statements. Fail to reject/reject null hypothesis. AND A statement as to what the results of the test indicate in terms of the question of interest.


9
New cards

Type I Error

The probability is denoted as X. When the null hypothesis is rejected by the statistical test, when in reality the null hypothesis is true.

R Code: p(TypeIerror)=x=p(rejectHo (line) Ho is true)

10
New cards

Type II Error

The probability is often denoted by B. When we fail to reject a false null hypothesis.

Note: We will generally never know if a statistical error has occured (such as a type I error) since the true population characteristics are never known.

11
New cards

What does QQ plot stand for? What is it used for?

Quantile-quantile plot, and used for testing if a data set is normal or not.

R Code: qqnorm(DNVN)qqline(DNVN) qqline(DNVN)

12
New cards

Normal Probability Plot

If the data points in the plot are close to a line, then the plot is providing evidence that the data follows a normal distribution. (qqplot/line)

Drawbacks: The interpretation of the plot may be subjective. But, it may provide useful information

13
New cards

Shapiro-Wilkes Test

Hypothesis test for normality, where null hypothesis is that the variable is normally distributed, and the alternative hypothesis is that the variable does not follow a normal distrubution.

p-value less than or equal to alpha, the null hypothesis is rejected and the conclusion is that the variable does not follow the normal distribution.

p-value greater than alpha, the null hypothesis is not rejected and there is not enough evidence to support that the variable does not follow the normal distribution.

R Code: shapiro.test(DN$VN)

14
New cards

Matched Pairs T Test

Common scenario for this is when the individually paired with themselves and observations are collected under two different conditions. This can be used to help control/account for a factor, that may impact the results.

15
New cards

Multiple Comparisons

When looking at datasets that are large with many variables a scenario that can occur is that need to conduct a large number of hypothesis tests is required.

16
New cards

Bonferroni Correction

To adjust the significance level for multiple comparisons using the Bonferroni correction, the alpha value is divided by the number of hypothesis tests being conducted.

alpha ad. Bon= alpha/n

Instead adjust the p value that results from the individual tests with

17
New cards

Sidak Correction

Adjust the significance level for multiple comparisons using Sidak correction.

It is possible for the value of the adjusted p value to be calculated as more than 1, when this occurs the adjusted p value is set to be equal to 1.

18
New cards

Chi Square Goodness of Fit test

Examines if the proportions of outcomes for a certain multinomial variable are different from specified values in data mining, this test is helpful when evaluating the data to determine if the sample that was collected reflects the population with respect to a certain categorical variable.

19
New cards

Hypothesis Test of interest for chi square goodness of fit test

Ho:p1=p10

Ha: At least one pi does not equal pio


Pio is the hypothized value of the population proportion of individuals in the ith category

20
New cards

Chi Sq test for homogenity proportions

multinomial variable being examined under two different conditions/scenarios.

The information used in this test can be presented as either the full dataset or presented as a summarized form in the form of a contingency table(two-way table).

21
New cards

Data Preparation for building mathematical models

Supervised methods: These are techniques where there is a primary response variable (target variable) and multiple observations where cases are available. Often allows for a prediction/classification of response variable values of an individual based on the predictor value.

Unsupervised methods: These are techniques where there is no primary response variable (target variable) available/identified. The statistical technique examines the data for paterns or structures within the variables. An example of this method that we will see would be a technique called cluster analysis.

22
New cards

Cross validation

Common technique implemented to help reduce the chances of drawing incorrect conclusions, or misclassifications, or producing misleading results.

General concept for cross validation: A partition of the data is used to build the mathematical model/alogirthm. Then, that model/algorithm is applied to a separate partition of the data, not used to develop the model algorithm to evaluate how well it performs.

23
New cards

Two fold Cross Validations

A method of cross validation where the data is divided into two proportions using a randomization technique.

One partition is known as the training set/data. The other partition is known as the test set/data.

Note: Another form of cross validation is k-fold cross validation. The data is partitioned into K independent subsets, the k-1 subsets are used as training sets to iteratively build a mathematical model/algorithm.

24
New cards

General Steps for building models/algorithms with cross validation

  1. Partition the data into a training and test set

  2. For categorical variables of interest(response variable) evaluate for differences in the proportions across the training and test set

  3. Build the model/algorithm based on the training data

  4. Evaluate the model/algorithm by applying it to the test data set.


25
New cards

Applications of chi-square test

It is possible to evaluate the data to examine if there is evidence that data does not reflect the population of interest with respect to a categorical response variable of interest.

This can be done with the chi-square goodness of fit test where the null hypothesis proportions are set to the population proportions for each category.

Note: A drawback is that it requires the population proportions to be known.

Once the data has been partitioned into the training set and test set it is possible to examine the data for a difference in proportions of categories for a categorical variable across the training and test set.

26
New cards

Overfitting

Occurs when the model built using the training set is overly complex and attempts to capture every aspect of the data(even the unimportant aspects, the noise in the data) This can increase the accuracy when applied to the test set, however can often result in a lack of general ability of the model outside the data set. Therefore, when only minimal accuracy is gained by increasing the complexity of the model, the model should no longer increase.

27
New cards

Regression

Examines the relationship between a response variable (target variable) and predictor variables (explanatory variable) This is an example of a supervised method.

28
New cards

Recall

Simple linear regression (a response variable and a single regression a response variable and a single explanatory variable)

29
New cards

Least Squares Regression

Slope!!!

y=bo+b1x

To produce estimated values for the response variable given a value of the predictor variable attach (DN)

R Code: reg.model=lm(RV~PV)

prediction.value=data.frame(PV=#)

30
New cards

Evaluating Model Assumptions

Normality:

The residuals follow a normal distribution. Can be evaluated with Shapiro-Wilkes test conducted on the residuals as well as the normal probability plot(q-q plot) of the residuals.

Constant Error Variance Assumption:

The variability of the residuals needs to be constant, can be evaluated using a plot of the residuals vs. the predicted values, examine if the vertical variability changes across the plot.

The mean of the residuals is zero:

Can be evaluated with a plot of the residuals vs. the predicted values to see if the residuals are centered around the line at 0.

Independence:

The residuals need to be independent of eachother.

31
New cards

Coefficient of Determination

Can be used to evaluate the usefulness of an estimated regression equation at predicting values of the response variable.

The closer r² is to 1, the more variability in y that is explained by the regression equation and the more evidence the regression model is a good fit and predictions would be reliable.

32
New cards

Interpretation of the slope coefficient

For a 1 inch increase in the base diameter of the tree the expected (mean/predicted) increase in the height is 2.1 feet for this species of tree. Interpretation of the y- intercept coefficient