Exploring and Pre-Processing Data

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/26

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 1:26 AM on 9/23/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

27 Terms

1
New cards

What are the diff Variable Types?

Data, Rows, and Columns

2
New cards

Rows can also be called

Record, Data Point, Instance, or Case

  • e.g. customer, tax return, applicant


3
New cards

Columns can also be called

Fields, attributes, Features, Dimensions

  • e.g. age, income, temperature


4
New cards

Nominal or Categorical Variables

Categorical - Not ordered can say X does not equal Y but not X>Y or X<Y

  • e.g. colors, phone numbers

  • NOMIAL DATA

Ordered(Ranked) - Ordered can say X>Y but not by how much

  • e.g. grades

  • ORDINAL DATA


5
New cards

Numerical or Continuous - Intervals

Ordered and can say the diff X-Y but not exactly the other operations such as multiplication

6
New cards

Numerical or Continuous - True Numeric

Support all mathematical operations, measures from a meaningful zero point

  • e.g. weight, height, length


7
New cards

Example of a variable that may look like numeric while it is categorical

Zip code, IP address

8
New cards

Some algorithms expect variables to certain types…

  • Statistical regression and Neural Networks expect NUMERIC inputs

  • Decision trees expect categorical or ordered inputs


9
New cards
<p>Numeric</p>

Numeric

  • Most algorithms can handle numeric data

  • May need to “bin” into categories


10
New cards
<p>Categorical </p>

Categorical

  • Naive Bayes can use it as-is

  • In most algorithms, must create n or n-1 binary numbers


11
New cards

Columns with unique values

  • Customer ID

  • Telephone number

  • Address, Zip Code

  • These categories are NOT that meaningful in mining data

    • Sometimes they can contain useful information

      • e.g. geographic, starting year, etc


12
New cards

Columns - Derived Variables

Results from calculations

  • e.g. total sales


13
New cards

Unary

  • Columns with one value

  • NO value for data mining

  • They should be ignored because they lack any information


14
New cards

Almost-unary

  • Columns with almost only one value

  • No VALUE for data mining

  • Rule of thumb: 95%-99% of the values are identical, the column should be ignored

  • Important to understand WHY the values are so skewed - why is it almost unary?


15
New cards

Discrete Values

  • Basic Plots

    • Line graphs

    • Bar Charts

    • Scatterplots

  • Distribution Plots

    • Histograms


16
New cards

Continuous Values

  • Mean, median, mode, range, variance, standard deviation

  • Boxplots

  • Correlation


17
New cards

Histograms

Shows the distributions of the outcome variable

  • e.g. median house value

  • bins of value - x, frequency counts - y


18
New cards

Range

Difference between the smallest and largest observation in the sample

  • Max - Min


19
New cards

Median

Midpoint of values after they have been ordered from smallest to largest, or largest to smallest

20
New cards

Mean V.S. Median

Median is not affected by the extreme values or outliers!

21
New cards
<p></p>


Symmetrical Distribution

  • Mean, median, and mode are symmetrical


22
New cards
term image

Positive Skew

  • Mean is right skewed


23
New cards
term image

Negative skew

  • Mean is left skewed


24
New cards
term image
  • Extreme Values → Very likely

  • Find outliers in age and income → remove them to clean data (not likely)

    • Outliers are an extreme values but extreme values are not outliers


25
New cards

Correlation

  • A measure to see the change in one variable to another

  • r is between [-1,+1]

    • The closer the value of r to 0, the smaller the correlation


26
New cards

What happens when you remove a variable - highly correlated

You can remove a variable and you will not MISS anything, if it is highly correlated

27
New cards

Exploring the data

Using statistical methods to gather and understand the dataset provided

  • Examine distributions

  • Study histograms

  • Investigate extreme values

  • Compare values with descriptions

  • Validate assumptions

  • Note prevalence of missing values