Data 101 Week 1

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/59

flashcard set

Earn XP

Description and Tags

Structural concepts, data types, summary metrics, and data quality issues

Last updated 5:35 PM on 9/28/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

60 Terms

1
New cards

Column

A vertical strip in a table representing one attribute

2
New cards

Attribute

What a column is called when a row is treated as an object — a property attached to that record

3
New cards

Feature

A column used as input to a predictive model

4
New cards

Variable

A column in a statistical model

5
New cards

Field

The database-world term for a column

6
New cards

Value

The smallest unit of data in a table — a single cell's content

7
New cards

Cell

The spreadsheet term for a value

8
New cards

Observation

A statistician's term for an entire row — everything measured about one subject

9
New cards

Data point

A value plotted on a chart

10
New cards

Tuple

The formal database-theory term for a row — an ordered sequence of values

11
New cards

Record

The everyday database/software-engineering term for a row

12
New cards

Instance

The machine-learning term for a row

13
New cards

Dataset

The umbrella term for the entire table — every row and column together

14
New cards

Data frame

A rectangular dataset with labeled columns

15
New cards

Schema

The blueprint listing every column's name and type

16
New cards

Cardinality

The count of distinct values in a column

17
New cards

Distribution

The overall shape formed by all values in a column

18
New cards

Missingness

How much of a column or dataset is empty or null

19
New cards

Granularity

The level of detail each row represents (e.g. daily vs. hourly)

20
New cards

Sparsity

How much of a dataset is empty

21
New cards

Dimensionality

The number of columns/features a dataset has

22
New cards

Integer

A whole-number numeric type (e.g. age)

23
New cards

Double/Float

A decimal-precision numeric type (e.g. GPA)

24
New cards

String

A text type that can't meaningfully be averaged but can be alphabetized or searched

25
New cards

Factor

A categorical type with a fixed set of ordered levels (e.g. class year)

26
New cards

Boolean

A type holding only true/false (or 1/0) values

27
New cards

NA/NULL

A recognized marker for a missing value

28
New cards

Mean

The sum of all values divided by count — highly sensitive to outliers

29
New cards

Median

The middle value when sorted — robust against outliers

30
New cards

Mode

The most frequent value — the only central-tendency measure for categorical data

31
New cards

Variance

The average of squared deviations from the mean

32
New cards

Standard deviation

The square root of variance

33
New cards

Range

Maximum minus minimum — relies only on the two extremes

34
New cards

Interquartile Range (IQR)

Q3 minus Q1 — the spread of the middle 50%

35
New cards

Skewness

The directional asymmetry of a distribution

36
New cards

Kurtosis

Tail weight and peakedness of a distribution relative to normal

37
New cards

Outlier

A value sitting unusually far from the rest of the dataset

38
New cards

Population

The entire group under study

39
New cards

Sample

A subset of the population actually observed

40
New cards

Sample size (n)

The number of observations in a sample

41
New cards

Projection

The formal term for selecting specific columns

42
New cards

Split-apply-combine

The pattern behind group-by operations: split into groups

43
New cards

Misspelling

A text-pollution issue where variant spellings split true counts during aggregation

44
New cards

Case inconsistency

When identical values appear with different capitalization across rows

45
New cards

Extra spaces

Invisible leading/trailing/internal whitespace that breaks exact-match filters

46
New cards

Encoding issue

When text is read with the wrong character encoding

47
New cards

Placeholder noise

Missing data disguised as valid text like "unknown" or "N/A"

48
New cards

Exact duplicates

Identical rows appearing multiple times byte-for-byte

49
New cards

Near duplicates

Rows referring to the same entity with slightly different text (e.g. "Jon" vs "John")

50
New cards

Redundant columns

Separate columns storing identical information under different names

51
New cards

Label variation

A single category recorded under multiple text representations (e.g. "USA" vs "US")

52
New cards

Contradictory fields

Two columns in the same row holding mutually exclusive value

53
New cards
plot()
Creates base R scatterplots and general plots
54
New cards
boxplot()
Creates box-and-whisker plots showing distributions across groups using formula syntax
55
New cards
barplot()
Generates bar graphs from a vector of height values
56
New cards
mosaicplot()
Creates a mosaic plot to display categorical data relationships
57
New cards
ggplot()
Initializes a ggplot2 visualization object with input data
58
New cards
geom_bar()
Adds a bar chart layer to a ggplot object
59
New cards
geom_boxplot()
Adds a boxplot layer to a ggplot object
60
New cards
aes()
Constructs aesthetic mappings connecting data columns to visual properties (x