1/59
Structural concepts, data types, summary metrics, and data quality issues
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Column
A vertical strip in a table representing one attribute
Attribute
What a column is called when a row is treated as an object — a property attached to that record
Feature
A column used as input to a predictive model
Variable
A column in a statistical model
Field
The database-world term for a column
Value
The smallest unit of data in a table — a single cell's content
Cell
The spreadsheet term for a value
Observation
A statistician's term for an entire row — everything measured about one subject
Data point
A value plotted on a chart
Tuple
The formal database-theory term for a row — an ordered sequence of values
Record
The everyday database/software-engineering term for a row
Instance
The machine-learning term for a row
Dataset
The umbrella term for the entire table — every row and column together
Data frame
A rectangular dataset with labeled columns
Schema
The blueprint listing every column's name and type
Cardinality
The count of distinct values in a column
Distribution
The overall shape formed by all values in a column
Missingness
How much of a column or dataset is empty or null
Granularity
The level of detail each row represents (e.g. daily vs. hourly)
Sparsity
How much of a dataset is empty
Dimensionality
The number of columns/features a dataset has
Integer
A whole-number numeric type (e.g. age)
Double/Float
A decimal-precision numeric type (e.g. GPA)
String
A text type that can't meaningfully be averaged but can be alphabetized or searched
Factor
A categorical type with a fixed set of ordered levels (e.g. class year)
Boolean
A type holding only true/false (or 1/0) values
NA/NULL
A recognized marker for a missing value
Mean
The sum of all values divided by count — highly sensitive to outliers
Median
The middle value when sorted — robust against outliers
Mode
The most frequent value — the only central-tendency measure for categorical data
Variance
The average of squared deviations from the mean
Standard deviation
The square root of variance
Range
Maximum minus minimum — relies only on the two extremes
Interquartile Range (IQR)
Q3 minus Q1 — the spread of the middle 50%
Skewness
The directional asymmetry of a distribution
Kurtosis
Tail weight and peakedness of a distribution relative to normal
Outlier
A value sitting unusually far from the rest of the dataset
Population
The entire group under study
Sample
A subset of the population actually observed
Sample size (n)
The number of observations in a sample
Projection
The formal term for selecting specific columns
Split-apply-combine
The pattern behind group-by operations: split into groups
Misspelling
A text-pollution issue where variant spellings split true counts during aggregation
Case inconsistency
When identical values appear with different capitalization across rows
Extra spaces
Invisible leading/trailing/internal whitespace that breaks exact-match filters
Encoding issue
When text is read with the wrong character encoding
Placeholder noise
Missing data disguised as valid text like "unknown" or "N/A"
Exact duplicates
Identical rows appearing multiple times byte-for-byte
Near duplicates
Rows referring to the same entity with slightly different text (e.g. "Jon" vs "John")
Redundant columns
Separate columns storing identical information under different names
Label variation
A single category recorded under multiple text representations (e.g. "USA" vs "US")
Contradictory fields
Two columns in the same row holding mutually exclusive value