1/49
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is statistics?
the science that deals with the collection, preparation, analysis, presentation, and interpretation of data
Descriptive Statistics
refers to the summary of important aspects of a data set
Includes collecting, organizing, and presenting the data in the form of charts
and tables
Often calculate numerical measures (typical value, variabiliy)
Inferential Statistics
refers to drawing conclusions about a larger set of data (population) based on a smaller set of data (sample).
Population
consists of all items/members of interest
Sample
A sample is a subset of population
Need for Sampling
to make inferences about various characteristics of the population.
Cross Sectional Data
refers to data collected by recording a characteristic of many subjects at the same point in time, or without regard to differences in time
Time Series Data
refers to data collected over several time periods focusing on certain groups of people, specific events, or objects.
• can include hourly, daily, weekly, monthly, quarterly, or annual observations
Structured Data
• Reside in a pre-defined, row-column format.
• Spreadsheet or database applications.
• Enter, store, query, and analyze.
• Numerical information that is objective and not open to interpretatio
Unstructered Data
• Do not conform to a pre-defined, row-column format.
• Textual and multimedia content.
• Do not conform to database structures.
• These data may have some implied structure.
• Do not conform to a row-column model required in most database systems
Big Data
• A massive volume of structured and unstructured data.
• Extremely difficult to manage, process, and analyze using traditional data processing tools.
• Presents great opportunities to gain knowledge and game-changing intelligence
Qualitative variables
• Also called Categorical Data.
• Represent categories.
• Labels or names to identify distinguishing characteristics.
• Can be defined by two or more categories.
• Coded into numbers for data processing
Quantitative variables
• Also called Numeric Data.
• Represent meaningful numbers.
• Either discrete or continuou
Nominal
• Least sophisticated.
• Represent categories or groups.
• Values differ by label or name.
• Example: marital status
Ordinal
• Stronger level of measurement.
• Categorize and rank data with respect to some characteristic.
• Cannot interpret the difference between the ranked values, numbers are arbitrar
Interval
• Categorize and rank, differences are meaningful.
• Zero value is arbitrary and does not reflect absence of characteristic.
• Ratios are not meaningful
Ratio
• Strongest level of measurement.
• A true zero point, reflects absence of characteristic.
• Ratios are meaningful
Inspecting and preparing the data for analysis
• Counting and sorting.
• Handling missing values.
• Subsetting
Frequency Distribution for Qualitative data
Frequency Distribution for Quantitative data
Constructing a frequency distribution for a numerical variable
details about classes and observations
Relative Frequency
proportion of items
Percent Frequencies
percentage of items
Pie Charts (uses)
a segmented circle whose segments portray the relative frequencies of the categories of a qualitative variable
Bar Charts
depicts the frequency or relative frequency for each category of the categorial variable
Series of either horizontal or vertical bar
bar lengths proportional to the values they are depicting
Histogram
the numerical variable counterpart to the vertical bar chart for a categorical variable
provides information on the shape of the distribution
Ogive
a cumulative frequency polygon or line graph used in statistics to show how many data values lie above or below a specific value
Polygon
a line graph that shows data distribution by connecting the midpoints of the tops of the bars
Scatterplots Positive Linear
displays two quantitative variables that increase together
Scatterplots Negative Linear
displays data points that slope downward from left to right. As the x-value increases, the y-value decreases in a straight-line pattern
Scatterplots Curvilinear
displays data points that follow a smooth curved pattern rather than a straight line, showing that two variables change together at a changing rate instead of a constant one
Scatterplots No Relationship
shows a random spread of dots with no upward trend, downward trend, or clear pattern
Arithmetic Mean Median
Add all numbers together.
Divide the total by how many numbers exist.
Effect: It is easily skewed by very large or very small numbers (outliers
Arithmetic Mean Average
Sort all numbers from least to greatest.
Pick the exact middle number.
If there is an even count of numbers, take the arithmetic mean of the two middle numbers.
Effect: It resists being skewed by extreme high or low values
Median
The middle point of a data set
Mode
The most frequent value in a data set; how many can be in a data se
Boxplot 5 value
Minimum, Quartile 1, Median, Quartile 3, Maximu
Interquartile Range
Quartile 3 minus Quartile 1
Percentiles
divides a variable into two parts
a technical measure of location and relative position
Geometric mean: the characteristics; when it is useful
measures the rate of change of a variable over time.
smaller than the arithmetic mean
less sensitive to outliers
relevant measure when evaluating investment returns over several years and calculating compoud growth rates
Variance and Standard Deviatio
Most widely used measures of dispersion
Square root of Variance =
Standard Deviation
Coefficient of Variation
A relative measure of dispersion.
Calculated as Std.
Dev. divided by the Mean
Chebyshev’s Theorem
applies to any data shape
Empirical Rule
applies to relatively symmetric and bell shaped date
Analysis of relative location (what capabilities does it provide; what statements does it make)
defines a place or data point by its spatial relationship to other areas, providing context on connectivity, accessibility, and economic potential
Covariance
Measures the direction of the linear relationship between two variables
Correlation coefficient shows
the direction and the strength of the linear relationship between two variables
Correlation coefficient values fall between
-1 (negative relationship) +1 (positive relationship), and 0 (no relation)
Correlation coefficient calculated as
the Covariance of two variables divided by the product of their Standard Deviation