1/111
All vocab (finish this week)
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Statistics
Data, functions of data (mean and range), techniques for collecting, analyzing, and interpreting data for subsequent decision making, the science of creating and applying such techniques
Population
Collection of ALL people, objects, or events having one or more specified characteristics.
Element
A SINGLE person, object, or event of a population
Observation/Datum
Number used to represent an element in the population (measurable characteristic of the popualtion)
Sample
Proper subset of the population
Descriptive Stats
Tools for depicting or summarizing data so that they can be more readily comprehended
Inferential Statistics
procedures for using sample data to make inferences about one or more population parameters
Sampling Fluctuation/Chance Variability
Elements obtained differ from sample to sample
Range
Set of elements for which the variable stands
Value
Each element of the range
Variable
A characteristic that can take on different values (sometimes denoted as X or Y)
Constant
A characteristic that does not vary (pi)
Qualitative Variable
Symbol whose range consists of attributes or non quantitative characteristics of people, objects, or events (hair color, men, women, etc.)
Quantitative Variable
A symbol whose range consists of a count or a numerical measurement of a characteristic
Discrete
Range can assume only a finite number of values or an infinite number of values that are countable (gap in between numbers)
Continuous
Variables range is uncountable infinite (always approximate because of earth limitations)
Unordered Qual Variable
Categories do not suggest an order or rank
Ordered Qual Variable
Categories suggest order or rank
Measurement
Process of assigning numbers or labels to characteristics of people, objects, or events according to a set of rules
Nominal Measurement
Assigning elements to mutually exclusive/exhaustive equivalence classes so that those in the same class are considered to be equivalent to one another. The classes are then denoted by a set of distinct labels, which make a nominal scale
Ordinal Measurement
Consists of assigning elements to mutually exclusive or exhaustive equivalence classes that are ranked or ordered with respect to one another, The classes are then denoted by numbers or other ordered symbols (i.e. letters in the alphabet), that reflect the rank of the classes. The labels assigned to equivalence classes in ordinal measurement have the properties of distinctness and order, which make the ordinal scale
Interval Measurement
The numbers assigned to EC have distinctness, order, AND equal differences between numbers reflect equal magnitude differences between the corresponding classes. The measurement procedure consists of defining a unit of measurement and determining the number of units required to represent the difference between equivalence classes, and these numbers make the interval scale.
Ratio Measurement
The numbers assigned to an EC are distinct, ordered, equivalent in intervals, and the origin of the scale represents the absence of the measured characteristic (true 0). Set of numbers makes a ratio scale.
Equivalence Classes
A single score value, a collection of score values, a qualitative category (all can be a type of equivalence class)
Frequency Distributions
A table showing the equivalence classes and the frequency with which their score values occur (does not show distinct original information though)
Class Intervals
Equivalence classes (where everything is equal in one) of frequency distributions (so a range of numbers)
Real Limits
Extend 0.5 below the nominal lower limit and approximately 0.5 above the nominal upper limit
Class Interval Size (i)
Constructed using Real Limits (i = Real upper limint - Real lower limit)
Relative frequency distribution
A distribution that shows the prop f and %f for each class interval (which express each frequency as either a proportion or percentage of the total number of scores)
Cumulative Frequency Distribution (Cum f)
Shows the number, proportion, or percentage of scores that occur below the real upper limit of each class interval
Kurtosis
Property of being peaked, flat, or somewhere in between (for graphs)
Mesokurtic
Normal distribution (normal kurtosis level)
Platykurtic
Flatter distribution
Leptokurtic
Slender, narrower, more peaked distribution
Dispersion
Extent to which scores are spread out around a central point
Bimodal
Two humps, same maximum frequency
Uniform Distribution
Rectangle, each class interval has the same frequency (percentiles)
Mode (Mo)
The score or qualitative category that occurs at the greatest frequency
Mean (x bar)
Sum of scores divided by the number of scores (average)
Statistic -
Descriptive measure of a sample
Parameter -
Descriptive measure of a population
Median (Mdn)
Point a distribution that divides the data into two groups having equal frequency
Interpolating
Dividing the class interval containing the median into subintervals and finding the point that represents the (n + 1)/2th score or the point that was midway between the (n/2)th and the (n2)/1th scores
Weighted Means
Finding the means from two means (class averages, etc.)
Terminal Statistics
Usefulness in advanced statistics procedures is limited (Mode and Median)
Range
The distance between the largest and smallest scores
Semi-Interquartile Range (Q)
One half of the distance between the first quartile point, Q1, and the third quartile point, Q3, with Q2 being the midpoint of the data.
Standard Deviation (S for a sample and sigma for a pop)
Distances from midpoint
Index of Dispersion
DP/DPmax = number of distinguishable pairs to the maximum possible number of distinguishable pairs (instead of representing distance)
Inflection Points
Points where the curve changes from convex to concave and reverse (on a distribution)
Outliers
Scores that are unusually large or small relative to other scores
Whiskers
Extend from either side of IQR to the outermost data points that fall within the distance computed (do not contain the outliers, if any)
Measures of Dispersion
Summarize the extent to which scores differ from one another (IQR, Rande, Index of D, etc.)
Independent Variable
Variable controlled or manipulated by a researcher
Dependent Variable
Effected by independent variable
(Pearson product-moment) Correlation Coefficient (rxy or r for sample, ppop is rho)
Measure of the linear relationship (degree of association) between two quantitative variables, X and Y
Cross Product (Xi - X bar)(Yi - Y bar)
Product of the two variations, shows which quadrant the point is located on
Covariance (Sxy)
The directional relationship between two variables (direction of line of best fit)
Coefficient of Determination (r^2)
Ranges from 0 to 1 that examines how the differences in one variable can be explained by the differences in the second variable, hwne predicting the outcome of a given event
Coefficient of nondetermination (K^2 = 1 – r^2)
Shows the amount of variation in a dependent variable that is not explained by a regression model (extraneous factors)
Correlation Ration or Eta Squared (n with a squiggle ^2)
Used to determine strenght of association between nonlinearly related variables
Truncated
Restricted
Heterogeneity of Array/Heteroscedasticity
Presence of skewed X any Y distributions is often accompanied by an unequal dispersion of Y scores for different values of X and vice versa
Homogeneity of Array Variances/Homoscedasticity
Dispersions of X and Y are uniform
Spearman rank correlation coefficient (rs)
Describes the degree of agreement between paired data that are in the form of ranks
Tied Ranks
Two individuals assigned the same rank
Concomitance -
state of existing or happening at the same time
Regression Analysis
A relationship used to estimate a relationship between a dependent outcome (Y) and an independent outcome (X)
Multiple Regression
Simultaneous use of two or more independent variables in predicting a dependent variable (exercise AND diet to predict future weight) (predicted on a regression plane)
Prediction Error/Residual
The difference between the actual score and a persons predicted score
Line of Best Fit -
a line that minimizes the sum of the squared prediction errors
Standard Error of Estimate (Sy.x)
Statistical measure that shows how much actual data points stray from predicted regression line (assumes values between 0 and Sy)
Coefficient of Multiple Correlation (R²)
Measures the strongest linear relationship between one dependent variable and a group of two or more independent variables (extension of r²)
Multicollinearity
When two or more input variables in a regression model are closely tied together/presence of nonzero correlations among the independent variables (can be bad)
Subjective Personalistic View of Probability
Probability is a measure of the strength of one’s expectations that an event will occur
Classical/Logical View of Probability
The probability of an event, A, is given by the number of events favoring A (priori knowledge of previous events)
Empirical Relative-Frequency View of Probability
Defines the chance of an event as the ratio of the number of times the event occurs to the total number of experimental trials or observations (infinity)
Event
An observable happening
Mutually Exclusive Events
When two events contain NO sample points in common
Exhaustive Events
Events for which the probability of their union equals 1 (one of the events must occur)
Random Variable
a numerical outcome of an experiment
Probability Distribution
A table showing the possible values of a random variable and the associated probaabilities
Sampling Distribution
Probability distribution of a statistic (shown as a random variable) which is based on the results of more than one trial
Random Sampling
Drawing samples form a population so every possible sample of a particular size has the same probability of being selected
Random Variables
A mathematical function that assigns a real number to each possible outcome of a random event or experience
Expected Value
Average outcome you expect to see if you repeat a random event many times
Bernoulli Trials
1) only one of two outcomes, 2) prob of success remains constant from trial to trial, 3) The outcomes of successive trials are independent
Multinomial Distribution
extension of binomial distribution: k > equal to 2 classes and probabilities associated with the classes remain constant (sampling with replacement or iinfinite pop) (3+)
Hypergeometric Distribution
Outcome of k > or equal to 2 BUT the probabilities DO NOT remain constant (sampling without replacement from finite pop)
Standard Score
A number that expresses the value of a score relative to the mean and the standard deviation of its distribution
Percentile Rank of a Score -
Indicates the percentage of the scores of the distribution that falls below the score
Point Estimate (Estimation [inferential stat procedure])
One number representing the estimate is associated with a point on a real number line
Interval Estimate (Type of estimate, 2 numbers)
2 numbers and associated points define an interval on the real number line
Estimator
A rule that tells you how to calculate an estimate of a population parameter using sample information
Law of Large Numbers
The larger the sample size, the more probable it is that the sample mean comes arbitrarily close to the population mean
Standard Error of a Statistics (sigma with subscript of statistic to which it applies)
Sampling distributions standard deviation
Unbiased Estimator
Statistic whose long-term expected value equals the true population parameter it estimates
Minimum Variance Estimator
Unbiased estimator with the smallest variance (thus highest precision) of any unbiased estimator for all possible values of a population parameter
Test Stat
A statistic that is used to test hypotheses about the values of population parameters
Statistical Inference
Making decisions abut the population by using samples that contain only a small portion of the population