PSY201

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/53

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 6:11 PM on 6/15/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

54 Terms

1
New cards

what are the two branches of statistics?

descriptive and inferential

2
New cards

what is the descriptive branch of statistics?

characterize attributes of samples and/or populations; cleans up piles of raw human data so that patterns can be seen; without it, psychological data is essentially useless; it can be used for variability, central tendency, and distribution shapes

3
New cards

what is the inferential branch of statistics?

generalize from a sample to an unknown population; use a sample to make smart guesses (estimates) about an unknown population; make leaps from data you don’t have; it is impossible to test all 8 billion people on earth; used for hypothesis testing and estimation

4
New cards

what are frequency distributions?

a method used to organize a pile of raw data to communicate the headcount at each category on the scale of measurement, or how many observations exist; take unreadable numbers and turn them into a pattern so you can visualize clusters of data; it takes table and graph/figure form

5
New cards

what is the table form for frequency distributions?

a structured tally list (like listing quiz scores on the left and counts on the right)

6
New cards

what is the graph/figure form for frequency distributions?

a visual picture of the data (like a histogram or a bar graph)

7
New cards

why is it important to be careful of instructions?

there is no universal design law, so you can’t know if you need a table or graph every time; always read the titles, column labels, and tiny captions on an exam question carefully to see how that specific table or graph is styled

8
New cards

what are the three types of frequency distributions?

simple, relative, cumulative

9
New cards

what is the simple frequency distribution?

it is used for raw counts or frequencies; use it when you have a small, tight range of scores and you need to know the exact headcount for each score; it gives the raw truth about the sample, seeing how many people scored a perfect 10 on their wellness check; it becomes useless when it covers a massive range of numbers, as it violates the “fewer than 20 rows” rule and becomes unreadable

10
New cards

what is the relative frequency distribution?

proportions and percentages relative to the total sample size (N); you use it when you need to compare two different studies with different sample sizes, or when the raw headcount doesn’t matter as much as the overall rate; it standardizes the data, allowing for a fair comparison

11
New cards

what is the cumulative frequency distribution?

used for running totals by adding from the bottom upward; when you need to determine relative standing, percentiles, or cutoffs; allows you to see where a subject sits from the rest of the crowd

12
New cards

what are the different options for different data?

numerical (quantitative) data and categorical data

13
New cards

what is numerical data?

data made of real numbers where the values tell you how much of something exists; it works with simple, relative, and cumulative tables; it works because numbers have a natural, mathematical order

14
New cards

what is categorical data?

data split into distinct groups, names, or labels; only works with simple and relative tables; you cannot make a cumulative table for unordered categorical data

15
New cards

what is the protocol for grouped frequency distributions?

when the data covers a wide range of scores, representing all possible numbers is not feasible; limit the number of rows to less than 20; group the scores into intervals that include a range of scores; assign frequencies to these intervals; all intervals must be the same width

16
New cards

what are relative frequency distributions?

each score represented as a proportion (or percent) of the total sample (N)

17
New cards

what is the equation for proportion?

f / N; column adds to 1.0; like a probability

18
New cards

what is percentile rank?

% of observations with the same or smaller value

19
New cards

what is an alternate formula for cumulative percent?

cf/N * 100

20
New cards

what are histograms?

bar charts where the vertical bars are jammed tightly together with absolutely no spaces between them; could plot simple f, p, or %; X values (or midpoints of intervals) appear on X axis; plot each f with a bar, equal size, touching; no gaps; in grouped tables, write the midpoint instead along the bottom

21
New cards

what are frequency polygons?

line graphs that connect dots with straight lines to look like a mountain range; a connect-the dots histogram; make a point representing f for each possible X value; connect the dots; anchor the line on X axis; useful for comparing distributions in two samples (in this case, plot p greater than f)

22
New cards

what are the graphs for numerical (continuous) data?

when your data is made of real numbers that flow sequentially; histograms and frequency polygons

23
New cards

why are there no gaps in between numerical data?

the lack of space represents that the scale is unbroken, meaning you can score anything

24
New cards

what are the graphs for categorical data?

when your data consists of distinct groups, names, or labels; bar graphs

25
New cards

what are bar graphs?

similar to histograms, but the bars must have physical gaps/spaces in between them; for categorical data; useful for showing two samples side by side

26
New cards

why does there have to be gaps in between categorical data?

a deliberate statistical signal; they are complete separate entities, with no smooth transition or numerical flow in between

27
New cards

how do you determine whether to use a histogram/frequency polygon vs. a bar graph?

look at the data type

28
New cards

what is a normal distribution?

a symmetrical, theoretical frequency distribution where scores cluster naturally around a central value; bell-shaped, unimodal, and symmetrical; accurately models most natural human traits, such as height, weight, intelligence, and personality traits

29
New cards

what is a bimodal distribution?

a frequency polygon with two clear peaks (humps) that is typically symmetrical; a major red flag that your sample contains two distinct subgroups hidden inside the data; combined biological sex data (height, weight) or proficiency data (mixing experts with beginners on a test)

30
New cards

what is a positive skew?

the massive peak is on the left (low scores), and a long, thing tail points to the right (high scores); always name a skew by its tail, not its hum; for variable where most people naturally score very low, but a small handful of outliers score extremely high (e.g. household wealth, frequency of criminal behavior, or clinical symptom check-lists like anxiety/depression)

31
New cards

what is a negative skew?

the peak is on the right, and the tail stretches to the left; it typically occurs for positively valenced variables in a general population; most people naturally score high on these traits, while only a tiny, extreme minority scores low

32
New cards

what are the uses for figures?

personal use for understanding data (scatterplots, histograms); professional publications like for conference presentations (bar graphs, line graphs); publications for general audience (more advanced infographics)

33
New cards

what are scatterplots?

used to display raw data for two continuous variables simultaneously to check for an underlying relationship or correlation; every dot represents a single observation mapped across both the x-axis and the y-axis; confirms if a relationship exists before running complex analysis; instantly exposes extreme data points or measurement anomalies that standard averages mask

34
New cards

what are line graphs?

ideally used for a truly continuous, numerical X-axis (like time or age); allowed for ordinal categories (like low vs. high training) because they possess a sequential mathematical direction; completely illegal for nominal categories (like blood type or country), because connecting unrelated labels with a line falsely implies a continuous transition between them

35
New cards

when must a graph’s axis include 0 vs. when can it be zoomed in?

line graphs - can omit 0 and focus on a specific section of the range when including 0 would completely flatten and obscure a misleading pattern of change (e.g. tracking small growths or subtle score changes); always check the y-axis baseline before interpreting any visual trendwha

36
New cards

what are the ten ways to a great figure?

minimize the junk; plan before you start creating; say what you mean, mean what you say; label everything; communicate ONE idea; keep things balanced; maintain the scale in the graph; simple is best and less is more; limit the number of words; the chart alone should convey what you want to say

37
New cards

what are error bars?

show measurement precision and variability within groups; may represent confidence interval, standard error, standard deviation

38
New cards

what is a central tendency?

a statistical measure that identifies a single score as the most typical or representative of an entire distribution; "what single score best represents the center of this dataset?"; mean (mathematical average), median (middle score by rank), and mode (most frequent score); when data is skewed, these three measures split apart, forcing the researcher to decide which "center" is truly representative.

39
New cards

what is the mean?

the arithmetic average (sum of all scores divided by number of scores); sample mean uses English letters - n = sample size; population mean uses Greek letter "mu" (\mu); N = total population size; the mathematical expression remains the same, but the notation changes to explicitly signal whether the data represents a subset (sample) or the entire group (population); critical for inferential statistics; most useful overall; most common descriptive statistics; used for interval and ratio data

40
New cards

what is the balance point property of the mean?

the sum of the deviations from the mean will always equal zero; how far a raw score (X) lands from the mean (M), calculated as (X - M); mean acts as the mathematical fulcrum of the dataset - the total distance of scores above the mean perfectly balances out the total distance of scores below it; serves as an excellent quality-control check when calculating advanced metrics later

41
New cards

what is the primary limitation of the mean, and how does it relate to distribution shapes?

the mean is highly sensitive to outliers, particularly in small samples; because the mean balances total distance weight rather than headcount, a single massive outlier forces the mean to shift positions to maintain balance; sensitivity is the direct cause of skewed distributions; the mean gets pulled away from the main cluster of data and drags the distribution's tail toward the outlier

42
New cards

what is the median?

the literal middle location in an ordered distribution; corresponds exactly to the 50th percentile (divides data into two equal 50% halves); you MUST arrange the raw data points from lowest to highest before locating the middle value; it is completely insensitive to outliers; because it only measures positional rank rather than mathematical weight, extreme scores on the ends cannot drag the median away from the true center

43
New cards

what is the precise median for continuous variables?

when you have multiple tied scores sitting directly in the middle of a continuous distribution; treats the tied score as an interval defined by its real limits (Lower Real Limit to Upper Real Limit)

44
New cards

what is the mode?

the score or category that occurs with the greatest frequency (f) in a distribution; two peaks do not have to be exactly equal in frequency to qualify a graph as bimodal or multimodal; a secondary, distinct peak that is slightly lower than the major mode but still represents a meaningful, popular cluster; exposes hidden subgroups within a single sample

45
New cards

what is the skew divergence rule?

normal curve - symmetrical, all three measures are approximately equal (mean = median = mode); positive skew - tail points right, the mean gets dragged highest (mode < median < mean); negative skew - tail points left, the mean gets dragged lowest (mean < median < mode); the mode always stays at the peak, the mean is always pulled closest to the tail, and the median always sits in the middle

46
New cards

what are the central tendency rules for nominal data?

mode is always appropriate; median is banned, unless the categories are ordinal and can be ranked; mean is never appropriate, because nominal categories have no numerical magnitude or inherent sequence; even if categories are assigned code numbers, calculating an average or midpoint is inherently meaningless

47
New cards

when to use the median?

it is the mandatory choice when a distribution is skewed or contains heavy outliers, because it tracks position rather than weight; used for continuous data; use for ordinal data if ranked, since data can be sorted sequentially; nominal use is strictly banned; it is not highly useful for inferential statistics because its positional nature makes difficult to use in advanced algebraic equations

48
New cards

what are undetermined values and open-ended distributions?

when a participant's exact score cannot be completed/measured (e.g., quitting a timed task); when a data category has no fixed mathematical boundary (e.g., scoring "5 or more"); the mean is strictly banned in both scenarios because it requires calculating an exact \Sigma X - the median is the mandatory choice because it handles placeholders easily via positional ranking

49
New cards

what is variability?

a statistical measure that describes the amount of spread, scatter, or distance between scores in a distribution; scores are tightly clustered around the center, so the mean is highly precise and accurately represents the individual scores; scores are widely spread out, so the mean carries a high margin of error and is a less reliable representation of individual scores

50
New cards

why is variance bad for descriptive stats, and what is the fix?

variance leaves scores in squared unites, which are impossible for humans to interpret descriptively; take the square root of the variance; converts the metric into the standard deviation, returning the score to its original, real-world unit of measurement so it can cleanly describe the average distance scores sit from the mean

51
New cards

what is the population standard deviation formula?

sigma = population standard deviation; mu = population mean; N = total number of scores in the entire population; measures the standard, real-world average distance that scores drift away from the central mean

52
New cards

what is the definitional formula for population variance?

the exact mathematical average of the squared deviations from the mean

53
New cards

definitional vs. computational formulas

definitional is designed for theoretical understanding, every part of the equation clearly shows where the numbers come from conceptually (distance from the mean); computational formula is designed as a calculation shortcut - it rearranges the algebra to make calculations faster and prevent decimal rounding errors, but its individual parts do not help you understand the concept; both yield the exact same numerical result

54
New cards

why are naturally less variable than populations? (the bias concept)

small samples naturally cluster around the center and rarely capture the extreme outliers that exist in a large population; because they lack outliers, calculating sample variance using the standard population formula will systematically underestimate the true population spread; we use an adjusted formula with a smaller denominator, dividing the sum of squares by n-1 instead of N, artificially boosting the score to give an unbiased estimate