Notes for Statistics: Displaying the Order in a Group of Numbers; Central Tendency and Variability and Inferential Statistics
Displaying the Order in a Group of Numbers Using Tables and Graphs
- Chapter framing: Statistics as a branch of mathematics focusing on organization, analysis, and interpretation of numbers; treated here as descriptive of data in psychology and related fields.
- The Two Branches of Statistical Methods
- Descriptive statistics: summarize and describe a group of numbers (tables, graphs, etc.).
- Inferential statistics: draw conclusions and make inferences beyond the observed data about a larger population.
- This chapter focuses on descriptive statistics to build intuition for later inferential methods.
- What statistics is for psychologists
- To describe data, test ideas, compare groups, and evaluate research reports in media.
- Use software (e.g., SPSS) but building hands-on understanding by hand reinforces procedure.
- The learning approach in the text
- Small, simple numeric examples per chapter to emphasize underlying logic.
- Emphasis on understanding steps, not just mnemonics; SPSS/end-of-chapter sections for computer-based practice.
- Descriptive vs Inferential recap
- Descriptive: summarize a group of numbers.
- Inferential: generalize beyond the observed data to a larger population.
- Key concepts introduced early
- Variables, values, and scores; an example using a stress rating scale from 0 to 10.
- A variable is a condition that can vary; a value is a numeric or categorical descriptor; a score is a person’s specific value.
- Levels of measurement influence what statistics are appropriate.
Some Basic Concepts
- Variables, Values, and Scores (example)
- Stress level variable on a 0–10 scale; a respondent with score 6 has value 6 on the variable stress level.
- Definitions:
- Variable: a condition or characteristic that can have different values.
- Value: a possible numeric or category/or category value on a variable.
- Score: a particular person’s value on a variable.
- Table 1 (terminology)
- Variable: stress level; age; gender; religion; etc.
- Value: 0,1,2,3,… or categories like male/female.
- Score: a person’s numeric or categorical value on a variable.
- Levels of Measurement (Table 2 summary in the text)
- Equal-interval (numeric): differences between values reflect equal amounts measured; Examples: stress level, GPA; can be treated as numeric.
- Rank-order / Ordinal: numeric values reflect relative ranking, not equal intervals; Examples: class standing, place finished.
- Nominal / Categorical: values are categories with no natural order; Examples: gender, major, diagnoses.
- A variable can be equal-interval, ratio, or ordinal/nominal depending on how it is measured.
- Equal-interval vs Ratio scale
- Equal-interval: differences between values are meaningful; not necessarily a true zero.
- Ratio: has an absolute zero; allows statements about multiplicative comparisons (e.g., twice as big).
- Examples: GPA (roughly equal-interval), stress ratings (approximate equal-interval); counts like number of siblings (ratio, has true zero).
- Rank-order vs Nominal variables (discussion of information content)
- Rank-order provides relative position but less information about magnitude differences.
- Nominal provides categories with no inherent order.
- Discrete vs Continuous variables
- Discrete: specific values only (e.g., number of dentist visits; number of kids: 0,1,2,…).
- Continuous: theoretically infinite values between any two values (e.g., height, time).
- The stress example and Level of Measurement discussion
- Stress ratings on 0–10 are treated as numeric; often approximated as equal-interval.
- Distinctions between measurement scales affect the statistics that can be used.
- Practice questions (quick checks in the book)
- Identify variable, score, range, and level of measurement for various examples.
- Box 1: Poetic/historical trivia for statistics (origin, development, and early uses).
- Discrete vs Continuous and Levels of Measurement are revisited with examples and warnings about interpretation.
Frequency Tables
- An Example with stress ratings (n=151 in the full study, subset shown n=30 for teaching ease)
- Stress scores: 8,7,4,10,8,6,8,9,9,7,3,7,6,5,0,9,10,7,7,3,6,7,5,2,1,6,7,10,8,8
- A frequency table lists each possible value and how many times it occurred.
- Frequency = count of occurrences; Percent = frequency/total × 100.
- Frequency table basics
- Step 1: List all possible values from lowest to highest, including values with 0 frequency.
- Step 2: For each score, mark its occurrence on the list.
- Step 3: Compute frequency for each value by summing marks.
- Step 4: Compute percentage for each value: extPercent=NextFrequencyimes100
- Example: Stress rating table (Table 3) and its interpretation
- The table shows that most students clustered around 7–8; few around very low values.
- Frequency tables for nominal (categorical) variables
- Example: Closest person in life (208 students): Family member, Nonromantic friend, Romantic partner, Other with corresponding frequencies and percentages (Table 4).
- Another example: Social interactions diary (94 students, 10 minutes interactions)
- The Data set is a mix of numeric values (counts) that can be tabulated similarly.
- Steps for a numeric variable vs nominal variable
- Numeric: steps 1–4 as above.
- Nominal: same four-step process, but values are categories; frequencies show how many fall into each category.
- Grouped frequency tables
- Used when many possible values; group adjacent values into intervals (e.g., 0–4, 5–9, etc.).
- Interval width chosen to produce about 5–15 intervals and to have clean starting points (multiples of interval size).
- Intervals example: In stress ratings with interval size 2, intervals 0–1, 2–3, 4–5, 6–7, 8–9, 10–11.
- For 10-interval example with interval size 5: 0–4, 5–9, 10–14, …, 45–49.
- Grouped frequency table advantages and trade-offs
- Pros: simpler pictures for many values; easy visualization.
- Cons: loses some detail about frequencies within intervals.
- Histograms
- A histogram is a bar graph where bar height equals the frequency for each value (or interval midpoint for grouped tables).
- Nominal variables: bars are separated (bar graph style).
- Numeric variables: bars touch (like a skyline) to emphasize continuous distribution.
- Guidelines for building histograms from grouped data
- If using grouped data, bottom axis should be interval midpoints.
- Last interval midpoint is halfway between the start of the last interval and the start of the next one (even if that next interval doesn’t exist in data).
- Heights correspond to frequencies; bars centered on the interval midpoints.
- Figure 3–4 reference points
- Histograms comparing different tables demonstrate how grouping changes perceived distribution.
Shapes of Frequency Distributions
- Describing the shape of a distribution
- Unimodal: one clear peak in the frequency distribution.
- Bimodal: two distinct high points; often indicates two subgroups.
- Multimodal: more than two peaks.
- Rectangular (square) distribution: frequencies roughly equal across values.
- Symmetrical vs skewed distributions
- Symmetrical: left and right halves mirror each other.
- Skewed: tail on one side longer (left-skewed = negatively skewed; right-skewed = positively skewed).
- Frequency polygons
- A line-based graph connecting points corresponding to frequencies at each value; another way to visualize distributions.
- Common shapes in psychology data
- Many studies approximate unimodal distributions; bimodal/multimodal occur but are less common.
- Ceiling and floor effects contribute to skewness (e.g., many scores pile up at minimum or maximum values).
- Floor and ceiling effects (examples)
- Floor effect: many scores pile up at the lower end (e.g., number of children with low counts).
- Ceiling effect: many scores pile up at the upper end (e.g., test with max score).
- Normal curve and kurtosis
- Normal curve: bell-shaped, unimodal, symmetric; the benchmark for many statistical techniques.
- Kurtosis: degree of peakedness or flatness relative to the normal curve; high kurtosis = more scores in tails; low kurtosis = fatter/ thinner tails.
- Typical shapes in psychology distributions
- Stress and social interactions distributions are often near normal but can be skewed, reflecting measurement limits or sample characteristics.
Controversy: Misleading Graphs
- Core concern: public misuses of tables and graphs can mislead beliefs about data.
- Common misuses discussed:
- Unequal interval sizes in grouped data distort perception of trends (Figure 12 example: NYT graph with half-year data misleads); the fix is to use equal intervals.
- Exaggerating proportions by not starting the vertical axis at zero (Figure 13a vs 13b); starting at zero provides a more accurate visual impression.
- Distorting overall proportions by altering width/height relationships (Figure 14): standard 1:1.5 width-to-height ratio used for visual consistency; changing it misleads perception.
- Guidance for fair graphs
- Use equal interval sizes for grouped data.
- Start vertical axes at zero when appropriate.
- Keep aspect ratio reasonable (not wildly tall or flat) to avoid distortion.
- Public vs scholarly use
- Researchers often use frequency tables and histograms as initial steps; public plots can differ in style or formatting.
Frequency Tables and Histograms in Research Articles
- Frequency tables and histograms are valuable for understanding distributions but are not always shown in articles.
- Examples cited:
- Hether and Murphy (2010): Ten Most Common Health Issues for Male vs Female TV Characters (Table 8); shows sex-specific frequencies and percentages.
- Maggi, Hertzman, & Vaillancourt (2007): Adolescent smoking by age (Figure 15); young adolescents show rising rates with age; used grouped data with percentages.
- Maggi et al. provided a histogram for age groups; some graphs included gaps or percentages instead of raw counts to facilitate comparison.
- Maggi et al. used percentages rather than raw counts to normalize for differing group sizes.
- Reporting norms in psychology
- Mean and standard deviation are most common; mode and median appear less frequently but can be included to describe distribution shape.
- Tables (like Table 5 and small descriptive tables) may include medians when relevant, especially in triangular/skewed distributions.
- Observations about reporting practice
- In many studies, distributions are described rather than plotted; graphs appear more often in statistics-heavy papers.
- When graphs exist, the form may vary; nonstandard formats are common in applied journals.
Learning Aids, Summary, and Key Terms
- Summary points highlighted in the chapter:
- Descriptive stats summarize a distribution; central tendency focuses on a “typical” value; variability describes spread.
- Most psychological data are numeric with roughly equal intervals; rank-order and nominal data use different statistics.
- Frequency tables and histograms are foundational for understanding distributions; shapes include unimodal, bimodal, rectangular, and skewed; normal curves approximate many real-world distributions.
- Misleading graphs arise from unequal intervals, nonzero baselines, and distorted aspect ratios.
- Random vs nonrandom sampling affects generalizability; SPSS is a tool to assist in computing frequencies, histograms, and basic descriptive stats.
- Key terms (selected):
- statistics, descriptive statistics, inferential statistics, variable, value, score, numeric variable, equal-interval, ratio scale, rank-order, nominal, levels of measurement, discrete, continuous, frequency table, interval, grouped frequency table, histogram, frequency distribution, unimodal, bimodal, rectangular, symmetric, skewed, floor effect, ceiling effect, normal curve, kurtosis, central tendency, mean, mode, median, variance, standard deviation, sum of squares (SS), standard deviation, definitional formula, computational formula, population, sample, random selection, nonrandom sampling, p (probability), normal curve table (Table A-1), Z scores.
Example Worked-Out Problems (illustrative)
- Example: Ten first-year students rated interest in graduate school on a 1–6 scale: 2,4,5,5,1,3,6,3,6,6.
- Step 1: Make a frequency table for possible scores 1–6.
- Step 2: Compute frequencies: 1→1, 2→1, 3→2, 4→1, 5→2, 6→3.
- Step 3: Compute percentages: total N=10; 1:10%, 2:10%, 3:20%, 4:10%, 5:20%, 6:30%.
- Step 4: Construct a histogram from the frequency table or grouped data.
- Another worked example (mean, median, etc.)
- Stress ratings example (30 students’ scores): 193 total; mean = $$ar{X}=rac{193}{30}\