1/78
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
STATISTICS
- refers to a range of techniques and procedures for analyzing, interpreting, displaying, and making decisions based on data.
- Numerical facts and figures
- Involves math and relies upon
calculations of numbers
- Relies heavily on how the numbers are
chosen and how the statistics are
interpreted
why study it?
- To organize massive amount of information into a more objective interpretable form
- To properly evaluate the data and claims that bombard you everyday
- To communicate results and research conclusions
- To learn to recognize statistical evidence that supports a stated conclusion

Descriptive
- provide ways of summarizing the information that we collect from a multitude of sources

Inferential
- confidence in which we can generalize from a sample to the entire population

Data Simplification/Data exploration/Data reduction
- to make sense of large amounts of data that otherwise would be too much confusing

VARIABLE
- Simply a characteristic or feature of the thing we are interested in understanding
- Any concept that we can measure and that varies between individuals or cases

INDEPENDENT
variable is manipulated by an experimenter; cause

DEPENDENT
effect in variable caused by the manipulation on IV

QUALITATIVE
- express a qualitative attribute; values of qualitative v do not imply a numerical ordering
- "Categorical"

QUANTITATIVE
- variables measured in terms of numbers
- Discrete and continuous

DISCRETE
possible scores are discrete points on the scale; countable

CONTINUOUS
scale is continuous; infinite

OBSERVABLE
can be directly measured or observed.

LATENT
not directly observed but are inferred from observable variables

Mediator
Explains how or why two variables are related.
Think of it as: The "middle step" in the
process.
Mechanism (how or why)

Moderator
- Changes the strength or direction of the relationship between two variables.
Think of it as: A "switch" or "volume knob" that makes the relationship stronger, weaker, or different
depending on its level.
- Modifier (when or for whom)

level of measurement
- determines what kinds of statistics are meaningful and valid.
- Using the wrong statistic can lead to
misleading conclusions.

(TRUE) EXPERIMENTAL DESIGN
- The use of random assignment to treatment conditions and manipulation of the independent variable

Random sampling
randomly chosen as samples

Random Assignment
sample is randomly assigned to a certain condition

QUASI-EXPERIMENTAL DESIGN
- Manipulating the IV but not randomly assigning people to groups
Why use this?
- It may be unethical to deny potential treatment to someone if there is good reason to believe it will be effective and that the person would unduly
suffer if they did not receive it
- It may be impossible to randomly assign people to groups

NON-EXPERIMENTAL DESIGN
- Correlational research
- Observing things as they occur naturally and recording our observations as data
- Reflects reality as it actually exists since we as researchers do not change anything
- Becomes a predictor

DESCRIPTIVE ANALYSIS
- Numbers that are used to summarize and describe data
- Just descriptive; they do not involve generalizing beyond the data at hand
- It is important to differentiate what we use to describe populations vs. samples

Population
- is described by a parameter
- the true value of the descriptive in the population, but one that we can never know for sure

sample statistic
- refers to the specific number we compute from the data (e.g. average)
- an estimate of the true population parameter, and if our sample is representative of the population, then the statistic is considered to be a good estimator of the parameter.

Sampling Error
discrepancy/difference between the parameter and the statistic we use to estimate it.

INFERENTIAL STATISTICS
- shows how our data behaves
- how we generalize from our sample back up to our population
- Correlational, comparative, and predictive

Graphing Qualitative Variables
- No pre-established ordering
- refers to using visual tools to present and describe categorical data that has no pre-established or numerical ordering.
RESEARCH DESIGN ISSUE: allowing participants in your research to give more than one answer to a question.
• Statistics in general do not handle multiple responses very well.
• The totals in a table exceed the number of participants in the research

FREQUENCY TABLES
- shows the frequencies of the various response categories
- also shows the relative frequencies, which are the proportion of responses in each category

PIE CHARTS
• Each category is represented by a slice of the pie.
• The area of the slice is proportional to the percentage of responses in the category.
• simply the relative frequency multiplied by 100
• effective for displaying the relative frequencies of a small number of categories
DON'TS:
- Too many small slices are identified by different shading patterns, and the legend takes time to decode.
- Don't use it to compare the outcomes of two different surveys or experiments
- Don't label the ___ with percentages if they are based on a small number of percentages

BAR CHARTS
- can also be used to represent frequencies of different categories
- have a standard space separating them--spaces indicate that the categories are not in a numerical order; they are frequencies of categories, not scores.
THINGS TO REMEMBER:
- The heights of the bars represent frequencies (number of cases)in a category. q Each bar should be clearly labeled as to the category it represents. - Too many bars make bar charts hard to follow.
- Avoid having many empty or near-empty categories, which represent very few cases.
- Nevertheless, if important categories have very few entries, then this needs recording.
- Make sure that the vertical axis (the heights of the bars) is clearly marked as being frequencies or percentage frequencies.
- The bars should be of equal width
- Don't set the baseline to a value other than zero!

Y-axis
shows the number of observations in each category.

COMPARING DISTRIBUTIONS
To compare the results of different surveys, or of different conditions within the same overall survey

horizontal format
- It is useful when you have many categories because there is more room for the category labels.
simplified explanation: turning that entire chart sideways! Instead of climbing up toward the sky, the bars grow sideways from left to right like cars zooming down a racetrack.

pictogram
It is a type of chart that uses meaningful images or little drawings instead of normal colored bars to show amounts.
For example, instead of a plain blue bar showing how many apples were sold, a pictogram might stack up drawings of actual apples.

Line graph
bar graph with the tops of the bars represented by points joined by lines (the rest of the bar is suppressed)

Descriptive statistics
- They are, by and large, relatively simple visual and numerical techniques for describing the major features of one's data
simplified explanation: It is a set of simple visual techniques (like drawing a colorful chart) and numerical techniques (like finding the average) that help you describe the major features of your massive pile of information.

QUANTITATIVE VARIABLE
- Variables measured on a numeric scale
Example: height, weight, response time, subjective rating of pain, temperature, score on an exam

STEM DISPLAYS
arranged as a column to the left of the bars; represents the tens digits (e.g. 3 stem = 30 to 39)

LEAF DISPLAYS
- Numbers to the right represent the ones digits
- Every leaf in the graph, therefore, stands for the result of adding the leaf to 10 times its stem.
Purpose: to clarify the shape of the distribution; the precise numbers can be determined by examining the leaves

STEM AND LEAF DISPLAYS
- Best-suited for small to moderate amounts of data (up to 200 observations)
- We can make our figure even more revealing by splitting each stem into 2 parts
- Splitting depends on the exact form of your data: if rows get too long with single stems, you can try splitting them into 2 or more parts
- they are placed back to back along a common column of stems: "Back-to-back stem and leaf display"
- Easy to graph when:
- data are whole numbers
- All numbers are positive
- Data with decimals can be rounded to 2-digit accuracy

Test Anxiety Questionnaire
(Scores range from 10 - 50)
Low: 10 - 19
Moderate: 20 - 35
High: over 35

HISTOGRAMS
- Used for displaying the shape of the distribution of a quantitative variable.
- It groups data into class intervals (bins) and represents the frequency (or relative frequency) of data within each interval using adjacent bars.
- Best-suited for large amounts of data
- can also be used when the scores
are measured on a more continuous scale, such as the length of time (in milliseconds) required to perform a task.
When to Use It:
- Best-suited for large data sets (typically more than 20-30 observations).
- Useful for identifying distribution shape, such as normality, skewness, or multimodality.
- Can be used for both discrete and continuous quantitative variables.

class frequencies
- In a histogram, are represented by bars.
- Height of each bar corresponds to its class frequency.

FREQUENCY POLYGONS
- Same purpose as histograms, but are especially helpful for comparing two or more distributions.
- Also a good choice for displaying cumulative frequency distributions (or trends in grouped data)
- are useful for comparing distributions
- Achieved by overlaying the frequency polygons drawn for different data sets
- Also possible to plot two cumulative frequency distributions in the same graph.
- Shows data distribution
- X-axis: class midpoints,
- Y-axis: frequencies
- Always uses straight line segments
- Comparing grouped frequency data

Cumulative Frequency
- Y-axis values represent not the individual class frequency
- Helps to visually identify medians, quartiles, and percentiles.
- The final point always shows the total number of observations.

PERCENTILES
- are based on cumulative frequency, indicating the value below which a given percentage of data falls.
- indicates the score below which a given percentage of scores fall.
- Commonly used in standardization tables of psychological tests and measures
- Describe a person's standing compared with the set of individuals on which the test or measure was initially researched
- Quick method of expressing a person's score relative to those of others.
Example:
- 90th percentile = the score is greater than or equal to 90% of the distribution.
- Neuroticism score is 90th percentile = person is more neurotic than about 90% of people

BOX PLOTS
- Box-and-Whiskers Plot
- Provide a visual summary of a distribution's central tendency and spread.
- Highlight the median, quartiles, and potential outliers.

positive skew
Longer whisker on the top
negative skew
Longer whisker on the bottom
positively skewed
If mean > median
LINE GRAPH
- Are appropriate only when both the X- and Y-axes display ordered (not qualitative) variables
- Generally better than bar charts when comparing changes over time
- Shows trends over time
- Both axes are ordered variables
- Can be smooth or segmented
- Time-series data

SHAPE OF A DISTRIBUTION
- Guides the choice of analysis.
- The overall pattern of data when plotted (e.g., histogram).
Why it matters in psychology:
- Guides choice of statistical measures.
- Helps detect unusual patterns (e.g., extreme anxiety scores).
- Influences the validity of statistical tests.

Symmetry
- Are both sides mirror images?
- Skewness = 0

Skewness
- Is one tail longer than the other?
- The degree of asymmetry in a distribution.
- It affects the placement of the mean, median, and mode.
- reveal data characteristics.

Kurtosis
- How peaked or flat is the curve?
- reveal data characteristics.
- Describes "peakedness" and "tailedness" of a distribution.

SYMMETRICAL DISTRIBUTIONS
- A distribution where left and right halves are mirror images.
- Rare in real-world psychological data but many are approximately symmetrical.
Psychology Example:
IQ scores in the general population are approximately symmetrical, making the mean, median, and mode nearly equal.

Normal distribution
is the classic example.

BIMODAL DISTRIBUTIONS
- A distribution with two distinct peaks.
Why it matters:
- Indicates the presence of two subgroups or underlying processes.
- Average (mean) may not represent either group well.
- Potential for overlooking important insights.
- Requires a much more sophisticated measure.
Psychology Example:
Stress levels among hospital staff may show one peak for nurses and another for doctors.

Positive skew
→ tail to the right.
- Tail on the right is longer.
- Mean > Median > Mode.
- Often caused by extreme high values.
- Skewness > +0.4
Psychology Example:
Therapy wait times — most clients start quickly, but a few wait months.

Negative skew
→ tail to the left.
- Tail on the left is longer.
- Mean < Median < Mode.
- Often caused by extreme low values.
- Skewness < -0.4
Psychology Example:
Memory recall scores in older adults — most score high, but a few with impairments pull the tail left.

Approximately Symmetrical
−0.4 ≤ Skewness ≤ 0.4
Excess kurtosis
(common in software)
Leptokurtic
- (steep curve)
- Tall peak, heavy tails (more outliers).
- Higher risk of extreme values.
- Kurtosis > 0
Example:
Test anxiety scores in a highly competitive school — most cluster tightly, but some extreme cases exist.

Mesokurtic
- (normal curve)
- Moderate peak, moderate tails.
-Shape similar to normal distribution.
- Kurtosis = 0
Example:
Height distribution of adult males.

Platykurtic
- (flat curve)
- Flat peak, light tails.
- Fewer extreme values.
- Kurtosis < 0
Example:
Satisfaction survey where responses are more evenly spread.

CENTRAL TENDENCY
- refers to the statistical measures that identify the central point or typical value of a dataset.
- It summarizes data by identifying a representative score around which other values cluster.

Mean
- pulled toward tail in skewed data.
- Is the average of a dataset.
- Sum of all of the scores in the distribution divided by the number of scores
Median
- more stable in skewed data.
- The middle score of a set if the scores are organized from the smallest to the largest.
- When there is an odd number of scores, it is still simply the middle number.
- When even numbers, it is still the mean of the two middle scores.
- When there are numbers with the same values, each appearance of that value gets counted.
Mode
- The most frequently occurring value in the dataset
- stays at the peak.
- Only measure that we can use on qualitative or categorical data as well as numerical score data
- A dataset can have one mode, more than one mode, or no mode at all.
Bimodal or multimodal distribution
- several modes
Outliers
- can represent rare but significant cases or errors.
- must be handled carefully in psychology.

Extreme outliers
are identified in much the same way but the interquartile range is multiplied by 3 (rather than 1.5)

SPREAD/VARIABILITY/DISPERSION
- How "spread out" or clustered a group of scores is?
- are developed which include the extent to which each of the scores in the set differs from the mean score of the set.
measures of variability
describe how scores in a given dataset differ from one another

VARIANCE
Calculated like the mean deviation, but we square each deviation from the mean before summing the total of these squared

STANDARD DEVIATION
To get back to the original units (e.g., seconds, scores), you take the square root

Low variance
→ most participants have very similar reaction times.
High variance
→ some are very fast, some are very slow
→ might indicate different strategies, attention issues, or outliers.
Estimated variance
is your best guess of the variance of the population if you only have the data from a small set of scores.