1/29
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Internal Validity: a spectrum
correlational studies are the weakest; true experiments are the strongest
even true experiments vary based on how well confounds are controlled
confounds weaken internal validity; they donât eliminate it
to strengthen internal validity, rule out as many alternative explanations as possible
âLess imperfectâ
Measurement
the process of systematically assigning values to variables (number, labels, or other symbols) in order to represent attributes of organisms, object, or events
most constructs of interest are multidimensional
Types of Measure
self-report
behavioral (observational measure)
physiological measure
Self-report
assesses respondentsâ thoughts, beliefs, and feelings
Rosenberg Self-Esteem Scale (RSES)
Behavioral (observational) measure
involves the direct observation of behavior
recording of behaviors (ex: reaction time, number of words remembered, etc)
Rater coding of behaviors (ex: number of âviolentâ acts of children in a video recording
Physiological measure
internal processes that are not directly observable (ex: eye movement, heart rate, brain activity, etc)
Levels of Measurement
Nominal - classification
nationality
Ordinal - relative standing (order to variables)
tennis ranking
Interval - equal intervals
IQ score
Ratio - true zero
# of phones
Categorical âFactorâ variables
Nominal
Ordinal
Numerical ânumericâ variables
Interval
Ratio
Level of measurement: T-Shirt (small/medium/large)
Ordinal
Level of measurement: Type of high school attended (public/private)
Nominal
Nominal scales
shuffled order makes sense
Bryn Mawr / Hav / Swat or Swat / Bryn Mawr / Hav
Ordinal scales
shuffled order hurts the meaning
there is a relative standing between levels
large/medium/small vs. small/large/medium
Level of measurement: zip code (ex: 19010)
Nominal
no zip code is âbetterâ than the other, so it is just classification
Level of measurement: Number of international trips taken in the past year
Ratio
Level of measurement: Fahrenheit
Interval
equal difference between 30 degrees F and 40 degrees F and the difference between 70 and 80
Ratio scales
For variable X that can take zero as a value, does âzero Xâ mean âno Xâ?
Zero means none (ex: âzero international trips = âno international tripsâ
Interval scales
Zero does not mean none
â0 degrees Fâ does NOT mean âno degrees Fâ
Level of measurement: Response to âI love PSYC 205â on a scale of 1 (Strongly Disagree) to 5 (Strongly Agree)
Interval
Likert Scale: measuring peopleâs attitude toward something by assessing their level of agreement with several statements about it
the interval between Strongly Disagree and Disagree is not necessarily the same as the interval between Disagree and Neither
**However, these scales are usually treated as if they have equal intervals by psychologists.
Assumption: humans respond as if they perceive equal distance between scale points
quasi interval scale
Practical Consideration: Which level of scale to use?
use the highest possible level of measurement for a given set of observations to store greatest amount of information
allows for more flexibility in later analysis
ex: individual income
could measure as ordinal (lowest to highest) or ratio (dollars/per)
ratio better
Assessing the Quality of Research
good research is reliable and valid
Research Reliability: reproducibility and replicability of findings
Research Validity: appropriateness of conclusions
Construct, statistical, external, and internal validity
Measurement Reliability: measurement should be consistent and stable
Measurement Validity: scores from a measure should represent the variable they intend to capture
**Reliability is necessary for validity but not sufficient for validity.
ex: using height as an intelligence measure (reliable but not valid)
Measurement Reliability
Measurement should be consistent and stable.
Test-Retest reliability
Internal consistency
Inter-rater reliability
Test-Retest reliability
stability over time
test correlation between time points
ex: A researcher wants to demonstrate that the RSES is a reliable measure. She asks the same participants to answer the scale twice a week apart and shows that Time 1 and Time 2 answers are correlated.
Internal consistency
consistency across items that are measuring the same construct
test correlation between items using Cronbachâs alpha
ex: A researcher wants to demonstrate that the RSES is a reliable measure. She shows that items within the scale are correlated with each other.
Inter-rater reliability
agreement among independent raters who code observed behaviors
test correlation between different raters
ex: A researcher wants to measure aggression. She hires two raters and asks them to rate aggression from a 5-minute video recording for 30 participants. Aggression rated by both raters are correlated, where participants who are rated as highly aggressive by the first rater tend to be rated highly aggressive by the second rater as well.
Measurement Validity
Scores from a measure should represent the variable they intend to capture.
Face validity
Content validity
Criterion validity
Discriminant validity
Face validity
appears to measure what it wants to measure
Content validity
comprehensive
Criterion validity
correlated with expected behaviors or other known measures (convergent validity)
ex: A researcher is developing a new measure of test anxiety. She tests the correlation between the new measure and a test score on an easy but important school exam, and finds a strong negative correlation.
Criterion validity is high.
Discriminant validity
distinctive from other conceptually distinct measures
ex: A researcher is studying self-esteem, which is a stable construct over time. She measures current mood and shows that the current mood and self-esteem are strongly correlated.
Discriminant validity is low.