1/125
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Correlational data
Data used to investigate relationships between variables. Examples include customer satisfaction surveys, political polls, government statistics, and digital behavioral data.
Examples of correlational data
Customer satisfaction ratings, political voting intentions, governmental economic statistics, social media use, shopping behavior, and information from panel surveys.
Designed data
Data deliberately collected for a specific research purpose, such as questionnaires, observations, or experiments.
Organic data
Data generated incidentally rather than specifically for research, such as social media activity, existing records, or big data.
Designed versus organic data
Designed data are deliberately generated for research; organic data arise from activities that take place independently of the research.
Three inferential goals of designed data
Description, causation, and prediction.
Description as a research goal
Describing the characteristics, behaviors, or circumstances of a population. Example: estimating the proportion of people who intend to vote for a political party.
Causation as a research goal
Investigating whether one variable causes changes in another. The lecture emphasizes that establishing causal relationships generally requires randomized experiments.
Prediction as a research goal
Using information about one or more variables to predict future events or outcomes, such as voting intentions or consumer behavior.
Can correlational research establish causation?
The lecture emphasizes that causal conclusions generally require randomized experiments. Correlational research can investigate possible causal relationships but cannot establish causation from an observed association alone.
Data collection methods in correlational research
Questionnaires, observations, and existing data such as study results, social media use, age, shopping behavior, or personal interactions.
Survey research
Research using questionnaires to measure opinions, behavior, or personal characteristics at the group level.
Questionnaires in mixed-methods research
Questionnaires can be combined with other methods, such as interviews, to obtain different types of information.
Questionnaires in experimental research
Questionnaires may be used before an experiment or after a manipulation to measure characteristics, feelings, or experiences.
Theory-data cycle
Concept/theory → research questions → research design → hypotheses → data collection → data analysis. Supporting data strengthen theories, while nonsupporting data can lead to revisions.
Survey lifecycle
A framework describing the stages through which a survey moves from theoretical concepts and target populations to responses, data adjustments, and final survey statistics.
Two dimensions of the survey lifecycle
Measurement and representation.
Measurement dimension of the survey lifecycle
Focuses on translating a theoretical concept into a measurement instrument, obtaining responses, editing those responses, and producing survey statistics.
Representation dimension of the survey lifecycle
Focuses on how the target population is represented through the sampling frame, sample, respondents, and post-survey adjustments.
Measurement versus representation
Measurement concerns whether the survey measures characteristics accurately; representation concerns whether the people included in the survey adequately represent the target population.
Steps in the measurement dimension
Theoretical concept → measurement instrument → response → edited response → survey statistics.
Steps in the representation dimension
Target population → sampling frame → sample → respondents → post-survey adjustments → survey statistics.
Theoretical concept in the survey lifecycle
The abstract characteristic the researcher wants to measure, such as depression, satisfaction, or political preferences.
Measurement instrument in the survey lifecycle
The questionnaire or other tool used to translate a theoretical concept into measurable responses.
Response in the survey lifecycle
The answer provided by a respondent to a question in the measurement instrument.
Edited response in the survey lifecycle
The response after any data editing, checking, cleaning, or processing.
Target population
The complete group of individuals about whom the researcher wants to draw conclusions, as defined in the research question.
Sampling frame
The list or source from which individuals can be selected for the sample. Ideally, it corresponds to the entire target population.
Sample
The subset of individuals selected from the sampling frame to participate in the research.
Respondents
The sampled individuals who actually participate in the survey or provide usable responses.
Post-survey adjustments
Procedures applied after data collection to improve the survey estimates or compensate for differences between respondents and the target population.
Survey statistics
The final numerical results calculated from survey responses after relevant processing and adjustments.
Total Survey Error Framework (TSE)
A framework identifying the different sources of error that can arise throughout a survey and affect the accuracy of its results.
Six errors in the Total Survey Error Framework
Coverage error, sampling error, nonresponse error, adjustment error, measurement error, and processing error.
Which four TSE errors concern representation?
Coverage error, sampling error, nonresponse error, and adjustment error.
Which two TSE errors concern measurement?
Measurement error and processing error.
Coverage error
Error occurring when the sampling frame does not accurately correspond to the target population, meaning some eligible individuals are missing or ineligible individuals are included.
Population versus sampling frame
The population contains everyone the study aims to investigate; the sampling frame is the actual list or source used to select respondents.
Coverage error: Utrecht students example
The population includes all students living in the province of Utrecht, but the sampling frame may contain students studying in Utrecht who live elsewhere or people who are no longer students.
Undercoverage
Occurs when some members of the target population are missing from the sampling frame and therefore cannot be selected.
Overcoverage
Occurs when the sampling frame includes individuals who do not belong to the target population.
Undercoverage versus overcoverage
Undercoverage means eligible population members are excluded; overcoverage means ineligible individuals are included in the sampling frame.
When does coverage error create bias?
When excluded individuals systematically differ from those included in the sampling frame on characteristics relevant to the study.
Random versus systematic coverage error
Random omissions do not necessarily produce systematic bias, whereas systematic differences between covered and uncovered individuals can bias the results.
Bias
A systematic difference between an estimate or measurement and the true value it is intended to represent.
Systematic error
Error that pushes results away from the truth in a consistent or nonrandom direction, potentially producing bias.
Random error
Unpredictable variation that causes estimates or measurements to differ from their true values without necessarily creating a consistent directional bias.
Why are systematic errors especially problematic?
They can consistently overestimate or underestimate a characteristic, meaning that results may remain biased even when many observations are collected.
A survey of Utrecht students excludes everyone living in student housing. What error might occur?
Coverage error through undercoverage. Bias may occur if students in housing differ systematically from included students.
A sampling list contains graduates who are no longer students. What problem is this?
Overcoverage, because the sampling frame includes people outside the intended population.
Sampling error
The difference between a statistic calculated from a sample and the corresponding true population parameter, arising because only part of the population is observed.
Sampling error in correlation
The difference between the sample correlation and population correlation, expressed as r − ρ.
What does r represent?
The correlation coefficient calculated from the sample.
What does ρ (rho) represent?
The true correlation coefficient in the population.
Why does sampling error occur?
Different samples contain different individuals and therefore may produce different statistics, even when they come from the same population.
Can the exact sampling error be known from one sample?
Generally no, because the true population parameter is usually unknown.
Standard error
A measure of how much a statistic varies across repeated samples, used to describe the typical size of sampling variation.
Sampling error versus standard error
Sampling error is the difference for one particular sample; standard error measures the typical variation of the statistic across repeated samples.
A sample correlation is 0.35 and the population correlation is 0.20. What is the sampling error?
r − ρ = 0.35 − 0.20 = +0.15.
Why does increasing sample size generally reduce sampling error?
Larger samples generally produce more precise estimates, reducing the typical sampling variation as measured by the standard error.
Nonresponse error
Error arising when selected individuals do not provide survey responses or omit particular questions.
Unit nonresponse
Occurs when a selected individual does not participate in the survey at all.
Item nonresponse
Occurs when an individual participates in a survey but does not answer a particular question.
Unit nonresponse versus item nonresponse
Unit nonresponse concerns failure to participate in the entire survey; item nonresponse concerns unanswered individual questions.
Reasons for nonresponse
Technical difficulties, lack of motivation or interest, and lack of trust, particularly when sensitive questions are asked.
Nonresponse bias
Bias occurring when nonrespondents differ systematically from respondents on characteristics relevant to the survey results.
When does nonresponse create bias?
When individuals who refuse the survey or skip a question systematically differ from individuals who respond.
Does every instance of nonresponse lead to bias?
No. The lecture emphasizes that bias arises from systematic differences between respondents and nonrespondents, not simply from a random number of people failing to respond.
A student ignores an entire questionnaire invitation. What type of nonresponse?
Unit nonresponse.
A student completes a questionnaire but skips a question about alcohol consumption. What type of nonresponse?
Item nonresponse.
Students with very low satisfaction are less likely to complete a satisfaction survey. What error may arise?
Nonresponse bias, because respondents and nonrespondents systematically differ in satisfaction.
Coverage bias versus nonresponse bias
Coverage bias arises when relevant people are missing from the sampling frame; nonresponse bias arises when selected people fail to respond and differ systematically from respondents.
Adjustment error
Error introduced by post-survey adjustments used to modify survey estimates or compensate for representation problems.
Example of adjustment error
If survey weights or other corrections are inaccurate, the adjusted results may become distorted rather than accurately represent the population.
Measurement error
Also called response error. It occurs when the recorded survey answer differs from the true value of the characteristic being measured.
Response error
Another name for measurement error: inaccurate answers produced by survey format, question wording, questionnaire design, or respondent behavior.
Three broad causes of measurement error
Mode effects, question bias, and response bias.
Mode effect
A change in survey responses caused by the questionnaire administration format rather than a genuine difference in the underlying characteristic.
Question bias
Measurement bias caused by the formulation, content, or ordering of questions.
Response bias
Measurement bias caused by respondents' tendencies or behavior when answering questions.
Mode effect versus question bias versus response bias
Mode effects result from administration format; question bias results from how questions are designed; response bias results from how respondents behave when answering.
Processing error
Error introduced while recording, coding, editing, transforming, or otherwise processing survey responses.
Example of processing error
A researcher incorrectly enters a response of 4 as 1 or accidentally reverse-codes the wrong questionnaire item.
Measurement error versus processing error
Measurement error arises when a response inaccurately represents the intended characteristic; processing error arises when the collected response is incorrectly handled afterward.
Tourangeau's response process
A model describing the mental stages respondents pass through when answering survey questions.
Four stages of Tourangeau's response process
Comprehension, retrieval, judgment, and response.
Comprehension
The stage in which respondents interpret the question and understand the meaning of its words and concepts.
Retrieval
The stage in which respondents search their memory for the information needed to answer the question.
Judgment
The stage in which respondents evaluate, combine, or estimate the retrieved information to form an answer.
Response
The stage in which respondents translate their judgment into one of the available answers and decide what to report.
Tourangeau's response process: order
Comprehension → retrieval → judgment → response.
Why is Tourangeau's response process important?
Errors can occur at any stage, meaning that apparently simple survey questions may produce inaccurate answers even when respondents intend to answer correctly.
Tourangeau alcohol-consumption example
A survey asks, “How many glasses of alcohol did you drink last week?” Respondents must understand the terminology, remember their drinking, estimate the number accurately, and be willing to report it.
Comprehension problem in the alcohol example
The respondent may interpret “one glass” or “last week” differently from what the researcher intended.
Retrieval problem in the alcohol example
The respondent may not remember exactly how many alcoholic drinks they consumed during the previous week.
Judgment problem in the alcohol example
The respondent may remember different occasions but make an inaccurate estimate when combining the information.
Response problem in the alcohol example
The respondent may know the answer but choose not to report it accurately because they feel embarrassed or want to appear socially responsible.
Question formulation bias
Bias resulting from how a survey question is worded, such as using difficult words, double negatives, or complicated sentence structures.
Jargon or difficult words
Specialized language that respondents may misunderstand, creating errors during comprehension.
Double negatives
Questions containing multiple negations that make their intended meaning difficult to interpret.