Statistics: Bias and Sampling Flashcards
Conceptual Foundations and Definitions of Bias
Bias is introduced into a statistical study when the methodology employed systematically favors certain outcomes or results over others.
The presence of bias suggests that the research method is highly likely to either underestimate or overestimate the actual value of the parameter intended for study.
In the context of Free Response Questions (FRQ), it is imperative to always select a specific direction for the bias (underestimation or overestimation) and provide a corresponding logical reason. This requirement must be fulfilled regardless of whether the question explicitly asks for "DIRE".
Bias Arising from Suboptimal Sampling Methods
Utilizing poor sampling techniques, particularly convenience sampling or voluntary response sampling, frequently results in statistical bias.
Voluntary Response Bias
Voluntary response bias occurs when individuals self-select to participate in a study or survey.
This methodology is considered flawed because it attracts participants with strong feelings about a specific issue; these individuals often share the same viewpoint, which is typically negative.
Because the respondents are self-selected, they are generally not representative of the broader population under study.
This form of bias is easily avoidable through the implementation of random sampling. Consequently, some instructional materials do not even categorize it as a primary type of bias because it is considered a complete failure of random selection principles.
Question Wording Bias
This form of bias occurs when survey questions are phrased in a confusing manner or are designed to lead the respondent toward a specific answer.
Designers must be cautious of words that influence perception. For example, asking "Do you think it's OK for toddlers to eat unhealthy candy?" is biased because it uses the leading term "unhealthy" to influence the respondent's judgment.
Specific Biases Occurring Post-Random Selection
Even when random selection is correctly implemented, several types of bias can still manifest during the data collection process.
Undercoverage
Undercoverage occurs when certain members of the population have a lower probability of being selected in a sample than others.
The process of sampling often involves a "sampling frame," which is an exhaustive list of individuals within the population. However, such lists are rarely completely accurate or exhaustive.
Performing random sampling using an incomplete or unrepresentative sampling frame inevitably results in undercoverage.
Specific examples of populations frequently omitted in traditional household sampling frames include:
The homeless population.
Prison inmates.
Students who reside in dormitories.
Nonresponse
Nonresponse bias arises when an individual who has been selected for a sample cannot be reached or explicitly refuses to participate in the study.
For this to be a source of bias, there is a functional assumption that the individuals who do not participate would have provided different responses compared to those who did participate.
There is a strict temporal distinction for nonresponse: it can only occur after a sample has been formally selected.
In contrast, voluntary response samples do not experience nonresponse because the participants choose to participate from the start rather than being selected from a frame.
Common practical examples of nonresponse include:
Surveyors attempting to call landline telephone numbers while residents are not at home.
Survey questionnaires sent via postal mail that are lost or ignored by the recipient.
Individuals refusing to answer telephone calls from unrecognized or unknown numbers.
Response Bias and Participant Inaccuracy
Response bias occurs when there is a systematic pattern of inaccurate answers provided to survey questions. This indicates that the responses provided do not align with the factual reality.
There are several common drivers for response bias, including:
Intentional Dishonesty: Participants may lie about personal data, such as their age or annual income.
Sensitive Topics: Respondents are more likely to provide false information regarding sensitive or stigmatized behaviors, such as the use of drugs or alcohol.
Lack of Comprehension: Individuals may fabricate answers to questions they do not understand rather than admitting they are confused.
Interviewer Influence: The behavior, appearance, or presence of the interviewer can inadvertently influence a person's responses, leading them to give answers they feel are more socially acceptable or expected.