Questionnaire Design 2: Comprehensive Study Notes (UNT)
Developing Good Research Instruments
- Each question must clearly communicate to the respondent what the researcher wants to know.
- Each answer must clearly communicate to the researcher what the respondent wants to say.
- Bad questions yield bad data and increase (unnecessary) costs.
- Key goals: developing reliable and valid research instruments that elicit accurate responses at reasonable cost.
Question wording and respondent capacity
- Question wording is critical for expressing meaning and intent.
- Must understand respondents' intellectual capacity and language ability.
- Consider education level (kids vs. adults) and whether respondents are native speakers.
- Use appropriate language in both questions and answers; avoid ambiguous terms; keep vocabulary simple.
- Simple, straightforward sentence construction improves understanding.
- Avoid double-barreled questions and difficult vocabulary.
Examples and linguistic considerations
- Parking question variants illustrate how wording changes perception:
- "Do you think anything could be done to make it more convenient for students to find parking?"
- "Do you think anything should be done to make it more convenient for students to find parking?"
- "Do you think anything will be done to make it more convenient for students to find parking?"
- Use objective, quantifiable response options when possible (everyday, twice a week, once a week, etc.).
- Clarify terms like "many" in questions to avoid vagueness (e.g., replace with a concrete number).
Concrete guidelines
- Use simple vocabulary; avoid loaded terms; avoid jargon.
- Ensure response categories are mutually exclusive.
- Ensure questions are meaningful and easily answerable.
- Avoid ambiguous terms and ensure wording communicates exactly what is measured.
Avoiding common pitfalls
- Avoid unanswerable questions where respondents lack information or no option applies (e.g., income ranges that don’t fit respondents).
- Avoid loaded questions that push toward a socially desirable answer or that involve emotionally charged issues.
- Avoid double-barreled questions that mix multiple issues in one item.
- Avoid response bias from order effects by randomizing option order when feasible.
Double-barreled questions (examples)
- Example 1 (owning Apple products):
- "Do you feel that owning Apple products has made you more creative and successful?"
- Variants show different degrees of concreteness and potential bias.
- Example 2 (orange juice):
- "Do you drink orange juice for breakfast, lunch, and dinner?" vs. individual meal variants.
- Takeaway: ensure each item targets a single construct only.
Unanswerable and loaded questions
- Unanswerable: when respondents cannot provide a meaningful answer due to lack of information or non-applicable options.
- Loaded: questions that prompt a socially desirable response or hinge on emotional or political framing (e.g., taxes, immigration, sensitive topics).
- Examples show how to structure questions to minimize bias and capture genuine attitudes.
Sensitive questions
- Some domains (income, sexual beliefs/behaviors, medical conditions, alcohol use) see lower honesty due to discomfort.
- Respondents may give socially desirable answers or refuse to answer.
- Strategies: place sensitive questions toward the end; reassure confidentiality; phrase questions about behavior of others to reduce self-consciousness while still predicting respondent behavior.
- Examples illustrate differential phrasing to reduce defensiveness.
Screening and layout in questionnaire design
- Screening questions ensure you talk to the right target population (e.g., recent car buyers, iPad owners).
- Screening is common in computer-assisted surveys (e.g., Qualtrics).
- Introductory section provides research overview and legitimizes the study; assure anonymity/confidentiality.
- The research questions should flow from general to specific questions to avoid priming and bias.
- Start with broad questions about preferences, then move to specifics (e.g., overall liking before camera features on an iPhone).
- Response order bias: randomize options when possible to prevent central tendency or anchor effects.
- Sensitive questions should be placed late; demographic questions at the end to maintain comfort.
- General flow: Intro → Screener → Research questions → General to specific questions → Sensitive questions → Demographics → Thank-you.
Introduction and screening details
- Introduction should certify legitimacy, identify involved organizations, and guarantee anonymity/confidentiality.
- Screening questions help ensure only qualified respondents participate; can be crucial for computer-assisted surveys.
- Checklist examples include ownership or usage questions (e.g., Do you own an iPad? Have you eaten at Eagle Landing in the past month?).
Question sequencing and transitions
- Begin with general questions; progress to more specific ones.
- Use warm-up questions to ease respondents into the survey; transition questions link sections to the research objectives.
- Sensitive questions should be placed after respondents have built comfort with the survey.
- Demographic questions appear at the end.
Designing good questionnaires: practical layout rules
- Think like a respondent and check clarity and answerability of questions.
- Use a logical order and a general-to-specific flow to reduce sequence bias.
- Consider adding skip patterns to avoid irrelevant questions for some respondents.
- Ensure questionnaire looks attractive and professional; decide on incentive use with the client.
The questionnaire design process (7 steps)
- Confirm research objectives.
- Select appropriate data collection method.
- Develop questions and scaling.
- Determine layout of the questionnaire.
- Obtain client approval.
- Pretest, revise, and finalize.
- Implement survey.
Questionnaire layout and client approval
- Copies of the questionnaire should be shared with all project stakeholders for feedback.
- The questionnaire should be attractive and professional; incentives can be used as agreed.
- Client approval is a formal step before fielding.
Pretesting and piloting the questionnaire
- Pretest with a small sample (10–20 participants) to identify difficulties or questions.
- Check for signs of boredom (e.g., skipped questions) and estimate completion time.
- Use pretests to revise the questionnaire before full deployment.
Data collection and implementation
- The data collection process varies by modality (online, mail, telephone).
- Ensure data collection procedures align with the study design and client agreements; incentives, if any, should be used as agreed.
Scale measurement and questionnaire integration
- Scale measurement: the process of assigning descriptors to represent the range of possible responses to a question about a construct.
- A questionnaire includes a set of scales and questions; each scale measures one construct.
- Scales measure both the independent variable(s) and dependent variable(s).
- The questionnaire should include measures (e.g., scales) for both the independent and dependent variables.
Reliability and validity: a quick framework
- Reliability: the extent to which a scale produces consistent results across repeated measurements.
- Validity: the extent to which a scale measures what it is supposed to measure.
- Not necessarily correlated: high reliability does not guarantee high validity, and high validity does not guarantee high reliability.
- Core techniques for reliability:
- Test-retest reliability
- Equivalent form reliability
- Internal consistency reliability (e.g., Cronbach’s alpha)
- Core validity types include face validity, content validity, convergent validity, and discriminant validity.
Reliability: key approaches
- Test-retest reliability: administer the same scale to the same respondents at two different times under similar conditions; similar results indicate stability.
- Equivalent form reliability: create two different but equivalent versions of the same scale (Version A vs Version B) and administer to the same or similar samples; high similarity indicates reliability.
- Internal consistency reliability: assess how well the items on a scale hang together to measure the same construct.
- Split-half method: compare scores from two halves of the scale (e.g., odd vs. even items or random halves); high correlation indicates internal consistency.
- Coefficient alpha (Cronbach’s alpha): the most common metric for multi-item scales; ranges from 0 to 1; higher is better up to a point.
Split-half example
- Example scale (customer satisfaction): items rated on a 1–10 scale (Strongly Disagree to Strongly Agree):
- Example items include statements about overall satisfaction, perceived value, delivery quality, defect absence, staff courtesy, care for business, facilities, website usefulness, product information, likelihood to recommend.
- A two-column presentation can illustrate two halves of the scale with the same item structure.
Cronbach’s Alpha and interpretation
- Cronbach’s alpha (a) is a measure of internal consistency reliability.
- Acceptable reliability: exta≥0.70
- Higher categories (from slides):
- Excellent: a≥0.90
- Good: 0.90 > a \ge 0.80
- Acceptable: 0.80 > a \ge 0.70
- Questionable: 0.70 > a \ge 0.60
- Poor: 0.60 > a \ge 0.50
- Unacceptable: a < 0.50
- When Cronbach’s alpha is too high (e.g., a > 0.95), it may indicate redundancy among items.
Validity: what it means to measure the right thing
- Validity assesses whether a scale measures the intended construct.
- Face validity: is the indicator intuitively reasonable as a measure of the construct? Based on expert judgement.
- Content validity: the items cover all relevant dimensions of the construct (e.g., satisfaction with a restaurant should include food quality, courtesy, wait time, ambience).
- Convergent validity: measures that should be related are indeed related (e.g., related forms of a mood scale correlate).
- Discriminant validity: measures that should not be related are indeed not related (e.g., brand authenticity vs. brand personality should diverge).
Reliability vs validity: key relationship
- High reliability does not guarantee high validity; high validity does not guarantee high reliability.
- Both are essential for sound measurement.
Other considerations in scale design
- Single-item vs. multiple-item scales:
- Single-item scale: describes only one attribute (e.g., Age).
- Multiple-item scale: covers multiple attributes of a construct (e.g., Brand loyalty: purchase frequency, brand liking, brand recommendation, amount spent).
- Many marketing measures use multi-item scales because they capture more dimensions.
- Discriminatory power of scale descriptors refers to how well scale points differentiate responses.
- Balanced vs. unbalanced scales:
- Balanced: equal number of positive and negative items.
- Unbalanced: more items on one side; can introduce bias unless attitudes are genuinely one-sided (e.g., tax increases).
- Forced vs. non-forced scales:
- Forced: no neutral option; pushes choice toward positive or negative.
- Non-forced: includes a neutral option or middle point (e.g., N/A).
- Unstructured questions vs. structured scales:
- Structured (close-ended): quick to answer and analyze; preferred for efficiency.
- Unstructured (open-ended): richer data but time-consuming to analyze; respondents may skip.
Discriminatory power and number of scale points
- More scale points increase variability and discriminatory power but may complicate responses.
- Common practice in marketing: use 5-point, 7-point, or 11-point scales; debate on extremely high scales (e.g., 500 points) is generally discouraged due to respondent fatigue.
- Example scale point discussion:
- Balanced, 5-point or 7-point scales are typical; central values tend to be chosen by respondents to save effort.
Practical questionnaire design rules
- Place sensitive questions toward the end; place demographic questions at the end.
- Ensure questions are clear, answerable, and aligned with research objectives.
- Use skip patterns to avoid irrelevant questions for some respondents.
- Consider introducing the survey with a legitimacy statement and confidentiality assurances.
The big picture: questionnaire design workflow
- Start with a clear objective and map questions to constructs/measures.
- Choose data collection method and sampling frame with the client.
- Draft questions and scaling; design the layout to guide respondents logically.
- Get client approval; pretest the questionnaire with a small sample.
- Revise based on feedback and pretest results.
- Implement the survey and collect data; monitor for issues during data collection.
Key takeaways for exam readiness
- Always link each question to a single construct; avoid ambiguity and multi-issue items.
- Use appropriate language and avoid biasing phrasing.
- Structure surveys to minimize biases: general-to-specific flow, randomize options when feasible, and place sensitive items late.
- Distinguish reliability (consistency) from validity (accuracy); use tests such as test-retest, equivalent form, and internal consistency (Split-half, Cronbach’s alpha) to evaluate scales.
- Be mindful of scale design choices: number of items, balance, forcing of choices, and mode of data collection.
- Use content, face, convergent, and discriminant validity to ensure constructs are measured as intended.
Quick reference: glossary of terms from the slides
- Reliability: consistency of a measure across time or items.
- Validity: accuracy of a measure in capturing the intended construct.
- Face validity: expert judgment on whether a measure appears to assess the intended construct.
- Content validity: coverage of all relevant aspects of the construct.
- Convergent validity: related measures converge.
- Discriminant validity: unrelated measures diverge.
- Cronbach’s alpha: internal consistency reliability metric for multi-item scales; typically interpreted with thresholds (Excellent, Good, Acceptable, Questionable, Poor, Unacceptable) as discussed above.
- Test-retest reliability: consistency over time.
- Equivalent form reliability: consistency across alternative forms of a scale.
- Split-half reliability: internal consistency via correlating halves of a scale.
- Likert-type scales: common multi-item response formats (e.g., 5-, 7-, or 11-point scales).
- Double-barreled question: item that asks about two issues simultaneously.
- Loaded question: item that presupposes a socially desirable response.
- Demographic questions: usually placed at the end of the survey.
- Pretest: pilot run with a small sample to identify problems before full deployment.
- Skip patterns: logic to bypass irrelevant questions based on earlier responses.