Questionnaire Design 2: Comprehensive Study Notes (UNT)

Developing Good Research Instruments

  • Each question must clearly communicate to the respondent what the researcher wants to know.
  • Each answer must clearly communicate to the researcher what the respondent wants to say.
  • Bad questions yield bad data and increase (unnecessary) costs.
  • Key goals: developing reliable and valid research instruments that elicit accurate responses at reasonable cost.

Question wording and respondent capacity

  • Question wording is critical for expressing meaning and intent.
  • Must understand respondents' intellectual capacity and language ability.
  • Consider education level (kids vs. adults) and whether respondents are native speakers.
  • Use appropriate language in both questions and answers; avoid ambiguous terms; keep vocabulary simple.
  • Simple, straightforward sentence construction improves understanding.
  • Avoid double-barreled questions and difficult vocabulary.
Examples and linguistic considerations
  • Parking question variants illustrate how wording changes perception:
    • "Do you think anything could be done to make it more convenient for students to find parking?"
    • "Do you think anything should be done to make it more convenient for students to find parking?"
    • "Do you think anything will be done to make it more convenient for students to find parking?"
  • Use objective, quantifiable response options when possible (everyday, twice a week, once a week, etc.).
  • Clarify terms like "many" in questions to avoid vagueness (e.g., replace with a concrete number).
Concrete guidelines
  • Use simple vocabulary; avoid loaded terms; avoid jargon.
  • Ensure response categories are mutually exclusive.
  • Ensure questions are meaningful and easily answerable.
  • Avoid ambiguous terms and ensure wording communicates exactly what is measured.
Avoiding common pitfalls
  • Avoid unanswerable questions where respondents lack information or no option applies (e.g., income ranges that don’t fit respondents).
  • Avoid loaded questions that push toward a socially desirable answer or that involve emotionally charged issues.
  • Avoid double-barreled questions that mix multiple issues in one item.
  • Avoid response bias from order effects by randomizing option order when feasible.
Double-barreled questions (examples)
  • Example 1 (owning Apple products):
    • "Do you feel that owning Apple products has made you more creative and successful?"
    • Variants show different degrees of concreteness and potential bias.
  • Example 2 (orange juice):
    • "Do you drink orange juice for breakfast, lunch, and dinner?" vs. individual meal variants.
  • Takeaway: ensure each item targets a single construct only.
Unanswerable and loaded questions
  • Unanswerable: when respondents cannot provide a meaningful answer due to lack of information or non-applicable options.
  • Loaded: questions that prompt a socially desirable response or hinge on emotional or political framing (e.g., taxes, immigration, sensitive topics).
  • Examples show how to structure questions to minimize bias and capture genuine attitudes.

Sensitive questions

  • Some domains (income, sexual beliefs/behaviors, medical conditions, alcohol use) see lower honesty due to discomfort.
  • Respondents may give socially desirable answers or refuse to answer.
  • Strategies: place sensitive questions toward the end; reassure confidentiality; phrase questions about behavior of others to reduce self-consciousness while still predicting respondent behavior.
  • Examples illustrate differential phrasing to reduce defensiveness.

Screening and layout in questionnaire design

  • Screening questions ensure you talk to the right target population (e.g., recent car buyers, iPad owners).
  • Screening is common in computer-assisted surveys (e.g., Qualtrics).
  • Introductory section provides research overview and legitimizes the study; assure anonymity/confidentiality.
  • The research questions should flow from general to specific questions to avoid priming and bias.
  • Start with broad questions about preferences, then move to specifics (e.g., overall liking before camera features on an iPhone).
  • Response order bias: randomize options when possible to prevent central tendency or anchor effects.
  • Sensitive questions should be placed late; demographic questions at the end to maintain comfort.
  • General flow: Intro → Screener → Research questions → General to specific questions → Sensitive questions → Demographics → Thank-you.
Introduction and screening details
  • Introduction should certify legitimacy, identify involved organizations, and guarantee anonymity/confidentiality.
  • Screening questions help ensure only qualified respondents participate; can be crucial for computer-assisted surveys.
  • Checklist examples include ownership or usage questions (e.g., Do you own an iPad? Have you eaten at Eagle Landing in the past month?).
Question sequencing and transitions
  • Begin with general questions; progress to more specific ones.
  • Use warm-up questions to ease respondents into the survey; transition questions link sections to the research objectives.
  • Sensitive questions should be placed after respondents have built comfort with the survey.
  • Demographic questions appear at the end.

Designing good questionnaires: practical layout rules

  • Think like a respondent and check clarity and answerability of questions.
  • Use a logical order and a general-to-specific flow to reduce sequence bias.
  • Consider adding skip patterns to avoid irrelevant questions for some respondents.
  • Ensure questionnaire looks attractive and professional; decide on incentive use with the client.

The questionnaire design process (7 steps)

  • Confirm research objectives.
  • Select appropriate data collection method.
  • Develop questions and scaling.
  • Determine layout of the questionnaire.
  • Obtain client approval.
  • Pretest, revise, and finalize.
  • Implement survey.

Questionnaire layout and client approval

  • Copies of the questionnaire should be shared with all project stakeholders for feedback.
  • The questionnaire should be attractive and professional; incentives can be used as agreed.
  • Client approval is a formal step before fielding.

Pretesting and piloting the questionnaire

  • Pretest with a small sample (10–20 participants) to identify difficulties or questions.
  • Check for signs of boredom (e.g., skipped questions) and estimate completion time.
  • Use pretests to revise the questionnaire before full deployment.

Data collection and implementation

  • The data collection process varies by modality (online, mail, telephone).
  • Ensure data collection procedures align with the study design and client agreements; incentives, if any, should be used as agreed.

Scale measurement and questionnaire integration

  • Scale measurement: the process of assigning descriptors to represent the range of possible responses to a question about a construct.
  • A questionnaire includes a set of scales and questions; each scale measures one construct.
  • Scales measure both the independent variable(s) and dependent variable(s).
  • The questionnaire should include measures (e.g., scales) for both the independent and dependent variables.
Reliability and validity: a quick framework
  • Reliability: the extent to which a scale produces consistent results across repeated measurements.
  • Validity: the extent to which a scale measures what it is supposed to measure.
  • Not necessarily correlated: high reliability does not guarantee high validity, and high validity does not guarantee high reliability.
  • Core techniques for reliability:
    • Test-retest reliability
    • Equivalent form reliability
    • Internal consistency reliability (e.g., Cronbach’s alpha)
  • Core validity types include face validity, content validity, convergent validity, and discriminant validity.

Reliability: key approaches

  • Test-retest reliability: administer the same scale to the same respondents at two different times under similar conditions; similar results indicate stability.
  • Equivalent form reliability: create two different but equivalent versions of the same scale (Version A vs Version B) and administer to the same or similar samples; high similarity indicates reliability.
  • Internal consistency reliability: assess how well the items on a scale hang together to measure the same construct.
    • Split-half method: compare scores from two halves of the scale (e.g., odd vs. even items or random halves); high correlation indicates internal consistency.
    • Coefficient alpha (Cronbach’s alpha): the most common metric for multi-item scales; ranges from 0 to 1; higher is better up to a point.
Split-half example
  • Example scale (customer satisfaction): items rated on a 1–10 scale (Strongly Disagree to Strongly Agree):
    • Example items include statements about overall satisfaction, perceived value, delivery quality, defect absence, staff courtesy, care for business, facilities, website usefulness, product information, likelihood to recommend.
    • A two-column presentation can illustrate two halves of the scale with the same item structure.
Cronbach’s Alpha and interpretation
  • Cronbach’s alpha (a) is a measure of internal consistency reliability.
  • Acceptable reliability: exta0.70ext{a} \ge 0.70
  • Higher categories (from slides):
    • Excellent: a0.90a \ge 0.90
    • Good: 0.90 > a \ge 0.80
    • Acceptable: 0.80 > a \ge 0.70
    • Questionable: 0.70 > a \ge 0.60
    • Poor: 0.60 > a \ge 0.50
    • Unacceptable: a < 0.50
  • When Cronbach’s alpha is too high (e.g., a > 0.95), it may indicate redundancy among items.

Validity: what it means to measure the right thing

  • Validity assesses whether a scale measures the intended construct.
  • Face validity: is the indicator intuitively reasonable as a measure of the construct? Based on expert judgement.
  • Content validity: the items cover all relevant dimensions of the construct (e.g., satisfaction with a restaurant should include food quality, courtesy, wait time, ambience).
  • Convergent validity: measures that should be related are indeed related (e.g., related forms of a mood scale correlate).
  • Discriminant validity: measures that should not be related are indeed not related (e.g., brand authenticity vs. brand personality should diverge).

Reliability vs validity: key relationship

  • High reliability does not guarantee high validity; high validity does not guarantee high reliability.
  • Both are essential for sound measurement.

Other considerations in scale design

  • Single-item vs. multiple-item scales:
    • Single-item scale: describes only one attribute (e.g., Age).
    • Multiple-item scale: covers multiple attributes of a construct (e.g., Brand loyalty: purchase frequency, brand liking, brand recommendation, amount spent).
  • Many marketing measures use multi-item scales because they capture more dimensions.
  • Discriminatory power of scale descriptors refers to how well scale points differentiate responses.
  • Balanced vs. unbalanced scales:
    • Balanced: equal number of positive and negative items.
    • Unbalanced: more items on one side; can introduce bias unless attitudes are genuinely one-sided (e.g., tax increases).
  • Forced vs. non-forced scales:
    • Forced: no neutral option; pushes choice toward positive or negative.
    • Non-forced: includes a neutral option or middle point (e.g., N/A).
  • Unstructured questions vs. structured scales:
    • Structured (close-ended): quick to answer and analyze; preferred for efficiency.
    • Unstructured (open-ended): richer data but time-consuming to analyze; respondents may skip.

Discriminatory power and number of scale points

  • More scale points increase variability and discriminatory power but may complicate responses.
  • Common practice in marketing: use 5-point, 7-point, or 11-point scales; debate on extremely high scales (e.g., 500 points) is generally discouraged due to respondent fatigue.
  • Example scale point discussion:
    • Balanced, 5-point or 7-point scales are typical; central values tend to be chosen by respondents to save effort.

Practical questionnaire design rules

  • Place sensitive questions toward the end; place demographic questions at the end.
  • Ensure questions are clear, answerable, and aligned with research objectives.
  • Use skip patterns to avoid irrelevant questions for some respondents.
  • Consider introducing the survey with a legitimacy statement and confidentiality assurances.

The big picture: questionnaire design workflow

  • Start with a clear objective and map questions to constructs/measures.
  • Choose data collection method and sampling frame with the client.
  • Draft questions and scaling; design the layout to guide respondents logically.
  • Get client approval; pretest the questionnaire with a small sample.
  • Revise based on feedback and pretest results.
  • Implement the survey and collect data; monitor for issues during data collection.

Key takeaways for exam readiness

  • Always link each question to a single construct; avoid ambiguity and multi-issue items.
  • Use appropriate language and avoid biasing phrasing.
  • Structure surveys to minimize biases: general-to-specific flow, randomize options when feasible, and place sensitive items late.
  • Distinguish reliability (consistency) from validity (accuracy); use tests such as test-retest, equivalent form, and internal consistency (Split-half, Cronbach’s alpha) to evaluate scales.
  • Be mindful of scale design choices: number of items, balance, forcing of choices, and mode of data collection.
  • Use content, face, convergent, and discriminant validity to ensure constructs are measured as intended.

Quick reference: glossary of terms from the slides

  • Reliability: consistency of a measure across time or items.
  • Validity: accuracy of a measure in capturing the intended construct.
  • Face validity: expert judgment on whether a measure appears to assess the intended construct.
  • Content validity: coverage of all relevant aspects of the construct.
  • Convergent validity: related measures converge.
  • Discriminant validity: unrelated measures diverge.
  • Cronbach’s alpha: internal consistency reliability metric for multi-item scales; typically interpreted with thresholds (Excellent, Good, Acceptable, Questionable, Poor, Unacceptable) as discussed above.
  • Test-retest reliability: consistency over time.
  • Equivalent form reliability: consistency across alternative forms of a scale.
  • Split-half reliability: internal consistency via correlating halves of a scale.
  • Likert-type scales: common multi-item response formats (e.g., 5-, 7-, or 11-point scales).
  • Double-barreled question: item that asks about two issues simultaneously.
  • Loaded question: item that presupposes a socially desirable response.
  • Demographic questions: usually placed at the end of the survey.
  • Pretest: pilot run with a small sample to identify problems before full deployment.
  • Skip patterns: logic to bypass irrelevant questions based on earlier responses.