Marketing Research Data Sources and Analytics

Foundations of Marketing Research Data: Primary vs. Secondary Data

Core Definitions

  • Primary Data: Original data collected directly by the researcher or user specifically to address the research question at hand.

  • Secondary Data: Data that has already been collected by another entity for a purpose other than the specific question currently being investigated.

Comparison Matrix of Primary vs. Secondary Data

  • Cost:

    • Primary Data: High cost (-)

    • Secondary Data: Low cost (+)

  • Time Required:

    • Primary Data: Slow / Time-consuming (-)

    • Secondary Data: Fast / Immediately accessible (+)

  • Breadth of Insights:

    • Primary Data: Limited initial breadth, but can be expanded through detailed cross-tabulations and advanced statistical modeling (+/-)

    • Secondary Data: High breadth (+)

  • Depth of Insights:

    • Primary Data: Deep insight into customer motivations, attitudes, and underlying reasons (+)

    • Secondary Data: Shallow depth (-)

  • Actionability:

    • Primary Data: High actionability because the study is custom-designed to address specific managerial decision requirements (+)

    • Secondary Data: Low or variable actionability depending on data alignment with the problem (-)

Empirical Demonstration: Primary vs. Secondary Data in Published Research

  • Study Details: Published in Sleep Health (Journal of the National Sleep Foundation) by Jason J. Jones, PhD, Gregory W. Kirschen, PhD, Sindhuja Kancharla, MS, and Lauren Hale, PhD (Stony Brook University).

  • Research Question: Association between late-night tweeting and next-day game performance among professional basketball players.

  • Data Classification: Secondary Data Analysis (utilizing existing public timestamped Twitter/X activity paired with official National Basketball Association [NBA] box scores).

  • Sample Size: 112112 professional NBA players.

  • Key Metrics & Empirical Findings:

    • Late-night tweeting was defined as Twitter activity occurring between 11:00 PM11\text{:00 PM} and 7:00 AM7\text{:00 AM}.

    • Within-person performance reductions were observed on the day following late-night tweeting, including fewer total points scored and fewer total rebounds.

    • Decreases were also observed in playing time per game, as well as decreases in negative output statistics such as turnovers and personal fouls (driven by reduced playing time).

    • Shooting Accuracy Penalty: Shooting accuracy—a critical metric that is independent of total minutes played—showed the clearest performance drop. Players successfully made shots at a rate 1.71.7 percentage points lower (1.7 percentage points1.7\text{ percentage points}) following late-night tweeting activity.

Real-World Data Classification Scenarios

  • Scenario 1 (Secondary Data): A financial firm reviews credit scores, debt levels, and payment histories purchased from consumer reporting agencies to analyze the financial profiles of prospective target customers.

  • Scenario 2 (Primary Data): A enterprise recruits a custom participant panel to track behavioral shifts and brand attitudes over time through recurring surveys and active logging.

Qualitative vs. Quantitative Research Methodologies

Methodological Dichotomy

  • Qualitative Research:

    • Data Format: Non-numeric, unstructured text, audio, or video.

    • Analytical Method: Subjective interpretation, contextual coding, and thematic analysis.

    • Objective: Provides deep, rich contextual understanding (e.g., exploring how consumers perceive a brand and why).

    • Sample Size: Small number of non-random cases.

  • Quantitative Research:

    • Data Format: Structured, numeric data.

    • Analytical Method: Formal statistical hypothesis testing and numerical modeling.

    • Objective: Recommends definitive courses of action (e.g., evaluating a go vs. no-go product launch decision).

    • Sample Size: Large, representative sample sizes.

Focus Groups: Diagnostic and Exploratory Applications

  • Role: Primary qualitative method used for generating hypotheses, probing consumer motives, testing new product concepts, and diagnostic research to lay the groundwork for quantitative studies.

  • Structural Limitations:

    • Sample Non-Representativeness: Conducted with small, non-random samples; results cannot be extrapolated to estimate overall market share.

    • Group Dynamics: Vulnerable to distortion through dominant vocal participants, group conformity pressures, and social desirability bias.

    • Moderator Effects: Framing, tone, and active probing by the moderator can inadvertently alter or bias participant responses.

Behavioral Distortions in AI Synthetic Respondents

When using generative AI models to simulate human consumer responses in qualitative research, five systemic distortions emerge:

  1. Insufficient Individuation: Simulated personas tend toward uniform average responses, smoothing out extreme or idiosyncratic viewpoints.

  2. Stereotyping: The model relies on broad cultural generalizations (e.g., "what a typical UT student would say") rather than authentic lived experience.

  3. Representation Bias: Synthetic response fidelity varies unevenly across demographic, socioeconomic, or cultural sub-groups.

  4. Ideological Bias: Responses drift systematically toward the embedded training values of the model (e.g., pro-social, pro-environmental, or tech-optimistic bias).

  5. Hyper-Rationality: AI respondents display unrealistic levels of logical consistency, thoroughness, self-awareness, and information processing compared to real humans.

Categories of Marketing Research Insights and Methodological Trade-offs

Key Consumer Insights Derived from Primary Data

Primary marketing research targets six distinct categories of consumer information:

  1. Demographic Characteristics: Who the customer is (e.g., age, income, education, gender).

  2. Attitudes and Opinions: What consumers think and feel regarding products, brands, or values.

  3. Awareness and Knowledge: What brands or solutions exist within the consumer's consideration set.

  4. Intentions: Planned or anticipated future purchase actions.

  5. Motivation / Protocol: The underlying psychological drive or process behind consumer actions (why they buy).

  6. Behavior: Direct, observable consumer actions (what, where, when, and how much they actually purchase).

Consumer Behavior Gap Analysis

  • Attitudes vs. Intentions Gap: Occurs when a consumer expresses positive brand attitudes but reports zero purchase intention (e.g., "I love this sports car, but I do not intend to buy it"). Key drivers include financial constraints (too expensive), lack of immediate operational need, or current ownership of an alternative.

  • Intentions vs. Behavior Gap: Occurs when stated purchase intentions fail to translate into actual transaction behavior (e.g., "I intended to purchase that item today, but I did not"). Key drivers include point-of-sale execution issues, out-of-stock inventory, localized marketing disruptions, or spontaneous impulse purchases of competing products.

Primary Data Collection: Questioning vs. Observation

  • Questioning Methods: In-depth interviews (IDIs), focus groups, and structured surveys.

  • Observation Methods: Direct physical observation, passive behavior tracking, and field experiments.

  • Comparative Trade-off Analysis:

    • Versatility: Questioning (+), Observation (-)

    • Realness / Accuracy: Questioning (-), Observation (+)

    • Respondent Convenience: Questioning (-), Observation (+)

    • Depth of Insights: Questioning (+), Observation (-)

Taxonomy of Marketing Research Designs

Research Classification Framework

Marketing research projects are structured based on the clarity of the problem definition:

  1. Exploratory Research:

    • Problem Definition: Highly ambiguous problem context (e.g., "Our product sales are declining and we have no understanding of why").

    • Primary Data Sources: In-depth interviews, focus groups, online qualitative research communities, qualitative observation, exploratory surveys.

    • Secondary Data Sources: Social media sentiment analysis, existing internal survey repositories, organizational records, syndicated research databases.

  2. Descriptive Research:

    • Problem Definition: Awareness of the specific problem context exists, requiring detailed profile quantification (e.g., "What are the specific demographic profiles of our customer base versus our main competitors' customers?").

    • Managerial Focus: Quantifying market composition, establishing share of wallet, identifying at-risk customer churn segments, evaluating buyer habits.

    • Collection Modes: Active Data Collection (surveys, participant diaries) and Passive Data Collection (POS terminal scans, media tracking, search engine telemetry, app telemetry).

  3. Causal Research:

    • Problem Definition: Problem context is clearly and precisely defined, seeking to isolate direct cause-and-effect relationships (e.g., "Will changing our product packaging design from blue to green produce a statistically significant increase in unit sales?").

    • Primary Methods: Formal field experiments, laboratory experiments, and randomized controlled trial (A/B testing) frameworks.

Point-of-Sale (POS) and Scanner Data Architecture

Structural Dimensions of POS Data

Point-of-Sale (POS) scanner data captures passive observational purchase data at store checkout terminals. The full analytical value of POS data is expressed by the multi-dimensional structure: POS Data Value=Geography×Product×Time×Variables\text{POS Data Value} = \text{Geography} \times \text{Product} \times \text{Time} \times \text{Variables}

  • Geography: Retail outlet identity, channel, region, city, zip code, store format (e.g., Target, Walmart, H-E-B).

  • Product: Stock Keeping Unit (SKU) details, brand name, sub-brand, package physical size, flavor/variant, container type.

  • Time: Transaction timestamp, hour, day of week, calendar date, week number, month, season.

  • Variables: Retail shelf price paid, feature advertisement placement, temporary price reduction (TPR) status, end-of-aisle display presence, in-store promotional placement.

The Consumer Packaged Goods (CPG) Data Ecosystem Scale

  • Manufacturers: Over 80 to 10080\text{ to }100 major Consumer Packaged Goods (CPG) corporate manufacturers.

  • Distribution Hubs: Dozens of major warehouse and regional distribution networks.

  • Retail Outlets: Approximately 45,00045{,}000 physical supermarkets and mass retail locations.

  • End-Consumers: Over 135,000,000135{,}000{,}000 individual domestic households.

Micro-Level Brand Switching and Household Panel Analytics

The Aggregate Sales Illusion

Aggregate retail sales tracking can mask substantial underlying shifts in consumer behavior. Consider a two-period retail market study tracking total unit sales across four competing brands:

  • Period 1 (t1t_1) Aggregate Sales:

    • Sudsy: 200 units200\text{ units}

    • Sunshine: 300 units300\text{ units}

    • Sparkle: 350 units350\text{ units}

    • Silky: 150 units150\text{ units}

    • Total Market Volume: 1,000 units1{,}000\text{ units}

  • Period 2 (t2t_2) Aggregate Sales:

    • Sudsy: 250 units250\text{ units} (+50 gain+50\text{ gain})

    • Sunshine: 270 units270\text{ units} (30 loss-30\text{ loss})

    • Sparkle: 330 units330\text{ units} (20 loss-20\text{ loss})

    • Silky: 150 units150\text{ units} (zero net change\text{zero net change})

    • Total Market Volume: 1,000 units1{,}000\text{ units}

Managerial Misinterpretation: Looking only at aggregate volume, the brand manager for Silky would conclude that their customer base is perfectly stable (150 units150\text{ units} in both periods), while Sunshine and Sparkle appear to be losing customers directly to Sudsy.

Disaggregating Sales via the Brand-Switching Matrix

By leveraging store loyalty cards and syndicated household panels (e.g., NielsenIQ), customer purchases can be tracked across consecutive time periods (t1t_1 to t2t_2):

Brand Bought in Period 1 (t1t_1)

Bought Sudsy (t2t_2)

Bought Sunshine (t2t_2)

Bought Sparkle (t2t_2)

Bought Silky (t2t_2)

Total Units (t1t_1)

Bought Sudsy

175175

2525

00

00

200200

Bought Sunshine

00

225225

5050

2525

300300

Bought Sparkle

00

00

280280

7070

350350

Bought Silky

7575

2020

00

5555

150150

Total Units (t2t_2)

250250

270270

330330

150150

1,0001{,}000

  • Matrix Mechanics:

    • Main Diagonal Elements (175,225,280,55175, 225, 280, 55): Represent repeat buyers who remained brand-loyal from Period 1 to Period 2.

    • Off-Diagonal Elements: Represent brand-switching movements between competitors.

Calculating Brand Loyalty and Switching Probabilities

Dividing each matrix cell by its corresponding row total yields the conditional switching probabilities P(Brandt2Brandt1)P(\text{Brand}_{t2} | \text{Brand}_{t1}):

  • Sudsy (t1=200t_1 = 200):

    • Loyalty Probability P(Sudsyt2Sudsyt1)=175200=0.875P(\text{Sudsy}_{t2} | \text{Sudsy}_{t1}) = \frac{175}{200} = 0.875

    • Switching to Sunshine P(Sunshinet2Sudsyt1)=25200=0.125P(\text{Sunshine}_{t2} | \text{Sudsy}_{t1}) = \frac{25}{200} = 0.125

  • Sunshine (t1=300t_1 = 300):

    • Loyalty Probability P(Sunshinet2Sunshinet1)=225300=0.750P(\text{Sunshine}_{t2} | \text{Sunshine}_{t1}) = \frac{225}{300} = 0.750

    • Switching to Sparkle P(Sparklet2Sunshinet1)=503000.167P(\text{Sparkle}_{t2} | \text{Sunshine}_{t1}) = \frac{50}{300} \approx 0.167

    • Switching to Silky P(Silkyt2Sunshinet1)=253000.083P(\text{Silky}_{t2} | \text{Sunshine}_{t1}) = \frac{25}{300} \approx 0.083

  • Sparkle (t1=350t_1 = 350):

    • Loyalty Probability P(Sparklet2Sparklet1)=280350=0.800P(\text{Sparkle}_{t2} | \text{Sparkle}_{t1}) = \frac{280}{350} = 0.800

    • Switching to Silky P(Silkyt2Sparklet1)=70350=0.200P(\text{Silky}_{t2} | \text{Sparkle}_{t1}) = \frac{70}{350} = 0.200

  • Silky (t1=150t_1 = 150):

    • Switching to Sudsy P(Sudsyt2Silkyt1)=75150=0.500P(\text{Sudsy}_{t2} | \text{Silky}_{t1}) = \frac{75}{150} = 0.500

    • Switching to Sunshine P(Sunshinet2Silkyt1)=201500.133P(\text{Sunshine}_{t2} | \text{Silky}_{t1}) = \frac{20}{150} \approx 0.133

    • Loyalty Probability P(Silkyt2Silkyt1)=551500.367P(\text{Silky}_{t2} | \text{Silky}_{t1}) = \frac{55}{150} \approx 0.367

Crucial Analytical Insights for Brand Managers

  • Silky's Vulnerability: Silky's aggregate sales stability (150 units150\text{ units}) masked a severe customer defection rate. Silky retained only 5555 of its original 150150 customers (36.7%36.7\% brand loyalty rate), losing 7575 customers (50.0\%$) directly to Sudsy and 20customers(customers (13.3\%$) to Sunshine.

  • Silky's Compensating Gains: Silky avoided an aggregate sales collapse only by poaching 7070 customers from Sparkle and 2525 customers from Sunshine.

  • Sudsy's Growth Driver: Sudsy's overall net volume gain of +50 units+50\text{ units} was driven entirely by acquiring 7575 former Silky buyers, despite losing 2525 of its own buyers to Sunshine.

Analytical Evaluation of Sales Promotions

Temporal Dynamics of Sales Promotions

When analyzing store-level promotional POS data over time (e.g., across an 8-week8\text{-week} observation horizon), sales volume displays three distinct structural phases:

  1. Anticipation Dip: A decline in baseline sales immediately preceding an expected promotional event, occurring because consumers delay planned purchases to wait for the discount.

  2. Promotional Peak: An immediate surge in volume during the active promotional period (e.g., Week 5 sales spike).

  3. Post-Promotion Dip: A drop in volume below baseline immediately following the promotion, caused by consumer stockpiling (pantry loading / forward buying).

Evaluating True Promotional Lift

  • Analytical Mistake: Calculating promotional effectiveness by subtracting baseline sales from peak promotional sales alone overstates true promotional lift.

  • Correct Calculation Framework: To determine net promotional lift, managers must sum total sales generated during the promotional spike and subtract both the pre-promotion anticipation deficit and the post-promotion stockpiling dip relative to non-promotional baseline volume.

Core Managerial Questions Addressed by POS Promotion Analytics

  • Promotional Audience Profile: Are discounted units being purchased by loyal buyers who would have bought at full price, or by deal-prone switchers ("cherry pickers")?

  • Forward Buying Effects: Are buyers borrowing volume from their own future demand, or generating true incremental volume?

  • Long-Term Conversion: Do promotional switchers transition into repeat, full-price loyal customers over time?

  • In-Store Merchandising Effectiveness: What is the relative ROI of end-of-aisle display placement versus feature ad placements?

  • Cross-Category Effects: Which product categories act as true operational substitutes versus complements during temporary price promotions?

Limitations of POS / Scanner Data

  • Incomplete Retail Channel Coverage: Syndicated POS data does not track every retail channel, independent dealer, or specialized e-commerce outlet.

  • Absence of Direct Causality: Observational correlation in POS trend lines does not establish definitive mathematical causality.

  • Missing Consumer Psychographics: POS transactions record what was bought, but contain zero psychographic measurement regarding consumer values, perceptions, or attitudes.

  • Unobserved Choice Set: Standard POS databases reveal what item was purchased, but cannot confirm what alternative competing items were physically present on the retail shelf at the exact moment of choice.

Strategic Rationale for Purchasing Syndicated POS Data

Commercial firms invest heavily in syndicated data vendors (e.g., NielsenIQ, Circana) due to:

  • Comprehensiveness: Ability to systematically link aggregate and household sales metrics directly to marketing mix instruments.

  • Timeliness: Weekly delivery schedules that support rapid operational and pricing adjustments.

  • Accuracy: Objective transactional recording that eliminates recall bias inherent in self-reported survey data.

The Marketing Data Provider Ecosystem

Specialized Commercial Data Vendor Classification

  1. Group 1: Internet Communities & Qualitative Tech

    • Vendors: Alida (formerly Vision Critical), C Space, Ipsos, itracks.

    • Capabilities: Digital qualitative co-creation platforms, proprietary online research panels, real-time video focus groups.

  2. Group 2: Syndicated Retail Data

    • Vendors: NielsenIQ (NIQ), Circana (formerly IRI + NPD), SPINS (natural and organic products focus).

    • Capabilities: Comprehensive store scanner aggregations, household panel tracking, CPG retail share audits.

  3. Group 3: Transaction & Purchase Intelligence

    • Vendors: Bloomberg Second Measure, Cardlytics.

    • Capabilities: Direct transaction monitoring via credit card/debit card processing data feeds and banking app integration.

  4. Group 4: Media & Audience Measurement

    • Vendors: Comscore, Nielsen.

    • Comscore Digital Measurement Footprint:

      • Desktop Screens: 193M193\text{M}

      • Connected TV (CTV) Screens: 140M140\text{M}

      • Mobile Phones & Tablets: 240M240\text{M}

      • Linear TV Screens: 75M75\text{M}

  5. Group 5: Social Media Analytics

    • Vendors: Hootsuite, Buffer Analytics, Sprout Social.

    • Capabilities: Multi-channel social listening, audience engagement benchmarking, post impression and sentiment analytics.

  6. Group 6: Search Data Intelligence

    • Vendors: Google Trends, Comscore, Semrush, SimilarWeb.

    • Search vs. Social Analytics Distinction: Search queries capture unvarnished consumer intent without the social desirability bias common in public social media posts ("Actions speak louder than words").

    • Real-World Case Example: In the immediate aftermath of the Brexit referendum, the top trending Google query across the United Kingdom was "What is the European Union?", revealing public knowledge gaps that were masked prior to the vote.

  7. Group 7: Web & App Analytics

    • Vendors: Google Analytics, Adobe Analytics, Amplitude.

    • Capabilities: Digital funnel tracking, user flow pathing, retention/churn analytics, conversion tracking.

  8. Group 8: A/B Testing & Field Experimentation

    • Vendors: Optimizely, CleverTap, Mixpanel.

    • Capabilities: Automated randomized digital experiment allocation, variable testing (UI layout, copy, pricing), causal effect isolation.

  9. Group 9: Location & Geospatial Intelligence

    • Vendors: Mogean, Foursquare.

    • Capabilities: Real-world consumer journey tracking utilizing anonymized geospatial temporal location streams from mobile devices.

Classroom Discussions and Applied Exercises

Group Data Source Exploration Activity

  • Task: Student teams analyze specific data vendor websites to determine: (1) core services provided, (2) specific data variables collected, (3) managerial questions answered, and (4) structural limitations.

  • Assigned Team Categories:

    • Team 1: Internet Communities (Alida, C Space, Ipsos, itracks)

    • Team 2: Syndicated Retail Data (NielsenIQ, Circana, SPINS)

    • Team 3: Transaction/Purchase Intelligence (Bloomberg Second Measure, Cardlytics)

    • Team 4: Media & Audience Measurement (Comscore, Nielsen)

    • Team 5: Social Media Analytics (Hootsuite, Buffer Analytics, Sprout Social)

    • Team 6: Search Data (Google Trends, Semrush, SimilarWeb)

    • Team 7: Web and App Analytics (Google Analytics, Adobe Analytics, Amplitude)

    • Team 8: A/B Testing and Experiments (Optimizely, CleverTap, Mixpanel)

    • Team 9: Location/Geospatial Intelligence (Mogean, Foursquare)

Student-Instructor Interaction Contexts

  • Prompt on Descriptive Research Examples: An illustrative campus dining research study is structured to evaluate price sensitivity for meat purchased off-campus versus home-cooked meals, evaluating demand for campus meal plans, hot vs. cold grab-and-go options, and brand-switching habits.

  • Prompt on Aggregate Data Interpretation: When reviewing aggregate sales for Silky showing 150 units150\text{ units} in Period 1 and 150 units150\text{ units} in Period 2:

    • Instructor Question: "If you were Silky's brand manager, what does this aggregate data tell you?"

    • Student Response: "Sparkle and Sunshine lost customers to Sudsy, but Silky looks stable."

    • Follow-Up after Disaggregation: "Looking at the brand switching matrix, should you still be happy?"

    • Student Response: "No, because only 55 of your original customers remained."

  • Prompt on Promotional Dip Calculations:

    • Instructor Question: "When evaluating a promotion spike, how should you account for surrounding volume drops?"

    • Student Response: "You have to sum the promotional period gains and subtract the pre-promotion anticipation dip and post-promotion stockpiling drop to find the true lift."