Study Notes: Data Analytics

BUSINESS IN PRACTICE: DATA ANALYTICS

SESSION OVERVIEW

  • Henley Business School, University of Reading

  • Course Objective: Learn to manage, visualize, analyze data sets, and infer meaningful conclusions.

MODULE INFORMATION

  • Essential Processes for Data Analysts

    • Descriptive Analytics: Analyzing past data.

    • Predictive Analytics: Predicting future events using historical data.

    • Prescriptive Analytics: Using data to alter outcomes.

  • Scope: The module covers foundational aspects of all three analytics processes.

DEFINITION OF DATA ANALYTICS

  • Data Analytics: A field that bridges statistics and software development to address data-related issues.

    • Key Steps:

      • Collect data

      • Organize and store it

      • Analyze it

      • Draw conclusions from findings

    • Outcome: Produces data-driven recommendations.

  • Business Data Analytics: Focuses on addressing business-related questions utilizing data analytics principles.

IMPORTANCE OF DATA ANALYTICS

  • Reasons for Importance: The importance of data analytics spans various industries and operations; specific reasons will be expanded in detailed notes.

MODULE LIMITATIONS

  • The module does not cover every facet of data analytics, as the field evolves continually.

    • Learning Outcome: Students should gain the competence and confidence to analyze data and perform independent research on methodologies not included in the course.

CONTACT INFORMATION

  • Instructor Email: m.kyritsis@henley.ac.uk

  • Dr. Nico Biagi Email: nicolo.biagi@henley.ac.uk

  • IT Team Email for Blackboard Queries: it@reading.ac.uk

ASSESSMENTS FOR THE MODULE

  • Types of Assessment:

    • Multiple Choice In-Class Quiz (Week 7) - 50% of grade

    • Data Analytics Project (Hybrid Online Assessment/Test) - 50% of grade

  • Recommendations: Attend all quizzes, workshops, and seminars; seek help early if needed.

CONTENT OF ASSESSMENTS

  • Focus: Assessments focus on seminar and workshop material and practical applications rather than lecture content.

    • Importance: Maintaining pace with seminar and workshop activities is critical for success.

TOOLS USED

  • Required Tools: R-commander

  • Optional Tools: Excel

    • Compatibility: Available on both Windows and Mac, accessible via apps.

    • Installation support: provided by Teaching Assistants during workshops.

DESCRIPTIVE STATISTICS AND ITS LIMITATIONS

  • Descriptive Statistics: Overview and importance discussed.

  • Caution Against Generalizing:

    • Descriptive statistics should not be used to generalize findings; limitations will be explained through practical examples.

MR. ROGERS' CHICKEN EXAMPLE

  • Scenario Setup: Mr. Rogers tests two feeds on chicken groups to assess which leads to higher weight gains to evaluate profit potential.

    • Variables:

      • Independent Variable: Type of feed (Feed A and Feed B)

      • Dependent Variable: Weight of chickens

    • Formal Hypothesis:

      • Null Hypothesis (H0): No difference in mean weights between groups

      • Alternative Hypothesis (H1): Difference exists in mean weights.

  • Data Access: Available as Roger.csv on Blackboard, part of the ChickWeight dataset.

EXPERIMENTAL RESULTS

  • Findings: Average weights show Feed A (142.95g) appears superior to Feed B (135.27g); however, this is inconclusive.

RANDOM GROUPING EXPERIMENT

  • Method: Mr. Rogers splits chickens fed on Feed A into two groups for further weight measurement.

    • Results: Group A1 has a mean weight of 149.62g, while Group A2 shows 136g; the differences exemplify natural variation rather than feed effect.

INFERENTIAL STATISTICS

  • Key Takeaway:

    • Certainty of findings is unattainable due to chance factors; instead, a probability % can indicate confidence in experimental outcomes.

  • Statements about significance reflect confidence (e.g., 90%, 95%, etc.) that differences observed are not due to chance.

DESCRIPTIVE STATISTICS SIGNIFICANCE

  • Utility of Descriptive Stats:

    • Summarizing data

    • Visualizing trends

    • Inputs for inferential statistics and modeling processes.

PARAMETER AND STATISTIC DISTINCTION

  • Parameter (Population):

    • Example: Average Weight of Cross: 125g

    • Average Price: £1.25

  • Statistic (Sample):

    • Example: Average Weight: 130.4g

    • Average Price: £1.44

    • Probability related to star (P(star)): figures exemplified.

SAMPLING METHODS

  • Sampling types covered:

    • Simple Random Sample

    • Systematic Sample

    • Stratified Sample

    • Cluster Sample

TYPES OF DATA

  • Qualitative Data:

    • Nominal: No order (e.g., eye color)

    • Ordinal: With order (e.g., Likert scale satisfaction levels)

  • Quantitative Data:

    • Discrete: Whole numbers (1,2,3…)

    • Continuous:

      • Interval: Lacks absolute zero (e.g., temperature in Celsius)

      • Ratio: Has absolute zero representing absence (e.g., weight).

MEASURES OF CENTRAL TENDENCY

  • Mean: Average value calculated using all data points.

    • Formulation: xˉ=racextSumofallobservationsn\bar{x} = rac{ ext{Sum of all observations}}{n}

  • Median: The middle value when data is sorted. Not affected by outliers.

  • Mode: Most frequent value; applicable in qualitative data analysis.

MEASURES OF DISPERSION

  • Variance: Indicates how far data points spread from the mean.

  • Standard Deviation (SD): Measure of the amount of variation in a set of values.

    • Formulations:

      • Population Variance: <br>u2=rac1NimesextSumof(xi−xˉ)2<br>u^2 = rac{1}{N} imes ext{Sum of } (x_i - \bar{x})^2

      • Sample Variance: s2=rac1n−1imesextSumof(xi−xˉ)2s^2 = rac{1}{n-1} imes ext{Sum of } (x_i - \bar{x})^2

  • Importance of SD: Vital for assessing distance from the mean and determining confidence levels pertaining to the mean.

EXAMPLE: IQ SCORES

  • Context of IQ Distribution:

    • Average IQ set at 100; standard deviation of 15.

    • Representation of IQ follow a normal distribution indicated as: IQhicksimN(100,15)IQ hicksim N(100, 15).

SUMMARY

  • Analytics Types: Descriptive, Predictive, and Prescriptive

  • Business Applications:

    • Utilizing data for customer behavior insights, classification, security, market analysis, forecasting, etc.

  • Descriptive Stats Value: Important for summarization, visualization, and foundational input into inferential statistics; caution in generalization urged.

  • Central Tendency Methods: Mean (outlier susceptible), Median (robust), Mode (qualitative relevant).

  • Broad Data Types: Continuous, interval, ordinal.

  • Dispersion Measures: Critical for understanding variance impact on mean confidence.