Introduction to Data Science and Statistical Analysis

Fundamental Concepts of Statistics

  • Definition of Statistics: Statistics is the science of data, encompassing both the science and the art of collecting and analyzing observations. Its primary goal is to facilitate learning about ourselves, our natural surroundings, and the entire universe.
  • Building Blocks: Data serve as the fundamental building blocks of statistics.
  • Definition of Data: Data are observations that are recorded by an observer. These can take various forms, including:
    • Numbers in context
    • Measurements
    • Lists
    • Images
  • Variation: This concept refers to different versions of the same type of data. For example, if one draws a circle multiple times, each circle will be slightly different; these differences represent variation.

Classifying Variables and Populations

  • Variables: A variable is a specific characteristic of people or things. Examples include:
    • The gender of a person.
    • the weight of a newborn puppy.
    • The concentration of CO2CO_2 in the atmosphere.
  • Numerical Variables: These describe quantities of the object of interest and are expressed as numbers. Examples include:
    • Weight of an infant.
    • Time required to run a mile.
    • The number of siblings a person has.
  • Categorical Variables: These describe qualities of interest and are typically expressed in words or categories. Examples include:
    • Academic major.
    • Eye color.
    • Type of precipitation.
  • Numerical vs. Categorical Classification Exercise:
    • Height of a bridge: Numerical
    • GPA: Numerical
    • Letter grade in class: Categorical
    • Hours worked each week: Numerical
    • Type of pets owned (cat, dog, etc.): Categorical
    • Flower varieties planted in a garden: Categorical
    • Living situation: Categorical
    • Commute distance: Numerical
    • Number of aunts: Numerical
  • Population vs. Sample:
    • Population: This includes everyone or everything within a defined group (e.g., all babies born in North Carolina in 20042004, or the entire population involved in a census).
    • Sample: A specific subset of the population (e.g., six specific babies selected from the North Carolina birth records, or a survey of 3030 SWTC students used to represent the student body).

Data Storage and Coding

  • Coding Data: This is the process of using numbers to record categorical data to make reading or analysis easier. For example, gender may be coded where 11 represents "Female" and 00 represents "Male" (or "No" to the condition of being female).
  • Nature of Coded Categorical Data: Even if a variable is recorded as a number (like 00 and 11), it remains categorical because mathematical operations like addition often do not make sense. For instance, adding labels for gender does not yield a meaningful numerical sum.
  • Stacked Data Format:
    • Data values are stored in a spreadsheet format where each row contains data for a single individual.
    • This format allows for the storage of many different variables simultaneously.
    • Example: A row for a single infant might include variables for weight (7.697.69), gender (FF), and maternal smoking status (00).
  • Unstacked Data Format:
    • Data values are stored in two separate columns, where each column represents a variable from a different group.
    • This format can only store data for two variables.
    • Critically, information in a single row does not correspond to the same individual. For example, separate columns for "Men's Heights" and "Women's Heights."

Organizing and Calculating Categorical Data

  • Two-Way Tables: These tables show how many times each combination of categories occurs within a dataset.
  • Frequency (Count): This is the number of times a specific value is observed. In a table regarding gender and seat belt use, if the cell for "Men" and "Not Always" is 22, it indicates two men in the sample do not always wear seat belts.
  • Calculating Proportions, Percentages, and Rates:
    • Proportion: Calculated as Number of observed unitsTotal population\frac{\text{Number of observed units}}{\text{Total population}}. For New York AIDS cases (75,25375,253 cases in a population of 19,297,72919,297,729), the proportion is 7525319297729=0.00390\frac{75253}{19297729} = 0.00390.
    • Percentage: Calculated as Proportion×100%\text{Proportion} \times 100\%. For the same NY data: 0.00390×100%=0.390%0.00390 \times 100\% = 0.390\%.
    • Rate: Often expressed per a standard unit, such as "per thousand." For the NY data, the rate is 3.903.90 per thousand.
  • Problem Applications:
    • In a class of 1515 men and 2323 women, the total is 3838. The percent of the class that is male is 1538×100%≈39.47%\frac{15}{38} \times 100\% \approx 39.47\%.
    • In a class of 234234 students where 64.1%64.1\% are men, the number of men is 234×0.641=149.994234 \times 0.641 = 149.994, which rounds to approximately 150150 men.
    • In a class that is 40%40\% women with 2020 women total, the total number of students (xx) is found via 0.40x=200.40x = 20, resulting in x=50x = 50 students.

Causality and Study Design

  • Establishing Causality: Showing that changes in one variable (the treatment variable) directly affect another variable (the outcome or response variable).
  • Study Components:
    • Treatment Variable: The variable assigned or manipulated by the researcher.
    • Outcome (Response) Variable: The variable that may respond to the application of the treatment.
    • Treatment Group: The group of subjects that receive the treatment.
    • Control (Comparison) Group: The group that does not receive the treatment.
  • Case Study: Peanut Milk and Gum Disease:
    • Treatment variable: Drinking Peanut Milk.
    • Response variable: Cure of gum disease.
    • Treatment group: Individuals drinking Peanut Milk.
    • Control group: Individuals not drinking Peanut Milk.
  • Observational Studies:
    • These record data without attempting to influence responses; no treatment is actively applied.
    • The purpose is to describe a group or situation.
    • Limitation: They typically cannot prove cause and effect.
    • Association: If an outcome occurs more often in one group than another without identifying a direct cause, the treatment and outcome are merely associated.
  • Association vs. Causation:
    • Association does not imply causation. For example, people with grey hair often have more wrinkles, but grey hair does not cause wrinkles.
    • Confounding Variable: A characteristic other than the treatment that causes both outcomes. In the hair/wrinkle example, old age is the confounding variable.
    • Another example: Wine drinkers may have better heart health than beer drinkers, but confounding variables (diet, lifestyle, socio-economic status) prevent concluding that switching to wine causes heart health.
  • Controlled Experiments:
    • Each individual is assigned either to the control group or the treatment group.
    • Sample sizes must be large enough to account for inherent variability.
    • Random Assignment: This is essential to minimize bias.
    • Controlled experiments can prove cause and effect.

Questions & Discussion

  • Fish Oil and Asthma Study: A medical journal reported a study on fish oil consumption in pregnant mothers and the subsequent development of asthma in their children.
    • Analysis: If mothers chose whether or not to take fish oil without researcher assignment, it is an observational study. If they were randomly assigned by researchers, it is a controlled experiment. Only a properly conducted controlled experiment would allow the conclusion that fish oil caused the lower asthma rate.
  • Salad/Vegetable Consumption Study: A study of 1,2261,226 older women over 1515 years found high vegetable consumption was associated with lower cardiovascular death risk.
    • Conclusion: We cannot conclude that the diet prevents the disease because this was an observational study. Confounding variables (e.g., overall healthier lifestyle, exercise habits) could be responsible.
  • Mice and Light Exposure Study (Baturin et al., 2001):
    • Data: 5050 mice in light/dark regimen (LD) and 5050 mice in 2424-hour light (LL).
    • Tumor counts: LD group = 44, LL group = 1414.
    • Calculations:
      • Percentage of LD group with tumors: 450×100%=8%\frac{4}{50} \times 100\% = 8\%.
      • Percentage of LL group with tumors: 1450×100%=28%\frac{14}{50} \times 100\% = 28\%.
    • Study Type: This was a controlled experiment because mice were randomly assigned to their regimens.
    • Causality: Because it was a controlled experiment with random assignment, researchers can conclude that 2424-hour light exposure causes an increase in tumors in these mice.