FYE110 Reasoning with Data - Unit 1: Developing Data Study Guide

Core Fundamentals of Data and Statistics

  • Ubiquity of Statistics: Statistical data permeates modern society and professional disciplines, appearing across social media platforms, news articles, corporate reports, meteorological forecasting, clinical medical trials, athletic tracking, demographic surveys, and commercial product labels.

  • Fields Utilizing Statistical Data:

    • Health: Clinical trial outcomes, epidemiological disease tracking, and treatment efficacy metrics.

    • Meteorology: Weather forecasting models, atmospheric pressure, and temperature trend monitoring.

    • Environment & Sustainability: Carbon footprint measurement, municipal waste management tracking, and resource conservation statistics.

    • Sports: Player performance metrics, game outcome probabilities, and athletic training analytics.

    • Education: Standardized testing results, graduation rate tracking, and institutional performance assessment.

    • Economics: Inflation rates, Gross Domestic Product (GDP\text{GDP}), labor market indicators, and financial index modeling.

  • Verbatim Definition of Statistics: Statistics is the art and science of designing studies and analyzing the data that those studies produce. Its ultimate goal is translating data into knowledge and understanding of the world around us. In short, statistics is the art and science of learning from data.

  • Three Core Pillars of Statistical Methodology:

    1. Design: Formulating the research plan, specifying sampling methods, and executing data collection protocols.

    2. Describe: Organizing, summarizing, tabulating, and visually displaying collected data.

    3. Inference: Formulating conclusions, predictions, or decision models about a wider population based on sample data.

Key Concepts: Populations, Samples, Parameters, and Statistics

  • Population: The complete, exhaustive group of individuals, subjects, or objects about which information or knowledge is sought.

  • Sample: The specific sub-collection or part of the population that is actually observed, measured, and analyzed.

  • Parameter: A specific numerical summary or description of a population characteristic. Parameters are fixed values but are typically unknown unless a full census is conducted.

  • Statistic: A specific numerical summary or description of a sample characteristic. Statistics are calculated from sample data and serve as estimates for unknown population parameters.

  • Structural Relationship:

    • Population characteristics correspond to Parameters (P\text{P} pairs with P\text{P}).

    • Sample characteristics correspond to Statistics (S\text{S} pairs with S\text{S}).

Population and Sample

Individuals and Variables in Data Structures

  • Individual (Subject): A person, entity, or object about which data is collected.

  • Variable: A specific attribute, characteristic, or measurement recorded for each individual.

  • Tabular Data Representation:

    • Rows: Represent distinct Individuals (e.g., individual patients or survey respondents).

    • Columns: Represent distinct Variables recorded across those individuals.

Variable and Individual in a dataframe
  • Data Matrix Layout Breakdown:

    • Row entries represent individual subjects (e.g., Patient #1\#1, Patient #2\#2, Patient #3\#3, …, Patient #75\#75).

    • Column headers define the measured variables:

    • Gender: Categorical demographic (M / F\text{M / F}).

    • Age: Expressed in years.

    • Weight: Expressed in pounds (lbs.\text{lbs.}).

    • Height: Expressed in inches (in.\text{in.}).

    • Smoking Status: Binary coded variable (0=No0 = \text{No}, 1=Yes1 = \text{Yes}).

    • Race: Categorical demographic (e.g., White, Black, Asian).

Core Branches: Descriptive vs. Inferential Statistics

  • Descriptive Statistics: Summarize and organize data using numerical measures (e.g., averages, proportions), tables, or graphical representations (e.g., scatter plots, bar charts, histograms).

  • Inferential Statistics: Utilize sample data to draw general conclusions, make predictions, or guide decision-making about an entire population while accounting for sampling variability.

Case Studies and Statistical Analysis

  • Case Study 1: Internet Habits Among UAE Students:

    • Context: Research published in the Intercontinental Journal of Social Sciences and reported in The National (June 01, 2024).

    • Sample Size: 600 university students600\text{ university students} in the UAE (including male and female participants).

    • Findings: UAE university students spend approximately the equivalent of a full working day online daily, with 84%84\% spending more than 7 hours7\text{ hours} online each day.

    • Applied Structural Breakdown:

    • Population: All university students in the UAE.

    • Sample: The 600 surveyed university students600\text{ surveyed university students}.

    • Parameter: The true proportion of all UAE university students who are online for more than 7 hours7\text{ hours} daily.

    • Statistic: The sample proportion of 84%84\% who reported being online for more than 7 hours7\text{ hours} daily.

    • Individual: Any single university student who participated in the survey.

    • Variables: Survey responses collected per participant, including age, major, and daily online time (e.g., responses to "How much time do you spend online daily?").

    • Design: Conducting the survey among the 600 university students600\text{ university students} in the UAE.

    • Describe: Summarizing data using reported daily online averages and proportions.

    • Infer: Generalizing sample findings to all university students across the UAE.

  • Case Study 2: Technology and Global Warming Attitudes:

    • Context: Survey conducted by duke+mir ahead of COP28 via YouGov.

    • Sample Size: Over 1,000 residents1{,}000\text{ residents} living in the UAE.

    • Findings:

    • 75%75\% of respondents think humans will find technology to solve global warming.

    • 52%52\% believe climate change is inevitable.

    • 50%50\% (half) of UAE residents recycle regularly.

  • Case Study 3: UAE Traffic and Road Rage:

    • Context: News article in The National (July 15, 2025).

    • Statistic: 8 in 108\text{ in }10 (80%80\%) drivers in the UAE experience road rage amidst increasing road traffic.

Practical Activity Analyses & Critiques

  • Analysis of Global Warming Attitudes Survey Breakdown:

    1. Individual / Subject: A single UAE resident who answered the survey.

    2. Variable: An individual's opinion on whether humans will find technology to solve global warming.

    3. Population: All people living in the UAE.

    4. Sample: The 1,000+ UAE residents1{,}000+\text{ UAE residents} surveyed through YouGov.

    5. Parameter: The true proportion of all UAE residents who think humans will find technology to solve global warming.

    6. Statistic: The 75%75\% sample proportion who expressed that opinion.

  • Critique and Correction of 2023 UAE National Reading Index Report:

    • Context: Report released on August 17, 2024, by the Ministry of Culture, surveying a national sample of approximately 3,900 citizens and residents3{,}900\text{ citizens and residents}, finding that 90.4%90.4\% use social media sites.

    • Analysis of Misconceptions and Correct Identifications:

    1. Claim: The population is the 3,900 citizens and residents3{,}900\text{ citizens and residents} contacted.

      • Correction: The population is all citizens and residents in the UAE. The 3,9003{,}900 surveyed individuals represent the sample.

    2. Claim: The sample is the 90.4%90.4\% who use social media.

      • Correction: The sample is the 3,900 citizens and residents3{,}900\text{ citizens and residents} surveyed. The 90.4%90.4\% value is a sample statistic.

    3. Claim: The variable is the 90.4%90.4\% who use social media.

      • Correction: The variable is whether a person uses social media (Categorical Yes/No response). The 90.4%90.4\% is the summarized proportion.

    4. Claim: The subjects in this survey are the social media sites.

      • Correction: The subjects/individuals are the people (citizens and residents) surveyed, not the websites or social platforms.

    5. Claim: The parameter consists of all citizens and residents in the UAE.

      • Correction: The parameter is the true proportion of all UAE citizens and residents who use social media (a numerical value summarizing the population).

    6. Claim: The statistic is the average number of citizens and residents who use social media sites.

      • Correction: The statistic is the sample proportion (90.4%90.4\%) of surveyed participants who use social media.

The Data Science Investigation Process: The PPDAC Cycle

  • PPDAC Model Overview: A structured framework for statistical investigations used to abstract real-world problems into statistical problems and generate data-informed solutions.

Cycle of Data Processing
  • Five Stages of the PPDAC Framework:

    1. Problem: Define the fundamental research question, specify objectives, and set the scope.

    2. Plan: Establish data collection methods, sampling protocols, measurement procedures, and tools.

    3. Data: Execute collection to gather relevant and accurate data following the plan.

    4. Analysis: Process, organize, graph, and examine data to identify patterns, associations, or anomalies.

    5. Conclusion: Interpret analytical results in context, address original questions, note limitations, and provide evidence-based recommendations.

Applied PPDAC Example: Height and Arm Span Relationship

  • Problem: Determine whether a statistical relationship exists between a student's height and their arm span.

  • Plan: Measure students' height and arm span in centimeters (cm\text{cm}) using a tape measure. For arm span, students stand facing a wall with arms stretched out horizontally while a partner measures fingertip-to-fingertip distance.

  • Data: Tabulated pairs of measured values:

    • Student 1: Height = 159 cm159\text{ cm}, Arm Span = 160 cm160\text{ cm}

    • Student 2: Height = 163 cm163\text{ cm}, Arm Span = 159 cm159\text{ cm}

    • Student 3: Height = 173 cm173\text{ cm}, Arm Span = 177 cm177\text{ cm}

    • Student 4: Height = 165 cm165\text{ cm}, Arm Span = 167 cm167\text{ cm}

    • Student 5: Height = 163 cm163\text{ cm}, Arm Span = 163 cm163\text{ cm}

    • Student 6: Height = 180 cm180\text{ cm}, Arm Span = 176 cm176\text{ cm}

  • Analysis: Construct a scatter plot of Height (cm\text{cm}) versus Arm Span (cm\text{cm}).

Height vs Arm Span
  • Analytical Observations: The plot shows a strong positive linear trend (taller students tend to have longer arm spans). One extreme data point (Height = 180 cm180\text{ cm}, Arm Span = 200 cm200\text{ cm}) acts as an outlier, indicating a potential measurement error.

    • Conclusion: Height is positively associated with arm span among students; taller individuals generally exhibit larger arm spans, while shorter individuals exhibit smaller arm spans.

Taxonomy of Data Types and Levels of Measurement

Data Types and Their Levels of Measurement
  • Qualitative (Categorical) Variables: Variables that represent non-numeric categories or labels. Categories are mutually exclusive, meaning each observation belongs to one and only one category.

    • Nominal Level: Categories with no intrinsic order or ranking.

    • Examples: Gender, Race, Nationality, Technology opinion (Yes / No\text{Yes / No}).

    • Ordinal Level: Categories that display a natural or logical order, though numerical distances between ranks cannot be measured.

    • Examples: Frequency of recycling (Regular, Sometimes, Never), Educational level (High School, Bachelor, Master, PhD), Income brackets (Low, Medium, High).

  • Quantitative (Numerical) Variables: Variables that take on numeric values where arithmetic operations (such as addition, subtraction, and averaging) are meaningful.

    • Discrete Level: Numeric values with distinct gaps or countable values (often whole numbers).

    • Examples: Number of times recycled per month, household count, patient volume.

    • Continuous Level: Numeric values that can take any real value within a continuous range or interval.

    • Examples: Age (years), Height (cm\text{cm}), Weight (lbs\text{lbs}), Time spent online (hours).

  • Applied Variable Classification Examples:

    • UAE Climate Survey Analysis:

    • Opinion on technology solution (Yes / No\text{Yes / No}): Qualitative →\rightarrow Nominal

    • Recycling frequency (Regular / Sometimes / Never): Qualitative →\rightarrow Ordinal

    • Age of respondent (years): Quantitative →\rightarrow Continuous

    • Recycling instances per month: Quantitative →\rightarrow Discrete

    • Mediterranean Diet Adherence Study Analysis:

    • Gender: Qualitative →\rightarrow Nominal

    • Age: Quantitative →\rightarrow Continuous

    • Nationality: Qualitative →\rightarrow Nominal

    • Educational Level: Qualitative →\rightarrow Ordinal

    • Income: Quantitative →\rightarrow Continuous

    • Height: Quantitative →\rightarrow Continuous

    • Weight: Quantitative →\rightarrow Continuous

Information Sources and Data Collection Methods

  • Methodological Flexibility: Most data collection techniques are applicable across both qualitative and quantitative research paradigms. Methodological differences stem from structured restrictions, procedural flexibility, sequence, and analytical depth.

  • Primary Data: Firsthand information collected directly by researchers from original sources for the specific purpose of the study.

    • Interviews: Direct, interactive, in-depth questioning to collect detailed qualitative or structured feedback.

    • Observation: Systematic monitoring and logging of behaviors or physical events.

    • Questionnaires / Surveys: Formally structured question sets distributed to participants.

    • Primary Data Example: The 20232023 UAE National Reading Index, where researchers gathered demographics, reading habits, and social media habits directly from 3,900 citizens and residents3{,}900\text{ citizens and residents} via questionnaires.

  • Secondary Data: Pre-existing information collected previously by external entities or researchers for another purpose, repurposed for current research.

    • Existing Records and Documents: Published books, academic articles, and institutional reports.

    • Official Statistics: Government archives, public registries, and census data.

    • Databases: Online research repositories and open-data platforms.

    • Secondary Data Example: The duke+mir climate survey using historical government recycling statistics to compare UAE performance against UK recycling rates.

Data Formation, Structuring, and Digital Formats

  • Converting Information into Data: The systematic process of transforming qualitative or unstructured information into a structured, machine-readable format suitable for computational processing, analysis, and modeling.

  • Steps of Information-to-Data Transformation:

    1. Identification: Target specific information elements required for analysis.

    2. Extraction: Isolate relevant details from source materials.

    3. Structuring: Organize extracted details into structured matrices, tables, spreadsheets, or relational databases.

    4. Quantification: Convert non-numeric information into numerical scale metrics where possible (e.g., rating satisfaction on a scale from 1 to 101\text{ to }10).

    5. Encoding: Format structured quantitative and qualitative values into standard machine-readable computer code.

  • Digital Data File Formats:

Samples of File Formats Used to Store Data
  • Categories of File Formats:

    • Text Formats:

    • Unformatted Text (.TXT)

    • Tabular Comma-Separated Values (.CSV)

    • JavaScript Object Notation (.JSON): Lightweight, human-readable, and machine-parseable structured data interchange format.

    • Binary Formats:

    • Multimedia image, audio, and video formats (.JPG, .PNG, .MPG, .AVI, .EPS) analyzed via specialized statistical and computer vision libraries.

    • Compressed File Archives (.ZIP, .RAR).

    • Document Formats:

    • Word Documents (.DOC), Portable Document Format (.PDF), Presentation Slides (.PPT).

    • Database & Spreadsheet Formats:

    • Spreadsheets (.XLS / Excel).

    • Relational SQL Databases and Scalable NoSQL Databases designed for large-scale data architecture.