FYE110 Reasoning with Data - Unit 1: Developing Data Study Guide
Core Fundamentals of Data and Statistics
Ubiquity of Statistics: Statistical data permeates modern society and professional disciplines, appearing across social media platforms, news articles, corporate reports, meteorological forecasting, clinical medical trials, athletic tracking, demographic surveys, and commercial product labels.
Fields Utilizing Statistical Data:
Health: Clinical trial outcomes, epidemiological disease tracking, and treatment efficacy metrics.
Meteorology: Weather forecasting models, atmospheric pressure, and temperature trend monitoring.
Environment & Sustainability: Carbon footprint measurement, municipal waste management tracking, and resource conservation statistics.
Sports: Player performance metrics, game outcome probabilities, and athletic training analytics.
Education: Standardized testing results, graduation rate tracking, and institutional performance assessment.
Economics: Inflation rates, Gross Domestic Product (), labor market indicators, and financial index modeling.
Verbatim Definition of Statistics: Statistics is the art and science of designing studies and analyzing the data that those studies produce. Its ultimate goal is translating data into knowledge and understanding of the world around us. In short, statistics is the art and science of learning from data.
Three Core Pillars of Statistical Methodology:
Design: Formulating the research plan, specifying sampling methods, and executing data collection protocols.
Describe: Organizing, summarizing, tabulating, and visually displaying collected data.
Inference: Formulating conclusions, predictions, or decision models about a wider population based on sample data.
Key Concepts: Populations, Samples, Parameters, and Statistics
Population: The complete, exhaustive group of individuals, subjects, or objects about which information or knowledge is sought.
Sample: The specific sub-collection or part of the population that is actually observed, measured, and analyzed.
Parameter: A specific numerical summary or description of a population characteristic. Parameters are fixed values but are typically unknown unless a full census is conducted.
Statistic: A specific numerical summary or description of a sample characteristic. Statistics are calculated from sample data and serve as estimates for unknown population parameters.
Structural Relationship:
Population characteristics correspond to Parameters ( pairs with ).
Sample characteristics correspond to Statistics ( pairs with ).

Individuals and Variables in Data Structures
Individual (Subject): A person, entity, or object about which data is collected.
Variable: A specific attribute, characteristic, or measurement recorded for each individual.
Tabular Data Representation:
Rows: Represent distinct Individuals (e.g., individual patients or survey respondents).
Columns: Represent distinct Variables recorded across those individuals.

Data Matrix Layout Breakdown:
Row entries represent individual subjects (e.g., Patient , Patient , Patient , …, Patient ).
Column headers define the measured variables:
Gender: Categorical demographic ().
Age: Expressed in years.
Weight: Expressed in pounds ().
Height: Expressed in inches ().
Smoking Status: Binary coded variable (, ).
Race: Categorical demographic (e.g., White, Black, Asian).
Core Branches: Descriptive vs. Inferential Statistics
Descriptive Statistics: Summarize and organize data using numerical measures (e.g., averages, proportions), tables, or graphical representations (e.g., scatter plots, bar charts, histograms).
Inferential Statistics: Utilize sample data to draw general conclusions, make predictions, or guide decision-making about an entire population while accounting for sampling variability.
Case Studies and Statistical Analysis
Case Study 1: Internet Habits Among UAE Students:
Context: Research published in the Intercontinental Journal of Social Sciences and reported in The National (June 01, 2024).
Sample Size: in the UAE (including male and female participants).
Findings: UAE university students spend approximately the equivalent of a full working day online daily, with spending more than online each day.
Applied Structural Breakdown:
Population: All university students in the UAE.
Sample: The .
Parameter: The true proportion of all UAE university students who are online for more than daily.
Statistic: The sample proportion of who reported being online for more than daily.
Individual: Any single university student who participated in the survey.
Variables: Survey responses collected per participant, including age, major, and daily online time (e.g., responses to "How much time do you spend online daily?").
Design: Conducting the survey among the in the UAE.
Describe: Summarizing data using reported daily online averages and proportions.
Infer: Generalizing sample findings to all university students across the UAE.
Case Study 2: Technology and Global Warming Attitudes:
Context: Survey conducted by duke+mir ahead of COP28 via YouGov.
Sample Size: Over living in the UAE.
Findings:
of respondents think humans will find technology to solve global warming.
believe climate change is inevitable.
(half) of UAE residents recycle regularly.
Case Study 3: UAE Traffic and Road Rage:
Context: News article in The National (July 15, 2025).
Statistic: () drivers in the UAE experience road rage amidst increasing road traffic.
Practical Activity Analyses & Critiques
Analysis of Global Warming Attitudes Survey Breakdown:
Individual / Subject: A single UAE resident who answered the survey.
Variable: An individual's opinion on whether humans will find technology to solve global warming.
Population: All people living in the UAE.
Sample: The surveyed through YouGov.
Parameter: The true proportion of all UAE residents who think humans will find technology to solve global warming.
Statistic: The sample proportion who expressed that opinion.
Critique and Correction of 2023 UAE National Reading Index Report:
Context: Report released on August 17, 2024, by the Ministry of Culture, surveying a national sample of approximately , finding that use social media sites.
Analysis of Misconceptions and Correct Identifications:
Claim: The population is the contacted.
Correction: The population is all citizens and residents in the UAE. The surveyed individuals represent the sample.
Claim: The sample is the who use social media.
Correction: The sample is the surveyed. The value is a sample statistic.
Claim: The variable is the who use social media.
Correction: The variable is whether a person uses social media (Categorical Yes/No response). The is the summarized proportion.
Claim: The subjects in this survey are the social media sites.
Correction: The subjects/individuals are the people (citizens and residents) surveyed, not the websites or social platforms.
Claim: The parameter consists of all citizens and residents in the UAE.
Correction: The parameter is the true proportion of all UAE citizens and residents who use social media (a numerical value summarizing the population).
Claim: The statistic is the average number of citizens and residents who use social media sites.
Correction: The statistic is the sample proportion () of surveyed participants who use social media.
The Data Science Investigation Process: The PPDAC Cycle
PPDAC Model Overview: A structured framework for statistical investigations used to abstract real-world problems into statistical problems and generate data-informed solutions.

Five Stages of the PPDAC Framework:
Problem: Define the fundamental research question, specify objectives, and set the scope.
Plan: Establish data collection methods, sampling protocols, measurement procedures, and tools.
Data: Execute collection to gather relevant and accurate data following the plan.
Analysis: Process, organize, graph, and examine data to identify patterns, associations, or anomalies.
Conclusion: Interpret analytical results in context, address original questions, note limitations, and provide evidence-based recommendations.
Applied PPDAC Example: Height and Arm Span Relationship
Problem: Determine whether a statistical relationship exists between a student's height and their arm span.
Plan: Measure students' height and arm span in centimeters () using a tape measure. For arm span, students stand facing a wall with arms stretched out horizontally while a partner measures fingertip-to-fingertip distance.
Data: Tabulated pairs of measured values:
Student 1: Height = , Arm Span =
Student 2: Height = , Arm Span =
Student 3: Height = , Arm Span =
Student 4: Height = , Arm Span =
Student 5: Height = , Arm Span =
Student 6: Height = , Arm Span =
Analysis: Construct a scatter plot of Height () versus Arm Span ().

Analytical Observations: The plot shows a strong positive linear trend (taller students tend to have longer arm spans). One extreme data point (Height = , Arm Span = ) acts as an outlier, indicating a potential measurement error.
Conclusion: Height is positively associated with arm span among students; taller individuals generally exhibit larger arm spans, while shorter individuals exhibit smaller arm spans.
Taxonomy of Data Types and Levels of Measurement

Qualitative (Categorical) Variables: Variables that represent non-numeric categories or labels. Categories are mutually exclusive, meaning each observation belongs to one and only one category.
Nominal Level: Categories with no intrinsic order or ranking.
Examples: Gender, Race, Nationality, Technology opinion ().
Ordinal Level: Categories that display a natural or logical order, though numerical distances between ranks cannot be measured.
Examples: Frequency of recycling (Regular, Sometimes, Never), Educational level (High School, Bachelor, Master, PhD), Income brackets (Low, Medium, High).
Quantitative (Numerical) Variables: Variables that take on numeric values where arithmetic operations (such as addition, subtraction, and averaging) are meaningful.
Discrete Level: Numeric values with distinct gaps or countable values (often whole numbers).
Examples: Number of times recycled per month, household count, patient volume.
Continuous Level: Numeric values that can take any real value within a continuous range or interval.
Examples: Age (years), Height (), Weight (), Time spent online (hours).
Applied Variable Classification Examples:
UAE Climate Survey Analysis:
Opinion on technology solution (): Qualitative Nominal
Recycling frequency (Regular / Sometimes / Never): Qualitative Ordinal
Age of respondent (years): Quantitative Continuous
Recycling instances per month: Quantitative Discrete
Mediterranean Diet Adherence Study Analysis:
Gender: Qualitative Nominal
Age: Quantitative Continuous
Nationality: Qualitative Nominal
Educational Level: Qualitative Ordinal
Income: Quantitative Continuous
Height: Quantitative Continuous
Weight: Quantitative Continuous
Information Sources and Data Collection Methods
Methodological Flexibility: Most data collection techniques are applicable across both qualitative and quantitative research paradigms. Methodological differences stem from structured restrictions, procedural flexibility, sequence, and analytical depth.
Primary Data: Firsthand information collected directly by researchers from original sources for the specific purpose of the study.
Interviews: Direct, interactive, in-depth questioning to collect detailed qualitative or structured feedback.
Observation: Systematic monitoring and logging of behaviors or physical events.
Questionnaires / Surveys: Formally structured question sets distributed to participants.
Primary Data Example: The UAE National Reading Index, where researchers gathered demographics, reading habits, and social media habits directly from via questionnaires.
Secondary Data: Pre-existing information collected previously by external entities or researchers for another purpose, repurposed for current research.
Existing Records and Documents: Published books, academic articles, and institutional reports.
Official Statistics: Government archives, public registries, and census data.
Databases: Online research repositories and open-data platforms.
Secondary Data Example: The duke+mir climate survey using historical government recycling statistics to compare UAE performance against UK recycling rates.
Data Formation, Structuring, and Digital Formats
Converting Information into Data: The systematic process of transforming qualitative or unstructured information into a structured, machine-readable format suitable for computational processing, analysis, and modeling.
Steps of Information-to-Data Transformation:
Identification: Target specific information elements required for analysis.
Extraction: Isolate relevant details from source materials.
Structuring: Organize extracted details into structured matrices, tables, spreadsheets, or relational databases.
Quantification: Convert non-numeric information into numerical scale metrics where possible (e.g., rating satisfaction on a scale from ).
Encoding: Format structured quantitative and qualitative values into standard machine-readable computer code.
Digital Data File Formats:

Categories of File Formats:
Text Formats:
Unformatted Text (
.TXT)Tabular Comma-Separated Values (
.CSV)JavaScript Object Notation (
.JSON): Lightweight, human-readable, and machine-parseable structured data interchange format.Binary Formats:
Multimedia image, audio, and video formats (
.JPG,.PNG,.MPG,.AVI,.EPS) analyzed via specialized statistical and computer vision libraries.Compressed File Archives (
.ZIP,.RAR).Document Formats:
Word Documents (
.DOC), Portable Document Format (.PDF), Presentation Slides (.PPT).Database & Spreadsheet Formats:
Spreadsheets (
.XLS/ Excel).Relational SQL Databases and Scalable NoSQL Databases designed for large-scale data architecture.