Week 2 (biostatistics) - graph and chart representations

Learning objectives

  • Identify data and choose appropriate graph
  • Describe the different summary statistics and rationalise what to report and when
  • Ability to interpret graphs and distributions

Graphs - Tell a Story

  • Graphs can contextualise data by showing trends and priorities over time or across groups
  • Example list of leading causes of death (from a public health perspective):
    • NHS areas: heart & circulatory disorders, cancer, respiratory disorders
    • Other topics: terrorism, war, pregnancy & birth, medical complications, murder, undetermined events, mental health disorders, transport accidents, suicide, musculoskeletal disorders, diabetes, non-transport accidents, infections, kidney disorders, digestive disorders, nervous system disorders

Injury in Australia (2011–12)

  • 454,000 Australians were hospitalised due to injury in 2011–12
  • Two main causes of injury: falls and transport accidents
  • Source: Australian Institute of Health and Welfare (AIHW) injury data

Road trauma in Victoria: timeline of legislative interventions

  • 1952–2012 era depicted in a timeline with milestone interventions
  • Key milestones and approximate years:
    • 1970: Compulsory seatbelt legislation introduced
    • 1976: Random breath testing introduced
    • 1983: Red light cameras introduced
    • 1989: Mobile speed cameras introduced
    • 1990: Booze buses introduced
    • 2000: Fixed speed cameras introduced
    • 2001: Victorian State Trauma System implemented
    • 2006: Random drug testing and impounding of hoons’ cars introduced
  • Additional notes: CT scans (x50k) referenced; a CT scan can take about 20 minutes for a full body scan; visual metaphor: "Donut of death → donut of truth" indicating interpretation of imaging data
  • Time axis spans 1952–2012; data source: Transport Accident Commission (TAC) and related Grand Round slides

Notes for TAC & CT Scan Graph (Trevor Fitzgerald talk, The Alfred Hospital)

  • Victorian Trauma Grand Round – 27 June 2017; overview of trauma care at The Alfred over the past 16 years
  • Highlights:
    • The road trauma graphs tell a story when milestones are indicated on the timeline
    • The rate of CT scans has increased as technology improves
    • Full body CT scan duration ~20 minutes
  • Donut metaphor: "Donut of death" vs. "donut of truth" for interpreting imaging data

Two cycling deaths in Victorian roads in 1 week (2015)

  • 17 June 2015: Died 29 June 2015; coverage in news outlets
  • Raises the question: how safe is cycling?

Notes for previous slide: implications of cycling deaths (2015)

  • June 2015: Two cycling deaths in one week
  • One victim: 17-year-old on a training ride
  • Other victim: wearing bright clothing, commuting to work as usual
  • Personal reflection on cycling safety and data interpretation

What data will be available? (Mortality and Morbidity)

  • Mortality: deaths while cycling
  • Morbidity: level of injury (serious, moderate, mild, no injury, near miss)
  • Data sources: death records, hospital records, doctor records, physiotherapy records, etc.

Raw mortality data for 2015: questions and utility

  • Question framing: what would you like to know?
  • Data usefulness: raw counts are not as useful as summary statistics for interpretation

MORTALITY - BY METHOD OF TRANSPORT AND YEAR (Australia, 1980–2013)

  • Graph type: All Road Fatalities for Australia (1980–2013)
  • Categories displayed: Drivers, Passengers, Pedestrians, Motorcyclists, Bicyclists
  • Data form: time series of fatalities by transport mode
  • Purpose: to illustrate trends and the contribution of each transport mode to total road fatalities over time

Campaigns for cyclist safety

  • Amy Gillett Foundation campaign: "A metre matters"
  • Message: motorists should give cyclists at least a metre when passing to reduce death/injury from multiple-vehicle crashes
  • Emphasises that behind every statistic is a person with grieving family and friends
  • Public health reminder: safety at home starts before you go on the road
  • Related campaign references: amygillett.org.au and Victoria’s Towards Zero program

Graphical & Statistical Presentation of Data (Overview)

  • Bar chart and Pie chart
  • Box-plot, Scatter plot, Histogram and Box-plot
  • Frequency tables, Relative frequencies
  • Summary statistics
  • Data types:
    • Categorical: Numerical, Ordinal, Nominal
    • Discrete vs Continuous
  • Distribution concepts:
    • Normal distribution: Mean, SD coincide (Mean = SD? actually Mean = Median = Mode in a perfect normal)
    • Skewed distributions: Left (Negatively) skewed and Right (Positively) skewed
  • Normal distribution relationships (as a guide):
    • For symmetric data: extMean=extMedian=extModeext{Mean} = ext{Median} = ext{Mode}
    • For left-skewed: extMeanextMedianextModeext{Mean} \le ext{Median} \le ext{Mode}
    • For right-skewed: extModeextMedianextMeanext{Mode} \le ext{Median} \le ext{Mean}

Frequency Tables (Example)

  • Example: Gender distribution among Monash University students (Nov 2015)
    • Categories: M = Male, F = Female, X = Indeterminate / Intersex / Unspecified, U = Unknown
    • Counts: M = 64,207; F = 89,806; X = 106; U = 2; Total = 154,121
    • Relative frequencies: M = 41.7%, F = 58.3%, X = 0.069%, U = 0.001%
    • Type of variable: categorical, nominal, NOT binary
  • Relative frequency formula (example):
    • ext{Relative frequency}{category} = rac{n{ ext{category}}}{N} imes 100 ext{ %}

Pie chart: gender distribution (Monash example)

  • Female: 58.3%; Male: 41.7%; Indeterminate / Intersex / Unspecified: 0.069%; Unknown: 0.001%

CROSS TABULATION OF TWO CATEGORICAL VARIABLES (example data)

  • Dataset: Respiratory symptoms in past 12 months (Yes/No) by Gender of child (Female/Male)
  • Table (example numbers):
    • Female: Yes = 81, No = 254, Total = 335
    • Male: Yes = 64, No = 237, Total = 301
    • Totals: Yes = 145, No = 491, Total = 636
  • Percentages shown in table:
    • Row percentages (by gender): Female: Yes 24.18%, No 75.82% (Total 100%); Male: Yes 21.26%, No 78.74% (Total 100%)
    • Column percentages (by symptom status): Yes column: Female 55.86%, Male 44.14%; No column: Female 51.73%, Male 48.27%
    • Overall totals: 145 Yes (22.80%), 491 No (77.20%), Total 636
  • Key questions:
    • What overall % are we interested in? e.g., % with respiratory symptoms: 145/636 = 22.80% (Row % perspective) and % by gender: 335/636 = 52.67% female (Column % perspective)
  • Interpretation guidance:
    • Row % answers: "What proportion of each gender had respiratory symptoms?"
    • Column % answers: "What proportion of those with/without symptoms are female?"
    • In this case, row % is often more informative for symptoms by gender; column % answers were used to compare gender distribution across symptom status
    • Conclusion: choose the percent type that matches your question

Repeated cross-tabulation discussion (interpreting percentages)

  • Repetition emphasizes two perspectives:
    • Row percentages: percent of the row (e.g., among each gender, what share have symptoms)
    • Column percentages: percent of the column (e.g., among those with/without symptoms, what share are female)
  • Practical takeaway: Always state which denominator you used when reporting percentages

Graphical & Statistical Presentation of Data (revisited)

  • Recap of types of graphs and data types
  • Common graphs by data type:
    • Categorical data: Bar charts, Pie charts
    • Numerical data: Scatter plots, Box plots, Histograms
  • Tables: summarise raw data; use frequencies and percentages (relative frequencies)
  • Numerical data: use summary statistics (mean, median, mode, SD, IQR, range)

Summary statistics (P. 21–25, P. 33–38)

  • Purpose: provides key information about the data; concise
  • Distribution of data: symmetric (normal) vs skewed
  • Central tendency:
    • Mean: extMean=1nextSumofallobservations=1n(<br/><em>ix</em>i)ext{Mean} = \frac{1}{n}\, ext{Sum of all observations} = \frac{1}{n}\,\bigg(<br />\nabla<em>i x</em>i\bigg)
  • Median: middle value; best for skewed data
  • Mode: most frequent value (or range)
  • Measures of dispersion:
    • Standard deviation (SD): s =
      oot extstyle{ rac{1}{n-1}\, extstyleig(igl(x_i - ar{x}igr)^2igr)}
    • Interquartile range (IQR): extIQR=Q<em>3Q</em>1ext{IQR} = Q<em>3 - Q</em>1
    • Range: difference between max and min

Histogram (data distribution example)

  • Example: Age group/class histogram with 636 observations
  • Data displayed as grouped into intervals (continuous data grouped into a frequency distribution)
  • Important note: bars in a histogram touch to indicate continuous data
  • Provided counts per interval (e.g., 7.00–7.25 years, etc.) culminating in Total = 636

Distribution of data and skewness recap

  • Normal distribution: symmetrical; mean = median = mode (in practice, approximately true for near-normal data)
  • Left-skewed (negatively skewed): extMeanextMedianextModeext{Mean} \le ext{Median} \le ext{Mode}
  • Right-skewed (positively skewed): extModeextMedianextMeanext{Mode} \le ext{Median} \le ext{Mean}

Box plot (Box-and-Whisker Plot)

  • Contains: Q1, Q2 (Median), Q3; upper and lower whiskers
  • Interquartile Range (IQR) = Q<em>3Q</em>1Q<em>3 - Q</em>1
  • Median = Q2 (middle value)
  • 25% of data in each quarter shown as the box boundaries
  • Whiskers extend to 1.5 x IQR beyond the quartiles
  • Outliers are values outside 1.5 x IQR from the quartiles; extreme values outside 3 x IQR

Box plot – key definitions (as in the slides)

  • 25% each region: lower whisker, box, upper whisker
  • IQR = upper quartile − lower quartile
  • Median divides the box into two halves
  • Outliers and extreme values identified by whisker rules

Summary statistics – Excel (2007) style (conceptual, not procedural code)

  • Mean: Xˉ=1nextRangesum\bar{X} = \frac{1}{n}\, ext{Range sum}
  • Median: value at middle rank
  • Standard Deviation (SD) for sample: s =
    \sqrt{\frac{1}{n-1}\sum{i=1}^n (xi - \bar{X})^2}
  • Quartiles: Q<em>1,Q</em>2,Q3Q<em>1, Q</em>2, Q_3 corresponding to 25th, 50th (Median), and 75th percentiles
  • Minimum and Maximum: range of data
  • Note: Many stats packages graph boxplots automatically; manual calculation not required for practical use

Graphical & Statistical Presentation – summary recap

  • Bar charts and pie charts for categorical data; scatter plots, box plots, histograms for numerical data
  • Frequency tables with relative frequencies; summary statistics (mean, SD, median, IQR, etc.)
  • Understanding distribution helps choose appropriate statistics (mean & SD for symmetric data; median & IQR for skewed data)

Lecture summary (key takeaways)

  • Graphs should tell a story and be able to stand alone
  • Categorical data: bar charts and pie charts; report frequencies and percentages
  • Numerical data: scatter plots, box plots, histograms; use appropriate summary statistics
  • Distribution guidance:
    • Symmetric distributions: use mean and SD
    • Skewed distributions: use median and IQR

Pre-tutorial video: Tables and summary statistics

  • Focus on types of data and corresponding summary statistics
  • Two categorical types of data to consider: Nominal and Ordinal; Numerical types: Discrete and Continuous
  • Resource links provided for further learning

Worked example datasets and interpretation (illustrative data)

  • Example data structure (Group, Consent, Gender, Shoulder ROM, Height):
    • Consent: n = 266 (96.0%); No consent: n = 11 (4.0%)
    • Gender: Male n = 131 (49.25%), Female n = 135 (50.75%)
    • Shoulder extension ROM: Mean = 65.4°, SD = 11.4, Range = 39° to 96°
    • Height: Mean = 170.3 cm, SD = 9.4 cm, Range = 145 cm to 193.2 cm
  • These illustrate reporting both counts and derived measures (means, SDs) for sample characteristics

Categorical data – Relative frequency example

  • Method of delivery for 600 babies: Normal 478 (79.7%), Forceps 65 (10.8%), Caesarean 57 (9.5%), Total 600 (100%)
  • Source: Essential Medical Statistics (Table 3.1)

Categorical data – Crosstab (summary concepts)

  • Crosstab tabulates counts in each combination of two categorical variables
  • Can compute:
    • Row percentages: percentages within each row (e.g., % ill by exposure within a row)
    • Column percentages: percentages within each column (e.g., % exposed among those with/without disease)
  • Example topic references: disease status by pesticide exposure (illustrative in the slides)

Numerical data – Descriptive statistics (continuous data)

  • Central tendency measures: mean, median, mode
  • Variation measures: SD, IQR, range
  • Choice depends on data distribution (normal vs skewed)
  • Note: The role of data checks and handling outliers is essential for credible statistics

Data checking and cleaning example

  • Real-world example: extreme BMI values (e.g., BMI > 48 kg/m^2) flagged as potential data entry errors
  • Process: verify height vs weight entries; correct misentries
  • Consequence of cleaning: improved plausibility of SD and other statistics

Types of graphs – practical guidance

  • Graphs should match data type and question:
    • Bar charts & Pie charts for categorical data
    • Scatter plots, Box plots, Histograms for numerical data
  • Understanding data type helps avoid misinterpretation and supports appropriate statistical summaries

End of pre-tutorial videos and notes

  • The slides are intended to prepare for table- and graph-based data interpretation in Monash Medicine, Nursing and Health Sciences
  • Emphasis on turning raw data into meaningful, reportable statistics

Quick reference: key formulas and definitions used in this deck

  • Mean: ar{X} = rac{1}{n}
    \sum{i=1}^n xi
  • Standard deviation (sample): s=1n1<em>i=1n(x</em>iXˉ)2s = \sqrt{\frac{1}{n-1}\sum<em>{i=1}^n (x</em>i - \bar{X})^2}
  • Interquartile range: extIQR=Q<em>3Q</em>1ext{IQR} = Q<em>3 - Q</em>1
  • Box-plot quartiles: Q<em>1,Q</em>2(Median),Q3Q<em>1, Q</em>2(\text{Median}), Q_3
  • 1.5 x IQR rule for whiskers and outliers:
    • Lower whisker: Q11.5×IQRQ_1 - 1.5\times\text{IQR}
    • Upper whisker: Q3+1.5×IQRQ_3 + 1.5\times\text{IQR}
    • Outlier: value outside [Q<em>11.5×IQR,Q</em>3+1.5×IQRQ<em>1 - 1.5\times\text{IQR}, Q</em>3 + 1.5\times\text{IQR}]
    • Extreme value: value outside [Q<em>13×IQR,Q</em>3+3×IQRQ<em>1 - 3\times\text{IQR}, Q</em>3 + 3\times\text{IQR}]
  • Relative frequency: extRF=ncategoryN×100%ext{RF} = \frac{n_{category}}{N} \times 100\%
  • Attack rate (illustrative epidemiology concept):
    • Attack rateexposed=aa+b\text{Attack rate}_{exposed} = \frac{a}{a+b} where a = ill and exposed, b = not ill but exposed
  • Relative risk (RR):
    • RR=Attack rate<em>exposedAttack rate</em>unexposed=a/(a+b)c/(c+d)\text{RR} = \frac{\text{Attack rate}<em>{exposed}}{\text{Attack rate}</em>{unexposed}} = \frac{a/(a+b)}{c/(c+d)} where c = ill and unexposed, d = not ill and unexposed

Note

  • The content above mirrors the structure and key concepts from the provided transcript, organized into comprehensive study notes with explicit emphasis on graph types, data types, and summary statistics, along with practical interpretation guidance and formulas where appropriate.