Categorical Data Representations, Contingency Tables, and Conditional Distributions
Categorical Data Representations and Misleading Graphs
Primary graphical displays for categorical data include pie charts and bar charts.
Visual Analysis Caution:
Graphs and charts can easily be visually misleading.
Data presented graphically must be critically analyzed to verify that visual representations accurately reflect underlying numeric values.
Example of a visually misleading graph:
A graph displaying run time ranges where a blue slice visually appears to be the largest section (occupying nearly a quarter or of the total display area).
Numeric analysis reveals that the yellow slice is actually the largest section at , or a slice labeled fails to match its visual proportions.
Frequency Distributions and Categorical Graphing Methods
Bar Charts:
Structure: Displays category names or labels on one axis (typically the x-axis) and frequency or relative frequency on the opposing axis (typically the y-axis).
Orientation: Can be rendered vertically or horizontally.
Structural Separation Rule: In a categorical bar chart, individual bars MUST NOT touch each other. This distinct gap differentiates bar charts from quantitative histograms, where bars touch.
Category Inclusion: Categorical datasets frequently incorporate an "Other" category to account for uncollected or remaining data points so that the total relative frequency equals .
Electricity Generation Example:
Categories include Coal, Hydroelectric, Nuclear, Natural Gas, Petroleum, and Other Renewable Sources.
Standard bar charts often require visual estimation of values relative to axis gridlines unless explicit numerical labels are provided.
Pie Charts:
Structure: Displays proportions as relative slices of a complete circle ().
Displays precise numerical percentages directly alongside each corresponding category slice (e.g., Coal at , Petroleum at ).
Desktop Conferencing Market Case Study
Dataset Context: A report detailing market share for desktop conferencing applications.
Provided Market Share Data:
Cisco Systems: share
Citrix: share
Microsoft: share
Constructing the Relative Frequency Table:
Sum of provided market shares:
Calculated "Other" app category:
Frequency Distribution Table:
Cisco: relative frequency (
Citrix: relative frequency (
Microsoft: relative frequency (
Other: relative frequency (
Total: (
Graph Construction Rules:
Bar Chart: Place market categories on the horizontal axis (Cisco, Citrix, Microsoft, Other). Set vertical scale up to at least (e.g., intervals of up to ). Ensure bars do not touch.
Pie Chart: Draw a circle with an approximate center point. Partition slices for Cisco (), Other (), Citrix (), and Microsoft ().
Examination Protocol Note: Exams consist of multiple-choice questions where manual graph construction is not required, but identifying correct graphs and reading graph data accurately is required.
Contingency Tables for Two Categorical Variables
Definition: A contingency table displays frequencies for two categorical variables simultaneously.
Structure:
Rows represent categories of one categorical variable.
Columns represent categories of the second categorical variable.
Cells contain joint frequencies (counts of cases meeting specific combined conditions of both variables).
Margins display row totals and column totals.
Proportions in Contingency Tables:
Joint Proportions (Relative Frequencies): Computed by dividing an individual cell frequency by the overall total sample size () of the entire table.
Conditional Proportions: Computed by restricting/conditioning analysis to a specific row or column. That specific row total or column total becomes the denominator.
Rounding Conventions:
Proportions: Round to decimal places (e.g., ).
Percentages: Round to decimal places (e.g., ).
Optometry Customer Dataset and Joint vs. Conditional Probabilities
Optometry Dataset Summary:
Collected from an optometry shop with variables Gender (Male, Female) and Eye Condition (Nearsighted, Farsighted, Needs Bifocals).
Raw Contingency Table Data:
Nearsighted Row: Male = , Female = , Row Total =
Farsighted Row: Male = , Female = , Row Total =
Needs Bifocals Row: Male = , Female = , Row Total =
Column Totals: Male Total = , Female Total = , Overall Sample Total () =
Step-by-Step Probability Calculations:
Question 1: What percent of all customers are farsighted?
Denominator / Condition: All customers ().
Numerator: Total farsighted customers ().
Proportion:
Percentage:
Question 2: What percent of nearsighted customers are females?
Denominator / Condition: Nearsighted customers only ().
Numerator: Female nearsighted customers ().
Proportion:
Percentage:
Question 3: What percent of all customers are nearsighted AND female?
Denominator / Condition: All customers ().
Numerator: Nearsighted females overlapping cell ().
Proportion:
Percentage:
Question 4: What percent of all customers are nearsighted OR female?
Denominator / Condition: All customers ().
Logical Operator Distinctions:
"AND" requires meeting both categories simultaneously.
"OR" includes meeting either category or both categories.
Double-Counting Warning: Adding total nearsighted () to total females () equals , which double-counts the nearsighted females.
Calculation Method 1 (Inclusion-Exclusion):
Calculation Method 2 (Sum of Non-Overlapping Cells):
Proportion:
Percentage:
Conditional Distributions by Category
Examination Calculator Requirements:
Calculators are required for exams to accurately compute decimals, proportions, and statistical functions.
Conditional Distribution of Eye Condition by Gender:
Conditioning by gender requires evaluating Male () and Female () columns independently as relative percentage distributions.
Male Conditional Distribution ():
Nearsighted:
Farsighted:
Needs Bifocals:
Column Sum Verification:
Female Conditional Distribution ():
Nearsighted:
Farsighted:
Needs Bifocals:
Column Sum Verification: