Week 2 - Data Visualization: Purpose, Chart Types, and Distribution Science
Learning Objectives for Data Visualization
Clear, Accurate, and Informative Visuals: Ability to identify the key elements that contribute to graphs being truthful and easily interpretable.
Selection and Interpretation of Charts: Developing the skill to select the most appropriate chart type based on the specific nature of the data (qualitative vs. quantitative) and the primary analytical goal.
Distribution Analysis: Describing and interpreting the shape and patterns of a distribution, specifically focusing on symmetry, skewness, modality (peaks), and the identification of unusual observations or outliers.
Purpose of Data Visualization
Data visualization serves distinct roles across analysis, communication, and management. These roles are categorized as follows:
Analysis & Research
Exploratory Data Analysis (EDA): Used to investigate patterns, trends, and anomalies within raw datasets.
Statistical Inference Support: Visualizing assumptions, relationships, and the results of tests used for scientific or statistical inference.
Communication & Insight Sharing
Communicating Insights: Translating complex technical findings into clear, accessible visuals tailored for different audiences.
Decision & Performance Management
Decision Support: Providing visual clarity to facilitate operational or strategic decision-making processes.
Monitoring & Tracking: Continuously following key metrics to ensure organizational alignment with specific goals or Key Performance Indicators (KPIs).
Key Elements of Effective Data Visualization
Accuracy: The visual must represent the data truthfully without distorting the underlying facts.
Consistency: Uniformity must be maintained in scales, labels, and units of measurement throughout the graphic.
Clear Labels: Axes, units of measurement, legends, and titles must be informative and explicitly stated.
Highlighting Key Insights: The design should naturally draw the viewer's attention to significant patterns or outliers.
Strategic Use of Color:
Categorical: Used to distinguish between different types of data.
Sequential: Using a gradient of colors to show contrast or magnitude.
Diverging: Used to highlight specific values or differences from a midpoint.
Monochromatic: Variations of lightness within a single hue.
Data Transformation: From Raw Tables to Visual Insights
The goal of visualization is to transform raw spreadsheet data into actionable insights.
Raw Data Sample (STR Hotel Data):
Hotel 110298: Des Voyageurs Hotel, Rue Grand St Jean 19, Lausanne (VD), 1003 Lake Geneva, Independent, 33 rooms.
Hotel 110299: Comfort Hotel Royal Zurich, Leonhardstrasse 6, Zurich (ZH), 1801 North-East Switzerland, Independent, 70 rooms.
Hotel 110826: Grand Hotel National Lucerne, Haldenstrasse 4, Lucerne (LU), 6006 Central Switzerland, Independent, 41 rooms.
Hotel 110827: Park Hotel Vitznau, Seestrasse 18, Vitznau (LU), 6354 Central Switzerland, Independent, 47 rooms.
Hotel 110829: Le Mirador Resort & Spa, Chemin Du Mirador 5, Mont Pelerin (VD), Lake Geneva Region, Independent, 60 rooms.
Hotel 110830: Suvretta House St Moritz, Via Chasellas 1, St Moritz (GR), 7500 South-East Switzerland, Independent, 181 rooms.
Hotel 110831: Grand Hotel Zermatterhof, Bahnhofstrasse 55, Zermatt (VS), 3920 South-West Switzerland, Independent, 77 rooms.
Hotel 110832: Baur Au Lac, Talstrasse 1, Zurich (ZH), 8001 North-East Switzerland, Independent, 119 rooms.
Hotel 110861: Moevenpick Hotel Lausanne, avenue de Rhodanie 4, Lausanne (VD), 1007 Lake Geneva, Moevenpick, 337 rooms.
Visual Representation (Swiss Hotels by Location):
Urban:
Small Metro/Town:
Suburban:
Resort:
Interstate:
Airport:
Pie Charts
Usage: Works specifically with nominal (categorical) variables.
Purpose: To show parts as a proportion of a whole, such as market shares.
Best Practices:
Limit to a maximum of 5 categories to maintain readability.
Do not use multiple pie charts for comparison.
Ensure the sum of all slices equals exactly .
Order slices logically (e.g., clockwise by size, or place the two most important slices at the top).
Bar Plots and Stacked Bar Plots
Bar Plot Usage: Works with qualitative (categorical) variables or quantitative discrete variables.
Bar Plot Best Practices:
Use horizontal labels for better legibility.
Space bars appropriately (not touching).
Always start the y-axis (or value axis) at .
Use consistent colors and order the data logically (e.g., descending frequency).
Stacked Bar Plot: Shows the relationship between two categorical variables. For example, it can show the count of locations broken down by operation type (Chain Management, Franchise, Independent).
Stacked Bar Plot: Normalizes the bars so each total is . This is used to compare the relative proportions of sub-categories within each primary category across a group.
Histograms and Distribution Shapes
Histogram Usage: Works best with continuous variables. It can also work with discrete variables if they are grouped into classes.
Process: You must define classes (bins) to extract meaningful insights from the data density.
Characteristics of Distribution:
Symmetry: Indicates if values are balanced or skewed. For instance, a right-skewed distribution in exam grades might suggest the test was relatively easy (more high scores).
Peak (Modality): Represents the most frequent value (the mode). A peak in an age histogram shows the most represented age group.
Shape: Reveals overall patterns. A single peak (unimodal) suggests one main group, while multiple peaks (multimodal) may indicate distinct subgroups in the data. Long tails or unusual values (outliers) are also identified through shape analysis.
Boxplots
Definition: A graphical representation of the "Five Number Summary":
Minimum
Quartile 1 ()
Quartile 2 ( or Median)
Quartile 3 ()
Maximum
Calculating Elements:
Interquartile Range: . This range contains the middle of the data values.
Lower Fence: .
Upper Fence: .
Whiskers: Extend from the box to the minimum and maximum values located before reaching the fences.
Outliers: Data points falling outside the fences (Minimum or Maximum values labeled as outliers).
Usage: Specifically for quantitative variables; excellent for detecting anomalies (outliers) and describing the relationship between one quantitative variable and one qualitative variable.
Line Plots and Scatterplots
Line Plot:
Works primarily with time series data.
Example: Tracking the average occupancy rates (e.g., ITA-Occupancy and FR-Occupancy) over time (years 2000 through 2016).
Scatterplot:
Used to visualize the relationship between two quantitative variables.
Fundamental for discussions on correlation and regression.
Best Practices: Enrich the visual by using different colors for additional variables and include trend lines (limit to 2 trend lines to avoid clutter).
Variable Summary Table
Number of Variables | Variable Type | Tables to Use | Graphs to Use |
|---|---|---|---|
1 Variable | Categorical | Frequency table (absolute, relative, cumulative) | Bar chart, Pie chart |
1 Variable | Numerical | Frequency table [ungrouped or grouped] (absolute, relative, cumulative) | Histogram, Boxplot |
2 Variables | Categorical + Categorical | Contingency table | Stacked bar chart (or ), Side-by-side bar chart |
2 Variables | Numerical + Numerical | Correlation summary | Scatter plot |
2 Variables | Categorical + Numerical | Group by category frequency tables | Boxplot by category |
Practice Exercises
Exercise 1: Analyze a provided graph indicating restaurant complaints to identify which causes contribute to of all issues.
Exercise 2: Construct a Histogram and a Boxplot from a random sample of 11 wine prices: . Discuss the difference in approach regarding class range and class frequency.
Exercise 3: Comparison of calories across similar food items (research-based activity).