Bivariate Data and Scatter Plots Study Guide
Learning Objectives for Bivariate Data and Scatter Plots
Identify bivariate data, which involves the study of two variables simultaneously.
Distinguish between independent and dependent variables within a bivariate data set.
Construct scatter plots to visualize and compare data from two variables.
Analyze scatter plots to determine the form, direction, and strength of the correlation between variables.
Define and identify outliers within a data set.
Utilize specific mathematical terminology to describe associations between variables.
Fundamental Concepts of Bivariate Data
Definition of Bivariate Data: This type of data includes measurements or observations for two variables collected from the same entity. These variables are typically related in some way, such as a person's height and their weight.
Independent Variable: Often referred to as the explanatory variable, this is the variable that is changed, controlled, or used to make predictions. It is systematically plotted on the horizontal -axis.
Dependent Variable: Often referred to as the response variable, this is the variable being measured, tested, or observed. Its value depends on the independent variable, and it is plotted on the vertical -axis.
Terminology for Relationships: The terms "relationship," "correlation," and "association" are used interchangeably to describe the nature and degree to which two variables are related.
Characteristics of Scatter Plots
Scatter Plot Construction: A scatter plot is a graph drawn on a number plane (Cartesian plane) where the two axes represent the two different variables. Each pair of data values is represented by a single point .
Direction of Correlation:
Positive Correlation: As the independent variable increases, the dependent variable also tends to increase.
Negative Correlation: As the independent variable increases, the dependent variable tends to decrease.
Strength of Correlation:
Strong: The data points fall very close to a straight line or a clear path, indicating a high degree of relationship.
Weak: The data points show a general trend but are more widely scattered, indicating a lower degree of relationship.
No Correlation: There is no apparent relationship between the variables; the points appear randomly distributed.
Outliers: An outlier is a data point that is clearly isolated or significantly distant from the rest of the data clusters. It represents an observation that deviates from the established pattern of the relationship.
Qualitative Analysis of Variable Relationships
When evaluating potential relationships between two variables, it is necessary to consider if one might influence the other and in what direction.
Height of person and Weight of person: A relationship is expected; as height increases, weight generally increases (positive correlation).
Temperature and Life of milk: A relationship is expected; as temperature increases, the life of the milk decreases (negative correlation).
Length of hair and IQ: No relationship is expected.
Depth of topsoil and Brand of motorcycle: No relationship is expected.
Years of education and Income: A relationship is expected; generally, as education increases, income levels tend to increase (positive correlation).
Spring rainfall and Crop yield: A relationship is expected; usually, increased rainfall leads to a higher crop yield (positive correlation), though this depends on the specific crop requirements.
Size of ship and Cargo capacity: A relationship is expected; as the size of the ship increases, the cargo capacity increases (positive correlation).
Fuel economy and CD track number: No relationship is expected.
Amount of traffic and Travel time: A relationship is expected; as traffic increases, travel time increases (positive correlation).
Cost of 2 litres of milk and Ability to swim: No relationship is expected.
Background noise and Amount of work completed: A relationship is expected; often, as background noise increases, the amount of work completed may decrease (negative correlation).
Likelihood of Strong Correlation
Determining if a strong correlation is likely between specific pairs of variables:
Height of door and Thickness of door handle: Unlikely to have a strong correlation.
Weight of car and Fuel consumption: Likely to have a strong correlation (heavier cars generally consume more fuel).
Temperature and Length of phone calls: Unlikely to have a strong correlation.
Size of textbook and Number of textbook chapters: Likely to have some correlation, though it may vary depending on the subject.
Diameter of flower and Number of bees: May have a correlation, but it is not necessarily strong without other factors.
Amount of rain and Size of vegetables in the vegetable garden: Likely to have a correlation, as water is essential for growth.
Quantitative Data Trends
Observing sets of bivariate data to determine if generally increases or decreases as increases:
Data Set A:
:
:
Trend: As increases, generally increases (positive trend).
Data Set B:
:
:
Trend: As increases, generally decreases (negative trend).
Comprehensive Numerical Example 1: Interpreting Scatter Plots
Consider the following bivariate data set:
Task A: Draw a scatter plot: The data is plotted on a grid where the -axis ranges from to and the -axis ranges from to .
Task B: Describe the direction: The correlation between and is negative. This is because the values of generally fall as the values of rise.
Task C: Describe the strength: The correlation is described as strong because the points (excluding outliers) follow a fairly clear downward linear path.
Task D: Identify outliers: The point is identified as an outlier because it is positioned significantly lower than the general trend established by the other data points.
Comprehensive Numerical Example 2: Interactive Data Set
Consider the following bivariate data set:
Task A: Draw a scatter plot: Points are plotted on a Cartesian plane where is the independent variable and is the dependent variable.
Task B: Describe the direction: The correlation is positive. As values move from toward , the corresponding values generally rise.
Task C: Describe the strength: The correlation is strong, as the majority of the points cluster along an upward sloping line.
Task D: Identify outliers: The point is identified as an outlier because it sits well below the rising trend formed by points like and .