Bivariate Data and Scatter Plots Study Guide

Learning Objectives for Bivariate Data and Scatter Plots

  • Identify bivariate data, which involves the study of two variables simultaneously.

  • Distinguish between independent and dependent variables within a bivariate data set.

  • Construct scatter plots to visualize and compare data from two variables.

  • Analyze scatter plots to determine the form, direction, and strength of the correlation between variables.

  • Define and identify outliers within a data set.

  • Utilize specific mathematical terminology to describe associations between variables.

Fundamental Concepts of Bivariate Data

  • Definition of Bivariate Data: This type of data includes measurements or observations for two variables collected from the same entity. These variables are typically related in some way, such as a person's height and their weight.

  • Independent Variable: Often referred to as the explanatory variable, this is the variable that is changed, controlled, or used to make predictions. It is systematically plotted on the horizontal xx-axis.

  • Dependent Variable: Often referred to as the response variable, this is the variable being measured, tested, or observed. Its value depends on the independent variable, and it is plotted on the vertical yy-axis.

  • Terminology for Relationships: The terms "relationship," "correlation," and "association" are used interchangeably to describe the nature and degree to which two variables are related.

Characteristics of Scatter Plots

  • Scatter Plot Construction: A scatter plot is a graph drawn on a number plane (Cartesian plane) where the two axes represent the two different variables. Each pair of data values is represented by a single point (x,y)(x, y).

  • Direction of Correlation:

    • Positive Correlation: As the independent variable xx increases, the dependent variable yy also tends to increase.

    • Negative Correlation: As the independent variable xx increases, the dependent variable yy tends to decrease.

  • Strength of Correlation:

    • Strong: The data points fall very close to a straight line or a clear path, indicating a high degree of relationship.

    • Weak: The data points show a general trend but are more widely scattered, indicating a lower degree of relationship.

    • No Correlation: There is no apparent relationship between the variables; the points appear randomly distributed.

  • Outliers: An outlier is a data point that is clearly isolated or significantly distant from the rest of the data clusters. It represents an observation that deviates from the established pattern of the relationship.

Qualitative Analysis of Variable Relationships

When evaluating potential relationships between two variables, it is necessary to consider if one might influence the other and in what direction.

  • Height of person and Weight of person: A relationship is expected; as height increases, weight generally increases (positive correlation).

  • Temperature and Life of milk: A relationship is expected; as temperature increases, the life of the milk decreases (negative correlation).

  • Length of hair and IQ: No relationship is expected.

  • Depth of topsoil and Brand of motorcycle: No relationship is expected.

  • Years of education and Income: A relationship is expected; generally, as education increases, income levels tend to increase (positive correlation).

  • Spring rainfall and Crop yield: A relationship is expected; usually, increased rainfall leads to a higher crop yield (positive correlation), though this depends on the specific crop requirements.

  • Size of ship and Cargo capacity: A relationship is expected; as the size of the ship increases, the cargo capacity increases (positive correlation).

  • Fuel economy and CD track number: No relationship is expected.

  • Amount of traffic and Travel time: A relationship is expected; as traffic increases, travel time increases (positive correlation).

  • Cost of 2 litres of milk and Ability to swim: No relationship is expected.

  • Background noise and Amount of work completed: A relationship is expected; often, as background noise increases, the amount of work completed may decrease (negative correlation).

Likelihood of Strong Correlation

Determining if a strong correlation is likely between specific pairs of variables:

  • Height of door and Thickness of door handle: Unlikely to have a strong correlation.

  • Weight of car and Fuel consumption: Likely to have a strong correlation (heavier cars generally consume more fuel).

  • Temperature and Length of phone calls: Unlikely to have a strong correlation.

  • Size of textbook and Number of textbook chapters: Likely to have some correlation, though it may vary depending on the subject.

  • Diameter of flower and Number of bees: May have a correlation, but it is not necessarily strong without other factors.

  • Amount of rain and Size of vegetables in the vegetable garden: Likely to have a correlation, as water is essential for growth.

Quantitative Data Trends

Observing sets of bivariate data to determine if yy generally increases or decreases as xx increases:

  • Data Set A:

    • xx: 1,2,3,4,5,6,7,8,9,101, 2, 3, 4, 5, 6, 7, 8, 9, 10

    • yy: 3,2,4,4,5,8,7,9,11,123, 2, 4, 4, 5, 8, 7, 9, 11, 12

    • Trend: As xx increases, yy generally increases (positive trend).

  • Data Set B:

    • xx: 0.1,0.3,0.5,0.9,1.0,1.1,1.2,1.6,1.8,2.0,2.50.1, 0.3, 0.5, 0.9, 1.0, 1.1, 1.2, 1.6, 1.8, 2.0, 2.5

    • yy: 10,8,8,6,7,7,7,6,4,3,110, 8, 8, 6, 7, 7, 7, 6, 4, 3, 1

    • Trend: As xx increases, yy generally decreases (negative trend).

Comprehensive Numerical Example 1: Interpreting Scatter Plots

Consider the following bivariate data set:

xx

yy

11

4.84.8

33

2.12.1

99

4.04.0

22

6.26.2

1717

1.31.3

33

5.55.5

66

0.90.9

88

3.53.5

1515

1.61.6

  • Task A: Draw a scatter plot: The data is plotted on a grid where the xx-axis ranges from 00 to 2020 and the yy-axis ranges from 00 to 77.

  • Task B: Describe the direction: The correlation between xx and yy is negative. This is because the values of yy generally fall as the values of xx rise.

  • Task C: Describe the strength: The correlation is described as strong because the points (excluding outliers) follow a fairly clear downward linear path.

  • Task D: Identify outliers: The point (6,0.9)(6, 0.9) is identified as an outlier because it is positioned significantly lower than the general trend established by the other data points.

Comprehensive Numerical Example 2: Interactive Data Set

Consider the following bivariate data set:

xx

yy

1212

4.04.0

22

1.31.3

1515

4.54.5

1010

3.63.6

44

1.81.8

55

2.02.0

88

2.52.5

1313

2.02.0

77

2.92.9

  • Task A: Draw a scatter plot: Points are plotted on a Cartesian plane where xx is the independent variable and yy is the dependent variable.

  • Task B: Describe the direction: The correlation is positive. As xx values move from 22 toward 1515, the corresponding yy values generally rise.

  • Task C: Describe the strength: The correlation is strong, as the majority of the points cluster along an upward sloping line.

  • Task D: Identify outliers: The point (13,2.0)(13, 2.0) is identified as an outlier because it sits well below the rising trend formed by points like (12,4.0)(12, 4.0) and (15,4.5)(15, 4.5).