DATA UNIT


TERMINOLOGY:

  • Mean Point / Coordinate:

    • Calculated by finding the mean of all x-data and y-data (i.e. x-data mean = 30.6, y-data mean = 50.6, SO mean coordiante is [30.6, 50.6]

  • Line of Best Fit (LOBF):

    • Has to pass through mean point

    • Must stay close to all data points in a fair manner

    • No extension outside data range

    • IMPORTANT NOTE: Not all data needs a LOBF; some data has too many outliers to have a LOBF

  • (2 types of) Correlations:

    • Positive (increasing slope)

    • Negative (decreasing slope)

  • Interpolating:

    • Making a prediction within the data range (refer to ā€˜LINEAR REGRESSION’ worksheet for visual clarification).

  • Expotrapolating:

    • Making a prediction outside of the data range (refer to ā€˜LINEAR REGRESSION’ worksheet for visual clarification).

  • Sample:

    • A portional representative of the population

  • Census:

    • A representation of the entire population

  • Types of Sampling:

    • Simple random sampling

      • ā€œInvolves choosing a set number of objects/people at random.ā€

    • Systematic random sampling

      • ā€œInvolves randomly choosing individuals/objects at a fixed interval (e.g. surveying every fourth customer in the store).

    • Stratified random sampling

      • ā€œInvolves diving the population into groups and randomly selecting members from each group (e.g. dividing a population into children, teens, adults, and seniors).ā€

    • Non-random sampling

      • ā€œInvolves using a method that does not randomly select a sample from a population; biasedly selects sample.ā€

    • Cluster sampling

      • ā€œUsually measures geographic zones; splits up zone into cluster and randomly chooses a cluster to survey ALL people within cluster.ā€

  • Difference between stratified random and cluster sampling:

    • Cluster = creating and choosing clusters (at random) to suvery ALL objects/individuals within.

    • Statified random = creating categories/clusters and surveying a fixed amount of objects/individuals within, based on given proportions.

  • What is Two-Variable Data?:

    • To collect data on two things; to answer 2 questions with data (i.e. hours of sleep vs. test marks); usually to find a connection/correlation. Has both an x and y axis with data.

  • What is One-Variable Data?:

    • To collect data on one thing; to answer 1 question with data (i.e. favourite/most popular fruit in a class); only one thing is being collected. Has only x-axis with data; y-axis is filled in based on anaylsis (how many per answer).

  • What is categorical data?:

    • Data that is not numerical; expressed in words/categories.

  • What is numerical discreet data?:

    • Data that is numerical, but not continous (meaning you CANNOT have points/data in between them; i.e. dollars on the x-axis that go up by 1’s. Refer to ā€˜Visual Displays of One-Variable Data’ sheet for visual representation.)

  • What is numerical continous data?:

    • Data that is numerical and continous (meaning you CAN have points/data in between them; i.e. points that go up by 10’s. Refer to ā€˜Visual Displays of One-Variable Data’ sheet for visual representation.)

  • What is a Bar Graph/When is it used?:

    • Made up of equally wide rectangles WITH SPACES to compare data (refer to ā€˜Visual Displays of One-Variable Data’ sheet for visual representation). Used when data is categorical or numercial discreet.

  • What is a Histogram/When is it used?:

    • Made of equally wide rectangles WITHOUT spaces to compare data (refer to ā€˜Visual Displays of One-Variable Data’ sheet for visual representation). Used when data is numerical continous.

  • What is frequency (in data)?:

    • How often something shows up repeatedly in a data set.

  • What is a ā€œshapeā€ on a bar graph?:

    • The shape the bars make on the graph. Can be normal (with one peak in the middle - SYMMETRIC), bimodal (with two peaks - SYMMETRIC), uniform (no peaks/bumps - SYMMETRIC) or left skewed (long tail, peak on right side - ASSYMMETRIC) / right skewed (long tail, peak on left side - ASSYMMETRIC). - [ REFER TO ā€˜Visual Displays of One-Variable Data’ SHEET FOR VISUAL REPRESENTATION. ]

  • What is a ā€œmodalā€ when talking about bar graph shapes/data?:

    • The amount of peaks in a bar graph.

  • What is ā€œModeā€ in data?:

    • ā€œData point that has the highest frequency.ā€

  • What is ā€œMedianā€ in data?:

    • ā€œThe middle data point if data is ranked from low to high (using frequency). ā€

  • What is the ā€œMeanā€ in data?:

    • ā€œThe average of all data points (using frequency).ā€

  • How do calculate / find ā€œMean, Median and Modeā€ in data?:

    • MEAN:

      • (sum of all data points) Ć· (# of data points)

    • MEDIAN:

      • (n+1) Ć· 2

        • n = # of data points; ā€œthis formula only helps us find the POSITION of the median (when all data is numerically organized).ā€

    • MODE:

      • whichever data (point) has the highest frequency

  • What is the Central Tendency?

    • It is the centre of the data; it depends on the interest of the investigation to figure out whether you use mean, median or mode to find the center of the data.

  • How do you find the Central Tendency?

    • Look towards the interest of the investigation. Is the investigation asking you to find the majority (mean), middle (median) or average (mean)? If there is an outlier, do not use mean, because it skews your data. Then look towards median and mode, and see which best accurately represents the centre of your data.ā€

  • When and when NOT are you allowed to use a break line for your data?

    • One-variable data (specifically histograms) ARE allowed break lines.

    • Two-variable data / y=mx+b lines are NOT allowed break lines (it messes up your proportions & interpolations/exprapolations).

  • What is a ā€œmodalā€ in data?

    • Whichever data value that appears the most frequently/has the highest occurence.

  • In what circimustances will mean, median and mode all be the same?

    • If your distrubution is symmetric on your bar graph (i.e. normal, bimodal, trimodal, uniform, etc.), then your mean, median, mode are the same.

  • In what circumstance can you NOT use a break line for your x-xis?

    • When doing y = mx+b / usually two-variable data sets.

  • When working with intervals (i.e. 32.5 - 35.5 cm), how do you find the median?

    • Find the average between the interval, and then use the bin width (+3) to find the rest of the ā€œinterval-averagesā€.

  • What is Measure of Spread?

    • How scattered/varied/spread a data set is.

  • What does it mean if a data set is more clustered?

    • It’s a small spread, has consistency in data, a concentrated data range.

  • What does it mean if a data set is more spread out?

    • It’s a large spread and has more variety.

  • What is one method of measuring spread in data, and what’s so special about it?

    • Quartiles and box-whisker plot. This measure of spread is specific to when the median is chosen to be the centre of the data.

  • How do we find quartiles / box-whisker plots?

    • First you find the main median, and the median of the two split up-halves number sets.

  • What are the two split-up halves / their medians called?

    • Quartile 1 / Q₁ & Quartile 3 / Qā‚ƒ

  • What is the range of a data set?

    • The maximum data value subtracted by the minimum data value.

  • What is the Interquartile Range (IQR) and how do you calculate it?

    • The distance spread out between the middle 50% of the data. Calculate by Q3 - Q1. Any data value that is 1.5 times more than the Q3 or Q1, is considered an outlier.

  • What is a box-whisker plot?

    • It places the quartile: Q1, Q2 (median) and Q3 on the x-scale (and created them into a box), whereas the minimum data value and maximum data values become points extended to the main box with lines (refer to worksheet for visual reference).

  • ACTIVITY:

    • Go back to ā€œVisual Displays of One-Variable Dataā€ and review the example on the last-and-a-half page in-depth!

  • YOU’RE DONE!

    • CONGRATULATIONS!