DATA UNIT
TERMINOLOGY:
Mean Point / Coordinate:
Calculated by finding the mean of all x-data and y-data (i.e. x-data mean = 30.6, y-data mean = 50.6, SO mean coordiante is [30.6, 50.6]
Line of Best Fit (LOBF):
Has to pass through mean point
Must stay close to all data points in a fair manner
No extension outside data range
IMPORTANT NOTE: Not all data needs a LOBF; some data has too many outliers to have a LOBF
(2 types of) Correlations:
Positive (increasing slope)
Negative (decreasing slope)
Interpolating:
Making a prediction within the data range (refer to āLINEAR REGRESSIONā worksheet for visual clarification).
Expotrapolating:
Making a prediction outside of the data range (refer to āLINEAR REGRESSIONā worksheet for visual clarification).
Sample:
A portional representative of the population
Census:
A representation of the entire population
Types of Sampling:
Simple random sampling
āInvolves choosing a set number of objects/people at random.ā
Systematic random sampling
āInvolves randomly choosing individuals/objects at a fixed interval (e.g. surveying every fourth customer in the store).
Stratified random sampling
āInvolves diving the population into groups and randomly selecting members from each group (e.g. dividing a population into children, teens, adults, and seniors).ā
Non-random sampling
āInvolves using a method that does not randomly select a sample from a population; biasedly selects sample.ā
Cluster sampling
āUsually measures geographic zones; splits up zone into cluster and randomly chooses a cluster to survey ALL people within cluster.ā
Difference between stratified random and cluster sampling:
Cluster = creating and choosing clusters (at random) to suvery ALL objects/individuals within.
Statified random = creating categories/clusters and surveying a fixed amount of objects/individuals within, based on given proportions.
What is Two-Variable Data?:
To collect data on two things; to answer 2 questions with data (i.e. hours of sleep vs. test marks); usually to find a connection/correlation. Has both an x and y axis with data.
What is One-Variable Data?:
To collect data on one thing; to answer 1 question with data (i.e. favourite/most popular fruit in a class); only one thing is being collected. Has only x-axis with data; y-axis is filled in based on anaylsis (how many per answer).
What is categorical data?:
Data that is not numerical; expressed in words/categories.
What is numerical discreet data?:
Data that is numerical, but not continous (meaning you CANNOT have points/data in between them; i.e. dollars on the x-axis that go up by 1ās. Refer to āVisual Displays of One-Variable Dataā sheet for visual representation.)
What is numerical continous data?:
Data that is numerical and continous (meaning you CAN have points/data in between them; i.e. points that go up by 10ās. Refer to āVisual Displays of One-Variable Dataā sheet for visual representation.)
What is a Bar Graph/When is it used?:
Made up of equally wide rectangles WITH SPACES to compare data (refer to āVisual Displays of One-Variable Dataā sheet for visual representation). Used when data is categorical or numercial discreet.
What is a Histogram/When is it used?:
Made of equally wide rectangles WITHOUT spaces to compare data (refer to āVisual Displays of One-Variable Dataā sheet for visual representation). Used when data is numerical continous.
What is frequency (in data)?:
How often something shows up repeatedly in a data set.
What is a āshapeā on a bar graph?:
The shape the bars make on the graph. Can be normal (with one peak in the middle - SYMMETRIC), bimodal (with two peaks - SYMMETRIC), uniform (no peaks/bumps - SYMMETRIC) or left skewed (long tail, peak on right side - ASSYMMETRIC) / right skewed (long tail, peak on left side - ASSYMMETRIC). - [ REFER TO āVisual Displays of One-Variable Dataā SHEET FOR VISUAL REPRESENTATION. ]
What is a āmodalā when talking about bar graph shapes/data?:
The amount of peaks in a bar graph.
What is āModeā in data?:
āData point that has the highest frequency.ā
What is āMedianā in data?:
āThe middle data point if data is ranked from low to high (using frequency). ā
What is the āMeanā in data?:
āThe average of all data points (using frequency).ā
How do calculate / find āMean, Median and Modeā in data?:
MEAN:
(sum of all data points) Ć· (# of data points)
MEDIAN:
(n+1) Ć· 2
n = # of data points; āthis formula only helps us find the POSITION of the median (when all data is numerically organized).ā
MODE:
whichever data (point) has the highest frequency
What is the Central Tendency?
It is the centre of the data; it depends on the interest of the investigation to figure out whether you use mean, median or mode to find the center of the data.
How do you find the Central Tendency?
Look towards the interest of the investigation. Is the investigation asking you to find the majority (mean), middle (median) or average (mean)? If there is an outlier, do not use mean, because it skews your data. Then look towards median and mode, and see which best accurately represents the centre of your data.ā
When and when NOT are you allowed to use a break line for your data?
One-variable data (specifically histograms) ARE allowed break lines.
Two-variable data / y=mx+b lines are NOT allowed break lines (it messes up your proportions & interpolations/exprapolations).
What is a āmodalā in data?
Whichever data value that appears the most frequently/has the highest occurence.
In what circimustances will mean, median and mode all be the same?
If your distrubution is symmetric on your bar graph (i.e. normal, bimodal, trimodal, uniform, etc.), then your mean, median, mode are the same.
In what circumstance can you NOT use a break line for your x-xis?
When doing y = mx+b / usually two-variable data sets.
When working with intervals (i.e. 32.5 - 35.5 cm), how do you find the median?
Find the average between the interval, and then use the bin width (+3) to find the rest of the āinterval-averagesā.
What is Measure of Spread?
How scattered/varied/spread a data set is.
What does it mean if a data set is more clustered?
Itās a small spread, has consistency in data, a concentrated data range.
What does it mean if a data set is more spread out?
Itās a large spread and has more variety.
What is one method of measuring spread in data, and whatās so special about it?
Quartiles and box-whisker plot. This measure of spread is specific to when the median is chosen to be the centre of the data.
How do we find quartiles / box-whisker plots?
First you find the main median, and the median of the two split up-halves number sets.
What are the two split-up halves / their medians called?
Quartile 1 / Qā & Quartile 3 / Qā
What is the range of a data set?
The maximum data value subtracted by the minimum data value.
What is the Interquartile Range (IQR) and how do you calculate it?
The distance spread out between the middle 50% of the data. Calculate by Q3 - Q1. Any data value that is 1.5 times more than the Q3 or Q1, is considered an outlier.
What is a box-whisker plot?
It places the quartile: Q1, Q2 (median) and Q3 on the x-scale (and created them into a box), whereas the minimum data value and maximum data values become points extended to the main box with lines (refer to worksheet for visual reference).
ACTIVITY:
Go back to āVisual Displays of One-Variable Dataā and review the example on the last-and-a-half page in-depth!
YOUāRE DONE!
CONGRATULATIONS!