Week 1: Data Levels & Descriptive Statistics (VOCABULARY)
Data types and levels of measurement
- Data can be qualitative (categorical) or quantitative (numerical).
- Qualitative data includes nominal and ordinal; Quantitative data includes discrete and continuous.
- Levels of measurement (the most informative level you can claim about the data):
- Nominal: categories with no intrinsic order (e.g., gender, color).
- Ordinal: categories with a natural order but not necessarily equal intervals (e.g., ratings: poor, fair, good).
- Interval: numerical differences are meaningful, but there is no true zero (e.g., Celsius temperature).
- Ratio: numerical differences and ratios are meaningful, and there is a true zero (e.g., height, weight, counts).
- Decision rule used in the lecture: you look for the highest level the data reaches and report that level (e.g., if data can be ordered and has meaningful differences, it’s at least ordinal; if it also has a true zero, it could be ratio).
- Examples from the transcript:
- Ratings like above average, just average, below average: qualitative (words) with an inherent order → ordinal data.
- Number of courses enrolled: quantitative, discrete (countable values) and, because zero means there are none, can be treated as ratio data under the right interpretation.
- Age: described as continuous in the discussion; while ages are measured on a continuum, rounding to whole numbers doesn’t change the underlying continuous nature. Depending on interpretation, age can be treated as interval or ratio in some contexts; the lecturer emphasizes continuous treatment and the nuance that you can’t always subtract/divide when data are word labels.
- GPA (e.g., 3.4): treated as interval data in the lecture; not considered a true ratio scale because a direct division/apportionment (e.g., double) isn’t interpretable on this scale.
- A practical takeaway: start by classifying the data by qualitative vs quantitative, then identify whether the data are nominal, ordinal, interval, or ratio, and finally consider whether they are discrete or continuous within the quantitative realm.
Central concepts for describing data (descriptive statistics)
- Three core characteristics to describe every data set:
- Central tendency: the typical or middle value.
- Dispersion: how spread out the data are around the center.
- Distribution (shape): the pattern of how data values are laid out across the range.
- A fourth, important feature to note: outliers (extremely unusual values that don’t fit the rest of the data).
- For qualitative data, the descriptive focus is on categories and frequencies; for quantitative data, numerical summaries apply.
Describing qualitative data (categorical data)
- Appropriate descriptive statistics:
- Mode: the most frequent category. Defined as extmode=extargmaxcfreq(c) for categories c.
- Variation ratio (VR): proportion of all observations not equal to the mode, i.e., the non-modal share. Formally, if fmode is the frequency of the mode and N is the total number of observations, then
VR=1−Nf</em>mode
and as a percentage,
VR<em>%=(1−Nf</em>mode)×100%=NN−fmode×100%.
- Practical workflow (as described in the transcript): use a descriptive statistics spreadsheet to compute the mode and VR, and to generate a distribution table (frequencies and percentages) for each category.
- Describing distribution for qualitative data:
- List categories and corresponding percentages (e.g., “28 of 44 students chose ‘just average’” → 63.64%).
- Example interpretation from the transcript (math-ability rating with categories above average, just average, below average):
- Counts: just average = 28, below average = 8, above average = 8; total N = 44.
- Mode: just average (the category with the highest frequency).
- VR: VR=1−Nfmode=1−4428=4416≈0.3636=36.36%.
- Percentages:
- just average: 4428×100%=63.64%.
- below average: 448×100%=18.18%.
- above average: 448×100%=18.18%.
- Narrative description rubric (sample phrasing):
- “Among the Sierra College Introduction to Statistics students, the most common rating of math ability was ‘just average’ (63.64%); 18.18% reported ‘below average’ and 18.18% reported ‘above average’. The variation ratio was 36.36%, indicating a moderate level of diversity around the modal category.”
Describing quantitative data (numerical data)
- Three-part framework (same three core characteristics applied to numbers):
- Central tendency: mean, median (and sometimes mode).
- Dispersion: range, variance, standard deviation, interquartile range; and concept of practical dispersion (how far data are spread from the center).
- Distribution (shape): pattern or shape of the data distribution (e.g., normal, skewed, multimodal).
- Patterns and shapes:
- Bell-shaped (normal-like) distribution: data concentrated around the center with symmetry.
- Skewed distributions: skew to the right (positive skew) or left (negative skew); tails extend farther in one direction.
- Uniform (even) distribution: data are spread relatively evenly across the range.
- Multimodal distributions: more than one peak (e.g., multiple bell shapes).
- Outliers: extremely unusual values that stand apart from the rest of the data; should be identified and described if present.
- Important caveat about interpretation: for interval data, differences between values are meaningful; for ratio data, both differences and ratios are meaningful due to a true zero. In the lecture examples, the distinction was used to decide whether a dataset could be treated as ratio (true zero) or interval/other; for some scales like GPA, the instructor treated it as interval rather than ratio because ratios (e.g., 2.0 is twice 1.0) are not meaningfully defined on that scale.
Worked qualitative example (math ability Rating) - detailed walkthrough
- Context: Population/sample from Sierra College Introduction to Statistics students; qualitative data (math ability rating) with categories: above average, just average, below average.
- Data: 44 students; counts: just average = 28, below average = 8, above average = 8.
- Calculations:
- Mode: just average (the most frequent category).
- Variation ratio: VR=1−Nfmode=1−4428=4416≈0.3636=36.36%.
- Percentages:
- just average: 4428×100%=63.64%.
- below average: 448×100%=18.18%.
- above average: 448×100%=18.18%.
- Spreadsheet workflow described:
- Use the designated Descriptive Statistics spreadsheet with a qualitative data tab.
- Enter data in the qualitative data column; the sheet computes mode and VR and creates a distribution table.
- If the table doesn’t auto-refresh, select the table and refresh to update counts and percentages.
- The worksheet shows counts and percentages for each category; ensure data entry reflects all observations.
- Because the file is owned (read-only) for the teacher, students must use File -> Make a Copy to create their own editable version in OneDrive.
- Copying online stores the file in OneDrive; alternatively, Download which saves a local copy (less ideal for ongoing coursework).
- After data are entered, you must share the link to the student’s OneDrive file with the instructor and ensure proper sharing permissions.
- Textual description you would write:
- “Descriptive statistics for math ability rating (qualitative data) show the modal category as ‘just average’. The variation ratio is 36.36%, indicating a moderate amount of dispersion around the mode. The distribution shows 63.64% of students reporting ‘just average’, with 18.18% reporting ‘below average’ and 18.18% reporting ‘above average’.”
- Notes on common issues:
- Ensure you refresh the distribution table after data entry; problematic refresh can leave old data showing.
- Ensure you copy the spreadsheet to your own OneDrive to avoid overwriting the instructor’s file.
- When sharing, adjust permissions so the instructor can view your work.
Worked quantitative example (movies: MPA rating) - descriptive walkthrough
- Context: Dataset of movies from the Internet Movie Database; qualitative data with MPA ratings: G, PG, PG-13, R.
- Reported results in the transcript:
- PG-13: 43.85%
- PG: 25.38%
- R: 23.08%
- G: 7.69%
- Variation ratio: 56.15%
- Interpretation guide:
- Mode is the most frequent category (the category with the highest percentage).
- VR indicates how diverse the data are around the modal category; higher VR means more diversity; lower VR means samples cluster around the mode.
- In the movie example, the distribution shows a fairly even mix with PG-13 as the largest category, followed by PG and R, with G being the least represented; VR > 50% suggests a moderate level of diversity.
- Textual description you would craft for this dataset:
- “For the sampled movies, the most common rating was PG-13 (43.85%), followed by PG (25.38%) and R (23.08%), with G representing 7.69%. The variation ratio is 56.15%, indicating a moderate amount of diversity in rating categories.”
- Practical workflow notes (movies):
- Create a new descriptive statistics workbook named after the dataset (e.g., ‘movies’).
- Enter sample data into the qualitative data tab and refresh the results to get the updated mode, VR, and distribution table.
- Share the workbook link with the instructor so the work can be reviewed along with the descriptive narrative.
Practical workflow and data-management notes (shared spreadsheets)
- Access and setup:
- Log in to Canvas with your Sierra College credentials.
- Open the Microsoft OneDrive integration within the Canvas course.
- Create or choose a folder (e.g., ‘statistics’) to organize all descriptive-statistics problems.
- Spreadsheet workflow:
- Each descriptive statistics problem uses a dedicated copy of the spreadsheet.
- To use a new problem, open the template and choose File -> Make a Copy (online) to save a working version to your OneDrive; alternatively, Download to save a local copy (not preferred for ongoing work).
- Enter your sample data in the designated column; the spreadsheet computes mode and VR (for qualitative data) or central tendency/dispersion/distribution (for quantitative data).
- If the automatic table doesn’t refresh, click anywhere on the table and use the refresh command to update values; do not rely on a naive reload.
- Describing work in words:
- You must write a short descriptive paragraph that ties your numerical results to the context and data collection:
- Include the population/sample source, the data type, and the three descriptive characteristics (central tendency, dispersion, distribution) plus any outliers.
- Collaboration and sharing:
- When sharing, ensure proper permissions so the instructor can view your workbook; use the share settings to allow access to the instructor via a link.
- Paste the link to your workbook in your assignment submission along with your descriptive narrative.
- Worked reminders:
- You may be asked to describe both qualitative and quantitative datasets; ensure you use the appropriate statistics (mode & VR for qualitative; mean/median, range/variance/std for quantitative).
- Exact wording matters: describe the central tendency (which value is typical), dispersion (how spread out), distribution (shape/pattern), and note any outliers.
- Variation ratio (VR) for qualitative data:
VR=1−Nf<em>modeVR</em>%=(1−Nf<em>mode)×100%=NN−f</em>mode×100% - Example calculation (qualitative):
- If N = 44, fmode = 28 (just average), then
VR=1−4428=4416≈0.3636⇒VR</em>%≈36.36%.
- Percent for just average: 4428×100%=63.64%.
- Data interpretation notes:
- For interval data: differences between values are meaningful; you can subtract values.
- For ratio data: both differences and ratios are meaningful; there is a true zero.
- Example distribution statements (qualitative):
- “63.64% just average, 18.18% below average, 18.18% above average.”
- Example distribution statements (quantitative):
- “The data are moderately dispersed around the mean, with a mix of values that form a roughly bell-shaped distribution and a few outliers.”
Quick-reference checklist for descriptive statistics tasks
- Identify data type and level of measurement (nominal, ordinal, interval, ratio).
- Decide whether data are qualitative or quantitative; if quantitative, note whether they are discrete or continuous.
- For qualitative data, compute and report:
- Mode
- Variation ratio (VR)
- Frequency table and percentages
- For quantitative data, compute and report:
- Central tendency: mean and/or median (and mode if applicable)
- Dispersion: range, variance, standard deviation, possibly IQR
- Distribution: shape (skewness, modality), and note any outliers
- Provide a contextual, sentence-level description that connects the statistics to the real-world setting and data collection.
- Use appropriate tools (spreadsheet, calculator) and ensure your workbook is properly saved, copied to your OneDrive, and shared with the instructor with the correct permissions.