Census at School 2025: Data Collection and Statistical Inference Notes on Statistical Inference
Overview of the Census at School Project
The primary objective is to collect data from students across New Zealand to create a comprehensive database.
This database serves as a resource for student investigations and internal assessments.
Participants range from to .
Current scale of the project as mentioned in the snapshots:
Data Exploration and Website Navigation
Students utilize the "Census at School" website to explore datasets.
The specific dataset for upcoming assessments is labeled "Census at School ending 2025."
Students are responsible for specifying their own parameters and datasets for their investigations.
Variable Selection: The list of available variables is extensive, unlike smaller preliminary exercises that may have only offered five variables (height, weight, species, gender, and location).
Comparison Variables: Options include comparing different year levels (e.g., vs. ), genders (male vs. female), or regions (e.g., Auckland vs. Canterbury).
Relationship Variables: Students can look for correlations between different measurements, such as:
Height vs. Left wrist circumference
Bag weight vs. Travel time to school
Height vs. Foot length
Sampling Methodologies
Random Sample: Data points are selected purely by chance from the database.
Stratified Sample: This involves taking the same number of samples from specific subgroups to ensure equal representation (e.g., taking an equal number of height measurements from and students).
Data Integrity and Collection Ethics
Anonymous Participation: No names are attached to the data entered into the website.
Potential for Data Errors: Errors frequently occur in the database due to:
Incorrect data entry (typing errors).
Inaccurate measurement techniques.
Limitations of Anonymity: Because the data is anonymous, if an error is identified, the student cannot be identified to re-measure. There are two theoretical ways to handle bad data:
Option 1: Re-measure the data (impossible with anonymous nationwide participants).
Option 2: If the identity were known, the individual would re-measure.
Responsibility: Students are urged to be precise because their data will eventually be used by other students across New Zealand for their own statistical studies.
Upcoming Classroom Data Collection
Next week, students will participate in practical data collection involving six distinct measurement methods.
Workflow:
Instructions will be posted at various stations around the room.
Students will work in pairs to rotate through stations.
Data will be recorded manually before being entered into the national website.
Types of Data Collected:
Physical measurements (e.g., foot length using specific instructions).
Performance data (e.g., a reaction game).
Categorical/Lifestyle data: Playing musical instruments (and which specific instrument), bedtimes from the previous night, and various opinions.
Social media/Technology usage: Data points include usage of YouTube, TikTok, Twitter, Discord, ChatGPT, and whether the student has a phone.
Software Tool: Students will practice downloading data and importing it into NZ Grapher for analysis.
Statistics Investigation: Population and Inference
Inference Definition: Using a sample to make a logical conclusion or statement about a broader population.
Defining the Population: The population must be explicitly defined based on the sample taken.
If a sample is taken only from students, the population is defined as students.
If a sample is taken only from students in Canterbury, the population is students in Canterbury.
Avoiding Bias: It is statistically incorrect to claim a sample represents the entire population of New Zealand if the sample was only drawn from one region (e.g., Canterbury). This introduces bias and lacks representation.
Investigation Question Format: A standard investigation question should look like: "I wonder if there is a relationship between [Variable A] and [Variable B] of students in New Zealand from the Census at School 2025 database?"