Census at School 2025: Data Collection and Statistical Inference Notes on Statistical Inference

Overview of the Census at School Project

  • The primary objective is to collect data from students across New Zealand to create a comprehensive database.

  • This database serves as a resource for student investigations and internal assessments.

  • Participants range from Year 3\text{Year } 3 to Year 13\text{Year } 13.

  • Current scale of the project as mentioned in the snapshots:

    • Schools participating: 421\text{Schools participating: } 421

    • Registered teachers: 901\text{Registered teachers: } 901

Data Exploration and Website Navigation

  • Students utilize the "Census at School" website to explore datasets.

  • The specific dataset for upcoming assessments is labeled "Census at School ending 2025."

  • Students are responsible for specifying their own parameters and datasets for their investigations.

  • Variable Selection: The list of available variables is extensive, unlike smaller preliminary exercises that may have only offered five variables (height, weight, species, gender, and location).

    • Comparison Variables: Options include comparing different year levels (e.g., Year 11\text{Year } 11 vs. Year 9\text{Year } 9), genders (male vs. female), or regions (e.g., Auckland vs. Canterbury).

    • Relationship Variables: Students can look for correlations between different measurements, such as:

      • Height vs. Left wrist circumference

      • Bag weight vs. Travel time to school

      • Height vs. Foot length

Sampling Methodologies

  • Random Sample: Data points are selected purely by chance from the database.

  • Stratified Sample: This involves taking the same number of samples from specific subgroups to ensure equal representation (e.g., taking an equal number of height measurements from Year 3\text{Year } 3 and Year 11\text{Year } 11 students).

Data Integrity and Collection Ethics

  • Anonymous Participation: No names are attached to the data entered into the website.

  • Potential for Data Errors: Errors frequently occur in the database due to:

    • Incorrect data entry (typing errors).

    • Inaccurate measurement techniques.

  • Limitations of Anonymity: Because the data is anonymous, if an error is identified, the student cannot be identified to re-measure. There are two theoretical ways to handle bad data:

    • Option 1: Re-measure the data (impossible with anonymous nationwide participants).

    • Option 2: If the identity were known, the individual would re-measure.

  • Responsibility: Students are urged to be precise because their data will eventually be used by other students across New Zealand for their own statistical studies.

Upcoming Classroom Data Collection

  • Next week, students will participate in practical data collection involving six distinct measurement methods.

  • Workflow:

    • Instructions will be posted at various stations around the room.

    • Students will work in pairs to rotate through stations.

    • Data will be recorded manually before being entered into the national website.

  • Types of Data Collected:

    • Physical measurements (e.g., foot length using specific instructions).

    • Performance data (e.g., a reaction game).

    • Categorical/Lifestyle data: Playing musical instruments (and which specific instrument), bedtimes from the previous night, and various opinions.

    • Social media/Technology usage: Data points include usage of YouTube, TikTok, Twitter, Discord, ChatGPT, and whether the student has a phone.

  • Software Tool: Students will practice downloading data and importing it into NZ Grapher for analysis.

Statistics Investigation: Population and Inference

  • Inference Definition: Using a sample to make a logical conclusion or statement about a broader population.

  • Defining the Population: The population must be explicitly defined based on the sample taken.

    • If a sample is taken only from Year 11\text{Year } 11 students, the population is defined as Year 11\text{Year } 11 students.

    • If a sample is taken only from students in Canterbury, the population is students in Canterbury.

  • Avoiding Bias: It is statistically incorrect to claim a sample represents the entire population of New Zealand if the sample was only drawn from one region (e.g., Canterbury). This introduces bias and lacks representation.

  • Investigation Question Format: A standard investigation question should look like: "I wonder if there is a relationship between [Variable A] and [Variable B] of students in New Zealand from the Census at School 2025 database?"