1/41
Flashcards defining core concepts, methods, and terminology in data science, levels of measurement, sampling, big data, storage options, security, and data sovereignty.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Quantitative Data
Numerical data that represents measurable quantities, such as counts, percentages, or statistical values, which can be analysed using mathematical and statistical methods.
Qualitative Data
Descriptive, non-numerical data that captures characteristics, opinions, or experiences, often collected through interviews, observations, or open-ended survey responses.
Discrete Data
A quantitative data type consisting of countable numbers whose values cannot be broken down any further.
Continuous Data
A quantitative data type consisting of measurable numbers that can take any value within a particular range, including fractions or decimals.
Categorical Data
A qualitative data type where data is sorted into distinct non-numerical groups or categories.
Ordinal Data
Categorical data that has a clear order or ranking, but where the difference between each level or rank is not measurable or consistent.
Binary Data
A type of qualitative data that has only two possible outcomes, such as yes/no or true/false.
Textual Data
Qualitative data expressed in text form, typically collected from open-ended survey responses, social media comments, or reviews.
Nominal Level of Measurement
A level of measurement that categorises items without any order or ranking, where each category is unique and no category is better or worse than another.
Interval Level of Measurement
A numerical level of measurement where intervals between values are consistent, but there is no true zero point.
Ratio Level of Measurement
A numerical level of measurement with consistent intervals and a true zero point, enabling meaningful comparisons such as 'twice as much'.
Active Data Collection
A data collection approach where the researcher directly engages with participants to gather data.
Passive Data Collection
A data collection approach where data is gathered in the background without direct interaction or input from a person.

Sampling Methods Overview
Techniques used to collect data from smaller representative groups of a population, including random, stratified, systematic, and convenience sampling.
Random Sampling
A sampling method where every individual in the population has an equal chance of being selected, reducing bias but potentially being time-consuming.
Stratified Sampling
A sampling method where the population is divided into groups based on shared characteristics, and a random sample is taken from each group.
Systematic Sampling
A sampling method where individuals are selected at regular intervals from a population list.
Convenience Sampling
A sampling method where data is collected from individuals who are easy to reach, which is quick but introduces bias.
Relevance (Data Quality)
The degree to which collected data pertains directly to the specific research problem or objective being addressed.
Accuracy (Data Quality)
The correctness and precision of data, ensuring it reflects actual values without errors.
Validity (Data Quality)
The extent to which a data collection tool or question measures what it was intended to measure.
Reliability (Data Quality)
The consistency of data results when collected under the same conditions across multiple attempts.
Primary Data
Information directly collected for a specific research project through methods like surveys, interviews, or experiments.
Secondary Data
Information that has already been collected by others for different purposes, such as database records, statistics, or publications.
Structured Data
Highly organized data formatted and stored in predefined rows and columns, such as SQL databases or Excel spreadsheets, making it easy to analyze quantitatively.
Unstructured Data
Data that lacks a predefined structure or layout, such as audio files, social media posts, or videos, making it more challenging to organize and analyze.
Raw Data
Unprocessed information collected directly from the source that may contain errors, inconsistencies, or missing values.
Processed Data
Data that has undergone screening, cleaning, transformation, and analysis to improve its quality, reliability, and usability.
Blockchain Technology
A digital system used to securely manage and verify data by creating decentralized, transparent public ledgers that record asset transaction histories.
Autofill
A software feature in browsers and apps that automatically completes form fields using stored user information.

Public vs. Private Connections
A comparison of network connections, where public networks lack robust encryption and carry higher risk, while private networks control access and enforce strong encryption standards.
Big Data
Large, complex datasets that grow continuously over time and cannot be easily managed or processed by traditional data-processing methods.
Three V's of Big Data
The defining attributes of big data: Volume (scale of data), Variety (different formats of data), and Velocity (speed of data generation and processing).
Data Warehouse
A centralized digital storage system that consolidates data from multiple sources into one location, optimized for efficient querying, reporting, and analysis.
Data Mining
The process of extracting, cleaning, and applying algorithms to large datasets in a data warehouse to discover hidden patterns, relationships, or trends.
Digital Footprint
The trail of data left behind by individuals through online activities, including social media posts, browsing history, and digital purchases.

Storage Types Comparison
The categorization of data storage systems into local storage, cloud storage, portable storage media, and data warehouses based on capacity, accessibility, and performance.

Data Issues and Responsible Alleviation Measures
Framework for examining data issues—such as bias, copyright, ICIP misuse, and security breaches—across social, ethical, and legal dimensions, along with methods to mitigate them.
Data Sovereignty of Indigenous Peoples
The right and authority of Aboriginal and Torres Strait Islander communities to control how their data—including cultural knowledge, traditions, and practices—is collected, stored, owned, and shared.
OAIC
Office of the Australian Information Commissioner; the authority responsible for regulating Australian privacy, freedom of information, and data protection compliance.
Data Swamp
A neglected, unmanaged, or unorganized data repository that has become inaccessible or unusable due to a lack of governance, metadata, or data cleaning.
Data Literacy
The ability to interpret, evaluate, and apply data effectively, enabling individuals to identify bias, misinformation, or manipulation in presented data.