1/23
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Define data mining and explain its primary objective in analysing large datasets.
Data mining is the process of extracting useful patterns, trends, relationships, and knowledge from large volumes of data using techniques from statistics, machine learning, and database systems. Its primary objective is to transform raw data into meaningful information for decision-making.
(Example: A supermarket analyses customer purchase records to identify products frequently bought together)
Explain the significance of data mining in decision-making and forecasting. Provide suitable examples.
Data mining improves business decisions, forecasting accuracy, risk identification, and operational efficiency by extracting actionable knowledge and identifying patterns within datasets.
Retail: Predicting sales
Banking: Fraud detection
Healthcare: Disease prediction
Telecommunications: Customer churn analysis
Discuss FOUR (4) main objectives of data mining and provide ONE (1) example for each objective.
Prediction: Utilising historical data to make informed predictions about future trends and outcomes. (Example: Predicting sales for the upcoming quarter based on past sales data).
Classification: Categorising data into predefined classes or groups based on its characteristics. (Example: Classifying emails as spam or not spam using features like keywords and sender information).
Clustering: Grouping similar data points together based on inherent patterns or similarities. (Example: Identifying customer segments based on purchasing behaviour for targeted marketing strategies).
Association Rule Mining: Discovering relationships or associations between variables in large datasets. (Example: Identifying the association between products frequently bought together, like bread and butter).
Explain the concept of prediction in data mining. How is it useful in business organisations?
Prediction uses historical data to estimate future outcomes. It is useful in business organisations because it assists with strategic planning, forecasting, decision-making, risk management, and resource allocation.
Outline the SIX (6) stages involved in the data mining process and explain the importance of each stage.
Data Collection: Identifying relevant structured and unstructured data across analytics applications.
Data Cleaning: Identifying and rectifying errors, inconsistencies, and missing values in the dataset.
Data Integration / Processing: Transforming and organising data for effective analysis.
Data Selection and Transformation (Model Building): Applying data mining algorithms and statistical models to discover patterns and relationships.
Evaluation: Assessing the accuracy and performance of the developed models.
Deployment: Integrating data mining results into operational systems or decision-making processes.
Explain any FOUR (4) popular data mining techniques and describe their applications in real-world scenarios.
Decision Tree: An interpretable predictive modelling algorithm used to classify data or make decisions.
Neural Network: Computational models inspired by the human brain that process information to analyse complex patterns.
Clustering Algorithm: Grouping similar data points for marketing segmentation.
Association Rule Mining: Identifying interesting relationships between variables to recommend products frequently purchased together.
Discuss THREE (3) real-world applications of data mining in industries such as healthcare, finance, marketing, or telecommunications.
Healthcare (Disease Prediction): Analysing patient records and medical histories to predict disease likelihood and assist early diagnosis.
Finance (Fraud Detection): Analysing financial transactions to identify anomalies and potentially fraudulent activity.
Marketing (Customer Segmentation): Analysing customer behaviour and preferences to create targeted advertising campaigns
Explain THREE (3) major challenges faced in data mining and discuss their impact on the effectiveness of data analysis.
Poor Data Quality: Inaccurate or incomplete data leads to flawed analyses, undermines model effectiveness, and produces unreliable results.
Scalability: Managing and processing large volumes of data efficiently; challenges can slow down processing times for big datasets.
Interpretability: Making complex models understandable to non-experts; lack of clarity causes mistrust in model decision-making.
Classify each attribute into the correct attribute type (Nominal, Ordinal, Interval, Ratio
Student ID (Nominal): Labels used for identification only; no natural order and no arithmetic operations.
Gender (Nominal): Categorical data without any natural order.
CGPA (Ratio): Numerical with a meaningful absolute zero point; ratios and comparisons are meaningful.
Course Rank (Ordinal): Indicates order or position, but differences between ranks are not equal.
Temperature in Celsius (Interval): Meaningful intervals, but zero does not signify complete absence of heat.
An online shopping company discovered problems in its customer database. Identify the data quality issue and suggest ONE solution for each: Some customers registered multiple accounts. Several records contain empty phone numbers. Incorrect ages, such as 250 years old, were found:
Duplicate Data: Caused by the same customer appearing multiple times. Solution: Perform data deduplication and enforce unique email addresses or customer IDs.
Missing Data: Essential information is unavailable. Solution: Make phone numbers mandatory during registration or prompt users to update.
Invalid / Inaccurate Data: Values fall outside realistic human boundaries. Solution: Implement validation rules (e.g., accepting age range 0 to 120).
A bank detects a transaction of RM15,000 in one day for a customer whose average spending is RM200/day.
What type of data issue does this represent?
Why is this data important for analysis?
State ONE real-world application of outlier detection.
Outlier (Anomaly): RM15,000 deviates significantly from the normal spending pattern of RM200/day.
Importance: Helps identify fraudulent transactions, detects unusual customer behaviour, improves security/risk management, and enables prompt bank intervention.
Real-world Application: Credit Card Fraud Detection in Banking.
Calculate the Euclidean Distance
d=X,Y=(x1−x2)2+(y1−y2)2
Manhattan Distance between
d(x,y)=∑∣xk−yk∣
A hospital found duplicated patient records, missing blood pressure readings, and different date formats. Discuss the operational impacts and solutions.
Duplicated Patient Records:
Causes confusion, repeated testing, and incorrect histories.
Solution: Conduct regular data cleaning and deduplication using unique patient IDs.
Missing Blood Pressure Readings:
Medical staff make decisions on incomplete data, risking misdiagnosis or improper treatment.
Solution: Make critical fields mandatory and set validation checks.
Different Date Formats:
Creates database inconsistency, complicates integration, and distorts timelines.
Solution: Standardise date formats across all departments (e.g., YYYY-MM-DD).
a) Define data preprocessing and explain its importance in data mining.
b) Describe any three activities involved in data preprocessing.
a) Definition: Data preprocessing is the process of cleaning, transforming, and organizing raw data into a suitable format for analysis and modelling. It is important because data quality directly dictates model accuracy and effectiveness.
b) Activities:
Data Cleaning / Handling Missing Data: Identifying and rectifying errors, inconsistencies, or missing values.
Data Normalisation: Scaling numerical values into a common range to improve model performance.
Feature Selection / Correcting Data Types: Choosing relevant features or converting features to correct formats (Date/Numeric).
a) Explain two methods used to identify missing values in a dataset.
b) Compare the use of mean, median, and mode for handling missing data. Provide a suitable scenario for each.
a) Identification Methods: Using the isnull() or isna() functions in Pandas to display missing counts per column.
Using visualisation libraries such as missingno to map missingness patterns visually.
b) Imputation Methods:
Mean: Used for numerical data that is normally distributed (e.g., replacing missing student test scores or heights).
Median: Used for numerical data that is skewed, as it is robust against extreme outliers (e.g., replacing missing customer income values).
Mode: Used for categorical variables (e.g., replacing missing gender or region values).
a) Discuss three problems caused by duplicate records in a dataset.
b) Explain why correcting data types is important before performing data analysis.
a)
Duplicate Problems: Inaccurate Analysis: Distorts statistical outputs and double-counts values.
Model Bias: Machine learning models give disproportionate weight to repeated data.
Storage Inefficiency: Unnecessarily inflates dataset size and consumes extra memory.
b)
Importance of Correcting Data Types: Prevents calculation errors (e.g., adding text strings instead of numbers), enables proper sorting/filtering (e.g., sorting dates correctly), improves memory efficiency, and ensures accurate machine learning operations.
a) Define an outlier and state two possible causes of outliers.
b) Explain three reasons why handling outliers is important in data analysis and machine learning.
a)
Definition: An outlier is a data point that differs significantly from other observations in a dataset. Causes: Data entry errors (e.g., typing 1000 instead of 100) or measurement errors (e.g., faulty sensor equipment).
b)
Importance: Prevents statistical metrics (like mean and variance) from becoming skewed.
Enhances model performance and accuracy.
Helps isolate valid critical anomalies (e.g., fraud detection or rare medical events).
a) Explain the purpose of data normalisation in machine learning.
b) Differentiate between One-Hot Encoding and Label Encoding with suitable examples.
a)
Purpose of Normalisation: Ensures fair comparison between features with different scales, speeds up gradient descent convergence, prevents large-scale attributes from dominating distance metrics, and boosts model performance in distance-based algorithms (k-NN, K-Means, SVM).
b)
Encoding Comparison:
One-Hot Encoding: Creates separate binary columns for each category.
Best for nominal data without order (e.g., Colour: Red, Blue, Green).
Label Encoding: Assigns sequential numerical integers to categories based on rank.
Best for Ordinal data with natural order (e.g., Level: Low=1, Medium=2, High=3).
a) What is feature selection?
b) Discuss four benefits of feature selection in data mining projects.
a)
Definition: Feature selection is the process of selecting the most relevant independent features contributing to predicting a target variable while dropping redundant or irrelevant attributes.
b)
4 Benefits:
Reduces overfitting by eliminating noise.
Improves model accuracy and performance.
Speeds up training and computation time.
Enhances model interpretability.
a) Define Exploratory Data Analysis (EDA).
b) Describe the six main steps or techniques involved in EDA.
a) Definition: EDA is a systematic process of examining, summarising, and visualising a dataset to understand its key characteristics before applying formal predictive modelling.
b) 6 Steps:
Descriptive Analysis: Summarising central tendency and spread.
Univariate Analysis: Examining single variables one at a time.
Bivariate Analysis: Exploring relationships between variable pairs.
Outlier Detection: Identifying extreme deviations using statistical or algorithmic bounds.
Feature Distribution: Checking shape, skewness, and modality.
Data Transformation: Applying scaling or logarithmic adjustments.
Mean Median Mode
…
a) Compare Univariate Analysis and Bivariate Analysis.
b) Suggest suitable visualisation techniques for:
i. Univariate numerical data.
ii. Bivariate numerical data.
a) Comparison:
Univariate: Analyses one variable in isolation to describe its central tendency, spread, shape, and distribution.
Bivariate: Analyses two variables simultaneously to discover correlations, associations, dependencies, or group differences.
b) Visualisations:
i. Univariate Numerical: Histogram, Boxplot, KDE Plot.
ii. Bivariate Numerical: Scatter Plot, Heatmap (Correlation Matrix).
a) Calculate the mean.
b) Determine Quartiles (Q1, Q2, Q3).
c) Calculate the Interquartile Range (IQR).
d) Calculate Sample Variance (s^2).
e) Calculate Sample Standard Deviation (s).
f) Identify potential outliers using IQR bounds
…