1/76
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is Machine Learning?
Machine Learning is a branch of AI where algorithms learn patterns from data, build a model, and use it to make predictions about future or unseen data.
What is the core idea of Machine Learning?
Teach a computer to learn patterns from data so it can make predictions or decisions on new/unseen data.
Why is manual email classification not scalable?
Millions of emails arrive, manual checking is expensive and time-consuming, and it does not scale efficiently.
What is the problem with a rule-based spam detection system?
Rules become large and difficult to maintain, new spam patterns can be missed, and rules require frequent manual updates.
How does Machine Learning differ from a rule-based system?
A rule-based system relies on explicitly written rules, while ML learns patterns from previously labelled data.
What is the basic Machine Learning learning pattern?
Past/present data → learn patterns → build model → predict future/unseen data.
What is a dataset?
A collection of data used for analysis or Machine Learning.
What are features in Machine Learning?
Features are the input variables/columns used by a model to make a prediction.
What is a target or label?
The target/label is the output that the model is trying to predict.
Given Marks, Attendance, Assignment, and Result columns for student prediction, which are the features?
Marks, Attendance, and Assignment.
Given Marks, Attendance, Assignment, and Result columns for student prediction, what is the target?
Result.
What is an ML pipeline?
A sequence of steps followed to build, train, evaluate, deploy, and maintain a Machine Learning solution.
What is the purpose of Problem Definition in an ML pipeline?
To clearly define the problem we want to solve or what the model should predict.
What happens during Data Collection?
The required data for solving the defined ML problem is collected.
What is Data Preprocessing?
The process of preparing and cleaning data so it is suitable for Machine Learning.
What is Feature Engineering?
Creating, selecting, or transforming useful features to improve the data used by an ML model.
What is EDA?
Exploratory Data Analysis; exploring and understanding patterns, relationships, distributions, and unusual values in data.
What is the difference between Data Preprocessing and EDA?
Preprocessing prepares or fixes the data; EDA explores and helps understand the data.
What is Model Selection?
Choosing a suitable ML algorithm/model for the problem and data.
What is Model Training?
Teaching the selected ML model to learn patterns from the training data.
What is Model Evaluation?
Measuring how well a trained model makes predictions.
What is Model Tuning?
Adjusting model settings to try to improve its performance.
What is Deployment in Machine Learning?
Putting the trained model into a real-world application so it can make predictions on new data.
Why is a deployed ML model monitored?
Real-world data and patterns can change, causing model performance to decrease and potentially requiring updating or retraining.
What are the four main types of Machine Learning?
Supervised Learning, Unsupervised Learning, Semi-supervised Learning, and Reinforcement Learning.
What are the two main types of problems in Supervised Learning?
Classification and Regression.
What is Classification?
A supervised learning problem where the model predicts which category or class a new input belongs to.
What are the two types of Classification covered?
Binary Classification and Multi-class Classification.
What is Binary Classification?
A classification problem where the output has exactly two categories.
Give three examples of Binary Classification.
Spam/Not Spam, Disease detection, and Pass/Fail student result prediction.
What is Multi-class Classification?
A classification problem where the output has more than two categories.
Give three examples of Multi-class Classification.
Weather prediction, student grade prediction, and movie genre classification.
What is Regression?
A supervised learning problem where the output is a continuous numerical value.
Give three examples of Regression.
House price prediction, salary prediction, and car price prediction.
What is Unsupervised Learning?
Training a model using input data without providing output labels.
What are the two areas of Unsupervised Learning covered?
Clustering and Dimensionality Reduction.
What is Clustering?
Group or segregate data based on similar patterns without predefined labels.
Give two practical examples of Clustering.
Customer segmentation and news article segmentation.
What is Dimensionality Reduction?
Reducing the number of columns or dimensions while preserving as much important information as possible.
What is Semi-supervised Learning?
Learning from a small amount of labelled data together with a large amount of unlabelled data.
Why is Semi-supervised Learning useful?
Labelling large amounts of data manually can be costly, so a small labelled portion can help make use of the unlabelled data.
What is Reinforcement Learning?
A type of ML where an agent learns through trial and error using rewards and penalties.
What is the basic feedback loop in Reinforcement Learning?
Agent takes an action → receives reward/penalty → learns from feedback → improves future actions.
What is Feature Encoding?
Converting categorical or text data into numerical representations so it can be used by an ML model.
Why is Feature Encoding needed?
ML algorithms perform mathematical operations, so categorical text such as Male/Female or Spam/Not Spam must be converted into numerical form.
What are the three Feature Encoding methods covered?
Label Encoding, One-Hot Encoding, and Ordinal Encoding.
What does Label Encoding do?
It assigns a unique integer to each category.
How does LabelEncoder assign numbers by default?
It orders categories alphabetically and assigns integers starting from 0.
What is the LabelEncoder mapping for Cloudy, Rainy, Sunny?
Cloudy → 0, Rainy → 1, Sunny → 2.
Given Climate = ["Sunny", "Rainy", "Cloudy"], what are the LabelEncoder values?
[2, 1, 0].
What is the important limitation of Label Encoding?
The assigned numbers can imply an order even when the categories have no meaningful ranking.
Why is Label Encoding problematic for Bad, Average, Good, Excellent?
The numerical values are assigned according to category ordering rather than the actual quality ranking, which can create a misleading order.
What does fit() do in an encoder?
It learns the categories and their mapping from the data.
What does transform() do in an encoder?
It applies the learned mapping to convert categorical values into encoded values.
What does fit_transform() do?
It combines fitting the encoder to the data and transforming the data using the learned mapping.
What is the general syntax for Label Encoding?
from sklearn.preprocessing import LabelEncoder; encoder = LabelEncoder(); df["col_name"] = encoder.fit_transform(df["col_name"]).
What does One-Hot Encoding do?
It creates a separate 0/1 column for each category.
Why is One-Hot Encoding useful for categories with no natural order?
It avoids falsely implying a ranking between categories.
For Climate categories Cloudy, Rainy, Sunny, how many One-Hot columns are created?
Three columns: Climate_Cloudy, Climate_Rainy, and Climate_Sunny.
In One-Hot Encoding, what does a 1 mean?
The row belongs to that category.
What input shape does OneHotEncoder expect?
A 2-D input, such as df[["Climate"]], rather than a 1-D Series such as df["Climate"].
What common error occurs when passing df["Climate"] directly to OneHotEncoder?
A ValueError occurs because OneHotEncoder expects 2-D input.
How can OneHotEncoder return a regular NumPy array instead of a sparse matrix?
Use OneHotEncoder(sparse_output=False).
How do you get readable column names after One-Hot Encoding?
Use encoder.get_feature_names_out().
What is the difference between Label Encoding and One-Hot Encoding?
Label Encoding converts categories into one integer column; One-Hot Encoding creates a separate 0/1 column for each category.
What does Ordinal Encoding do?
It converts categories into numbers according to an explicitly defined order.
When should Ordinal Encoding be used?
When categories have a genuine meaningful ranking, such as Low < Medium < High.
Why would Label Encoding be inappropriate for Low, Medium, High?
Its category ordering may produce High → 0, Low → 1, Medium → 2, which does not represent the meaningful Low < Medium < High order.
What is the correct Ordinal Encoding mapping for Low, Medium, High?
Low → 0, Medium → 1, High → 2.
What is the key difference between Label Encoding and Ordinal Encoding?
Label Encoding assigns numbers based on the encoder's category ordering; Ordinal Encoding allows you to explicitly define the meaningful category order.
What is the general syntax for specifying an OrdinalEncoder category order?
OrdinalEncoder(categories=[["Low", "Medium", "High"]]).
Given Salary = ["Low", "Medium", "High"] and categories = [["Low", "Medium", "High"]], what does OrdinalEncoder produce?
[0, 1, 2].
What is the Feature Encoding hierarchy?
Feature Encoding → Label Encoding, One-Hot Encoding, Ordinal Encoding.
Where does Feature Encoding belong in the ML structure?
Feature Encoding belongs under Data Preprocessing.
Where do Classification and Regression belong in the ML structure?
They belong under Supervised Learning, which is one of the four main types of Machine Learning.
What is the high-level ML hierarchy learned so far?
Machine Learning → Supervised/Unsupervised/Semi-supervised/Reinforcement; Supervised → Classification/Regression; Classification → Binary/Multi-class; Unsupervised → Clustering/Dimensionality Reduction.
What is the high-level relationship between ML Pipeline and ML Types?
ML Pipeline describes the process of building and using an ML solution, while ML Types describe the learning approach used by the model.