Machine Learning Notes

Introduction to Machine Learning

Arthur Samuel (1959) defined Machine Learning as giving computers the ability to learn without explicit programming.

Tom Mitchell (1998) defined it as algorithms improving performance PP at task TT with experience EE, a well-defined learning task is given by .

Well-Posed Learning Problem

A computer program learns from experience EE regarding task TT and performance measure PP if its performance improves with EE.

Traditional Programming vs. Machine Learning

Traditional programming involves inputting data and a program into a computer to produce output; machine learning involves inputting data and output into a computer to generate a program.

Machine Learning Algorithms

  • Supervised learning

  • Unsupervised learning

Supervised Learning

Learns from being given "right answers". Consists of:

  • Regression: Predicting a number with infinitely many possible outputs, such as housing price prediction.

  • Classification: Predicting categories with a small number of possible outputs, such as breast cancer detection.

Unsupervised Learning

Finds something interesting in unlabeled data. Two main types:

  • Clustering: Grouping similar data points together, like grouping news articles or customers.

  • Anomaly detection: Finding unusual data points.

Terminology

  • Training Set: Data used to train the model.

  • xx = "input" variable / feature.

  • yy = "output" / "target" variable.

  • (x,y)(x, y) = single training example.

  • mm = number of training examples.

  • (x(i),y(i))(x^{(i)}, y^{(i)}) = ithi^{th} training example.

During supervised learning:

  • A learning algorithm is fed with training set, features, and targets.

  • The algorithm produces a function ff called the hypothesis.

  • Given a new input xx, the hypothesis outputs an estimate or prediction y^ŷ.

Linear Regression

The hypothesis function is represented as: fw,b(x)=wx+bf_{w,b}(x) = wx + b, where ww and bb are parameters.

  • ww: weight / coefficient (slope).

  • bb: bias (y-intercept).

This represents a linear regression with one variable, also known as univariate linear regression.

Cost Function

Used to measure the accuracy of the model's predictions.

  • Model: fw,b(x)=wx+bf_{w,b}(x) = wx + b

  • Parameters: w,bw, b

  • Cost function: J(w,b)=12m∑<em>i=1m(f</em>w,b(x(i))−y(i))2J(w,b) = \frac{1}{2m} \sum<em>{i=1}^{m} (f</em>{w,b}(x^{(i)}) - y^{(i)})^2

  • Goal: Minimize J(w,b)J(w,b) by tuning ww and bb

Simplified cost function (with b=0b=0): J(w)=12m∑<em>i=1m(f</em>w(x(i))−y(i))2J(w) = \frac{1}{2m} \sum<em>{i=1}^{m} (f</em>{w}(x^{(i)}) - y^{(i)})^2