Statistical Learning, Linear Models, and Non-Linear Decision Theory
Terminology and Notation in Statistical Learning
Supervised Learning Fundamentals:
Supervised learning is a paradigm where a predictive model is trained using observed inputs and desired outputs. The central objective is to use input variables to predict target output variables .
Inputs: Also referred to as features, predictors, or independent variables.
Outputs: Also referred to as responses or dependent variables.
Variable Types:
In most practical statistical learning scenarios, inputs are quantitative (i.e., continuous or real-valued numbers ).
Outputs can be quantitative (e.g., in regression problems where ) or qualitative/categorical (e.g., in classification problems where output ).
Target Encoding:
Qualitative output variables are typically encoded into numeric representations for computational convenience. For example, binary outcomes such as "success" versus "failure" are mapped to and , respectively.
Numeric code values representing categories are formally termed targets or labels.
Comprehensive Mathematical Notation Guide:
: The generic input variable representing a vector of features.
: The -th component variable or feature of
: The -th observed instance of , structured as a -vector.
: The specific value of feature within the -th observed sample vector
: Quantitative output variable ().
: Qualitative or categorical output variable ().
: An feature matrix containing observed sample vectors for
: An -vector column within matrix comprising all observed measurements for feature variable
: The -th row of matrix (written as a row vector, since all standalone vectors default to column vectors).
: The model's predicted quantitative output value ().
: The model's predicted categorical class label ().
Indexing and Boldface Conventions:
Index is strictly used for rows (sample observations ranging from to ).
Index is strictly used for columns (input features ranging from to ).
Boldface Rule: Bold typography is reserved exclusively for matrices (e.g., ) and -vectors representing entire feature columns (e.g., ). Boldface is not applied to individual observation -vectors (such as ).
Training Set Representation: A dataset comprising observations used for training is expressed as

Linear Regression Models
Formulation of the Linear Model:
Linear regression models form a foundational baseline for statistical learning.

Two-Feature Model Formulation:
Given an input dataset with features , the model estimates scalar coefficients to predict values:
Generalization to Features:
For an arbitrary number of features , the prediction formula for observation is:
Generic Variable Form:
Dropping the specific observation index yields the generic equation over continuous input spaces:
Vectorized Representation and Bias Vector Augmentation:
Using explicit vector inner products, the equation simplifies to:
Constant Augmentation Trick:
By appending a constant term to the feature vector (making ), the intercept is seamlessly folded into the coefficient vector :
This compact vector form simplifies underlying mathematics and mirrors exact computational implementations.


Python Computational Implementation:
Numerical linear algebra packages execute matrix operations efficiently via matrix multiplication operators or dot products:
import numpy as np
# X: shape (p,), beta_hat: shape (p,)
Y_hat = X.T @ beta_hat # returns scalar
import numpy as np
# X: shape (p,), beta_hat: shape (p,)
Y_hat = np.dot(X.T, beta_hat) # returns scalar
Extension to Multiple Outputs (-Vector Response):
When predicting distinct targets simultaneously (), the model expands to:
Each individual output variable within requires its own dedicated set of parameter coefficients. Otherwise, identical predictions would be generated across all outputs.

Geometric Interpretation and Fitting of Linear Models
Geometric Interpretations:
One Feature (, ):
The system defines a straight line in two-dimensional space .
Two Features (, ):
The system forms a flat plane residing in three-dimensional space .
General Case ( Features, ):
The model produces a -dimensional hyperplane embedded within a -dimensional coordinate space.
Multiple Target Outputs ():
The overall space remains -dimensional, containing distinct hyperplanes. Each input instance maps to independent values across these planes.
Least Squares Fitting on Training Data:
Given a design matrix representing training observations, the prediction vector across all training instances is .

The objective is finding optimal parameters making predicted outputs as close to observed targets as possible.
Residual Sum of Squares (RSS):
The estimation process minimizing RSS is termed the method of least squares.
Analytical Derivation of the Least Squares Parameter Vector:
Express RSS in matrix-vector notation:
Differentiate with respect to vector and set derivative equal to zero:
Solve for :
Classification via Linear Regression Thresholding:
Consider a dataset of 100 sample points with two feature dimensions and a binary qualitative outcome , encoded numerically as Blue = 0 and Orange = 1.
Fitting a linear regression model yields continuous estimates , which are transformed into categorical predictions using a threshold decision rule at :
The resulting decision boundary in feature space is linear, defined precisely by the line

Non-Linear Models: -Nearest Neighbor () Methods
Core Principles:
Nearest-neighbor methods estimate output at target point by averaging response values of training observations located closest to in input feature space.
Formal Definition of : where specifies the neighborhood surrounding point , defined by the nearest training instances
Classification Decision Rule:
Given threshold (typically for binary outcomes):
Distance Metric:
Closeness is evaluated using Euclidean distance standard metric:
Behavior Under Different Values of :
Classifier:
Determines class membership via majority vote across the 15 nearest neighboring points.
Generates a smooth, non-linear decision boundary.

Classifier:
Sets prediction equal to the exact label of the single nearest point to
Creates a highly complex, non-linear decision boundary containing enclosed decision islands.
Training Error Rate: Exactly misclassifications on training data. Because every training sample instance is its own closest neighbor (), the model evaluates , yielding zero prediction errors during training.

Comparison: Linear vs. Non-Linear () Models
Model Flexibility, Complexity, and Trade-offs:
Linear Model Properties:
Makes higher misclassifications on non-linearly structured training data.
Highly rigid parameterization. Represented compactly by only scalar values (parameters ).
Low model complexity yields strong stability and low variance when evaluating unseen test data.
Model Properties:
Perfectly classifies training data ($0\% error rate).\n - Controlled by a single global hyperparameter k \in \mathbb{N}\frac{N}{k} local regions.\n - Extremely fragile and highly sensitive to training data noise (high variance), leading to severe overfitting on new data points.\n\n- **Synthetic Data Benchmark Setup**:\n - Synthetic simulation demonstrating non-linear cluster overlap:\n - Total sample size: 200 observations (100 Blue, 100 Orange).\n - Blue class centers: 10 mean vectors m_k\mathcal{N}\left( (1,0)^T, \mathbf{I} \right).\n - Orange class centers: 10 mean vectors m_k generated from bivariate Gaussian distribution $ NORTH \mathcal{N}\left( (0,1)^T, \mathbf{I} \right).
Individual sample generation: Select cluster mean uniformly at random with probability , then sample data point from .
This generative process produces overlapping multi-modal clusters that linear decision boundaries cannot separate cleanly.
Statistical Decision Theory Framework
Theoretical Framework Principles:
Let denote a real-valued random input vector, and represent a real-valued target random variable, governed by joint probability distribution .
The goal is finding a prediction function mapping inputs to predicted targets
Loss Function:
Quantifies penalization incurred by prediction errors. Using squared error loss:
Mathematical Foundations of Expectation:
Expectation of continuous variable :
Expectation of function :
Discrete:
Continuous:
In decision theory, the evaluated function is squared loss
Expected Prediction Error (EPE):
Objective criterion for selecting optimal function :
Applying conditional probability factorization , EPE rewrites as:
The inner expectation calculates expected squared error over all outcomes conditional on a specific input . The outer expectation averages these errors across input space
Pointwise Minimization and the Regression Function:
To minimize total , it suffices to minimize expected error conditional on every individual input value : where represents a fixed predicted scalar value.
The condition does not isolate a single training instance; rather, it refers to all possible population instances sharing feature values . For a given , target remains a random variable.
The theoretical minimizer of expected squared error is the conditional expectation . Thus, the optimal theoretical target solution is: which is formally termed the regression function.

Conceptual Application (Real Estate Evaluation):
Consider predicting home values using features (bedroom count), (square footage), and (neighborhood quality rating).
The regression function does not represent the average price of a single home, nor does it mean average price across the full housing market.
It equals the population average price across all homes possessing identical features .
Theoretical Target vs. Empirical Reality:
In practical applications, the true joint probability distribution is unknown.
Continuous input features mean exact feature matches appear at most once in finite datasets. Evaluating empirical expectations over a single observation yields no smoothing capability.
Parametric and non-parametric models represent practical computational attempts to estimate using finite data.
Model Alignment with Decision Theory:
Linear Regression: Assumes is approximated by a global linear structure . Solves by substituting this structure into the theoretical EPE equation.
: Relaxes conditioning at an exact point to conditioning over a localized neighborhood region , replacing theoretical population expectations with sample arithmetic averages: Assumes is approximated locally by a constant function.
Asymptotic Convergence of :
As sample size and neighborhood size such that :
Under asymptotic infinite sample conditions, operates as a universal function approximator.
Decision Theory for Categorical Outputs (Classification & Bayes Classifier)
Decision Theory Formulation for Classification:
Let target class labels belong to finite categorical set with set size . Prediction model outputs labels in .
Loss Matrix :
Costs are specified by a matrix , where entry defines the loss incurred by classifying an observation belonging to true class as predicted class
Correct classifications along the matrix diagonal incur zero cost ().

0-1 Loss Function:
The standard classification loss penalizes every misclassification with a unit cost of 1:
Categorical Expected Prediction Error:
Integrating misclassification loss over joint distribution yields:
Pointwise minimization at specific feature vector requires:
Derivation of the Bayes Classifier Under 0-1 Loss:
Substituting 0-1 loss simplifies the summation:
Minimizing misclassification cost simplifies directly to maximizing class probability: \hat{G}(x) = \operatorname{argmin}_{g \in \mathcal{G}} \left[ 1 - \operatorname{Pr}(g \mid X = x) \right] = \operatorname{argmax}_{g \in \mathcal{G}} \operatorname{Pr}(g \n\mid X = x)
Bayes Optimal Decision Rule:
The strategy assigning observations to the most probable class conditional on is termed the Bayes classifier. The minimal error rate achieved by this optimal boundary is the Bayes rate.
classification directly approximates the Bayes classifier by replacing exact conditional class probabilities with empirical training class frequencies inside local neighborhood .

The Curse of Dimensionality in Local Methods
Breakdown of Local Methods in High Dimensions:
Although asymptotic convergence holds theoretically as , high feature dimensionality () degrades performance drastically. This phenomenon is known as the curse of dimensionality.
Geometric Distortion 1: Loss of Neighborhood Locality:
Assume input variables are uniformly distributed within a unit hypercube .
Consider a sub-cube neighborhood capturing a fractional proportion of total sample volume. The required edge length of this sub-cube is:
Evaluating required edge lengths to capture a small volume fraction () across dimensions:
($10\% of feature range)\n - p = 2 \implies e_2(0.1) = 0.32\n - p = 3 \implies e_3(0.1) = 0.46 ($46\% of feature range)
($80\% of feature range)\n - **Implication**: To collect even 10\%80\% of the full coordinate range along every input dimension. Consequently, neighborhoods cease being local, invalidating local constant assumptions.\n\n\n\n\n\n- **Geometric Distortion 2: Vanishing Inscribed Spherical Volume**:\n - Consider a hypersphere S_pC_pp dimensions.\n - The ratio of hypersphere volume to hypercube volume approaches zero rapidly as dimension increases:\n \lim_{p \to \infty} \frac{\text{Vol}(S_p)}{\text{Vol}(C_p)} = 0\n - **Implication**: As feature dimensions grow, uniformly distributed sample mass concentrates almost entirely in the corners of the bounding hypercube. Data points become sparse and isolated near domain boundaries. Nearest neighbor estimation converts from local interpolation to unreliable boundary extrapolation.\n\n\n\n# Supervised Learning as Function Approximation\n\n- **Additive Error Statistical Models**:\n - Quantitative response data are modeled assuming an underlying statistical structure:\n Y = f(X) + \varepsilon\n where systematic function f(X) = E(Y \mid X)\varepsilonE(\varepsilon) = 0\n - Unexplained variance \varepsilon reflects unobserved input features or inherent stochastic system variability.\n - Non-deterministic systems allow multiple distinct target outputs YX = x\n\n- **Categorical Response Surface Modeling**:\n - Additive error formulations cannot represent qualitative targets (e.g., adding numerical noise to categorical text labels is undefined).\n - Instead, categorical responses are modeled by approximating conditional class probability distributions p(x) = \operatorname{Pr}(G \mid X = x).\n - For binary outcome coding (0/1p(x) = E(Y \mid X = x).\n\n\n\n- **Supervised Function Approximation Framework**:\n - Learning algorithm parameterize approximators f_\theta(x)\theta.\n - **Linear Basis Expansion Models**:\n - A flexible class of approximators expands functions as weighted sums of basis functions h_k(x):\n f_\theta(x) = \sum_{k=1}^K h_k(x) \theta_k\n - Basis choice dictates model architecture:\n - h_k(x) = x_k \implies Standard Linear Model\n - h_k(x) = x_k^m \implies Polynomial Surface Model\n - h_k(x) = \operatorname{ReLU}(w_k^T x + b_k) \implies Rectified Linear Neural Network\n - h_k(x) = \frac{1}{1 + \exp(-x^T \beta_k)} \implies Sigmoidal Neural Network\n\n\n\n- **Maximum Likelihood Estimation (MLE)**:\n - Parameters \thetaN independent samples:\n L(\theta) = \sum_{i=1}^N \log \operatorname{Pr}\theta(y_i)\n - Log-probabilities \log \operatorname{Pr}\theta(y_i)(-\infty, 0]L(\theta) selects parameters under which observed training data exhibits maximum occurrence probability.\n\n\n\n- **Equivalence of Least Squares and Maximum Likelihood**:\n - Assume additive error follows independent Gaussian noise \varepsilon \sim \mathcal{N}(0, \sigma^2), giving conditional likelihood:\n \operatorname{Pr}(Y \mid X, \theta) = \mathcal{N}\left( f_\theta(X), \sigma^2 \right)\n - The conditional log-likelihood equation expands as:\n L(\theta) = C - \frac{1}{2\sigma^2} \sum_{i=1}^N \left( y_i - f_\theta(x_i) \right)^2\n - Maximizing log-likelihood L(\theta)\thetaRSS(\theta) = \sum_{i=1}^N (y_i - f_\theta(x_i))^2$$