Computing for Business Analytics and Machine Learning Vocabulary Flashcards

Python Core Programming Concepts

  • Fundamental Python Data Types

    • int: Represents whole numbers without fractional components (e.g., a = 10).

    • float: Represents floating-point real numbers containing decimal points (e.g., b = 3.14). Standard division / always yields a float (e.g., 10 / 4 = 2.5).

    • str: Represents sequence of characters wrapped in single or double quotes (e.g., c = 'AI').

    • list: Represents ordered, mutable sequences of elements that can hold heterogenous data types (e.g., d = [1, 2]).

    • dict: Represents key-value mappings contained within curly braces (e.g., e = {'a': 1}).

  • Assignment vs. Comparison Operators

    • Assignment Operator (=): Assigns a value or evaluated expression on the right-hand side to a variable on the left-hand side.

    x = 7  # Assigns integer 7 to variable x
        ```
    * **Equality Comparison Operator (`==`)**: Evaluates whether two values or expressions are equal, returning a Boolean (`True` or `False`).
    

    python print(x == 7) # Returns True     ```

    • Relational Comparison Operators (<, <=, >, >=, !=):

    • <: Strictly less than.

    • <=: Less than or equal to.

    • >: Strictly greater than.

    • >=: Greater than or equal to.

    • !=: Not equal to.

    • Example Execution:

      a = 5
      b = 10
      print(a < b, a <= 5, b > 12)  # Outputs: True True False
    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;```
    &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;*Explanation*: `5 < 10` is `True`, `5 <= 5` is `True`, and `10 > 12` is `False`.
    

    python x = 15 y = 20 print(x != y and x >= 10) # Outputs: True       `` &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;*Explanation*:15 != 20evaluates toTrue, and15 >= 10evaluates toTrue.True and TrueyieldsTrue`.

  • Arithmetic Operators & Order of Operations

    • Addition (+) and Subtraction (-): Evaluated from left to right at equal precedence.

    x = 15
    y = 4
    print(x + y - 2)  # Outputs: 17
    &nbsp;&nbsp;&nbsp;&nbsp;```
    &nbsp;&nbsp;&nbsp;&nbsp;*Explanation*: Evaluates left-to-right: 15+4=1915 + 4 = 19, then 19−2=1719 - 2 = 17.
    * Multiplication (`*`) and Division (`/`): Take precedence over addition and subtraction.
    

    python a = 10 b = 4 print(a / b + a * 2) # Outputs: 22.5     `` &nbsp;&nbsp;&nbsp;&nbsp;*Explanation*:10 / 4evaluates to2.5,10 * 2evaluates to20`. Addition yields 2.5+20=22.52.5 + 20 = 22.5.

    • Modulo Operator (%): Returns the integer remainder of division between two numbers.

    remainder = 7 % 3  # Evaluates to 1
    &nbsp;&nbsp;&nbsp;&nbsp;```
    * **Exponentiation Operator (`**`)**: Raises the base value to the power of the exponent.
    

    python result = 2 ** 3 # Evaluates to 8     ```

    • In-Place Operators (e.g., *=): Performs an arithmetic operation on the variable's current value and updates the variable in place.

  • Decision Structures & Logical Evaluation

    • Two-Way Decisions: Employs if and else branches to direct control flow based on a single Boolean condition.

    • Multi-Branch Decisions: Uses if, elif, and else structures to test multiple non-overlapping conditions sequentially.

    • Boolean Logic & Short-Circuit Evaluation: Logical operators (and, or, not) evaluate Boolean expressions. In and expressions, if the first operand is False, evaluation stops immediately because the result must be False.

  • Loops & Control Flow

    • Fixed-Count Loops (for): Iterates over a sequence or range a predetermined number of times.

    • Conditional Loops (while): Repeatedly executes a block of code as long as a Boolean condition remains True.

    • Sequence Generator (range(stop)): Generates an arithmetic progression starting at 00, incrementing by 11, and stopping right before the specified stop integer.

    for i in range(5):
        print(i)  # Prints 0, 1, 2, 3, 4
    &nbsp;&nbsp;&nbsp;&nbsp;```
    * **While Loop State Updates**: A state variable within a `while` loop must be updated in the loop body to avoid infinite execution loops.
    

    python x = 1 while x < 8: x *= 2 # Updates x to prevent an infinite loop     ```

    • Nested Logic Iteration: Pattern where an inner loop completes its full cycle of iterations for every single iteration of the outer loop. python count = 0 for i in range(3): for j in range(2): count += 1 # Executes 3 x 2 = 6 times &nbsp;&nbsp;&nbsp;&nbsp;

  • Accumulator Pattern

    • Pattern used to compute running totals or cumulative values.

    • Requires explicit initialization of a variable (e.g., total = 0) before entering the loop, followed by an update statement (e.g., total += value) inside the loop body.

  • Functions & Execution Flow

    • Purpose: Encapsulate reusable blocks of logic, improve code modularity, and reduce redundancy.

    • Defining and Calling Functions: Defined using def function_name(parameters): and invoked using function call syntax.

    • print() vs. return:

    • print(): Only displays output text to the console; does not produce a value that can be captured or reused by caller scripts.

    • return: Immediately terminates function execution and sends a value back to the caller for assignment or further processing.

    def add(a, b):
        return a + b
    result = add(4, 8)  # Stores 12 in result
    &nbsp;&nbsp;&nbsp;&nbsp;```
    * **Function Program Flow & Parameter Mapping**:
    

    python def double_val(x): return x * 2 val = 5 res = double_val(val) print(res) # Outputs: 10     `` &nbsp;&nbsp;&nbsp;&nbsp;*Program Execution Flow*: Variablevalis assigned5. Callingdouble_val(val)binds parameterxto5. The expressionx * 2evaluates to10and is returned to assign variableres, which prints as10`.

    • Hard-Coding Errors in Functions: Occurs when a function uses fixed literal values instead of passing its declared parameters, causing invalid behavior regardless of input arguments. python def calc_total(price, tax_rate): total = 100 * 0.05 # Error: hard-codes 100 and 0.05 instead of using price and tax_rate return total print(calc_total(50, 0.10)) # Always returns 5.0 regardless of arguments passed &nbsp;&nbsp;&nbsp;&nbsp;

  • Random Number Generation Concepts

    • Used in algorithms for simulation, sampling, and initialization.

    • Distinguishes between uniform integer generation within ranges and continuous floating-point generation.

  • Module Imports and Namespaces

    • import module vs. from module import function:

    • import numpy as np: Loads the library into its own namespace under the alias np, requiring module prefix notation (e.g., np.sqrt(16)).

    • from math import sqrt: Imports sqrt directly into the local scope, eliminating the need for a module prefix (e.g., sqrt(16)).

    • Import Alias (as): Syntax assigning shorthand names to imported libraries for cleaner calls. python import numpy as np print(np.sum([1, 2, 3])) # Uses 'np' alias &nbsp;&nbsp;&nbsp;&nbsp;

  • Console Input & Type Conversion Errors

    • input() Function: Reads user input from console strictly as a string datatype (str). String inputs must be explicitly cast using functions like int() or float() before executing numeric operations. python user_val = input('Enter score: ') score = int(user_val) # Converts string input to integer &nbsp;&nbsp;&nbsp;&nbsp;

    • Concatenation TypeError: In Python, attempting to concatenate a string and an integer using the + operator causes a TypeError (can only concatenate str (not "int") to str). python age = 20 # print('Age is ' + age) # Raises TypeError print('Age is ' + str(age)) # Fixed via explicit string conversion print('Age is', age) # Fixed using comma separation in print() &nbsp;&nbsp;&nbsp;&nbsp;

Python Data Structures & Manipulations

  • Strings

    • String Slicing ([start:stop]): Extracts a substring starting from index start up to but excluding index stop.

    • Negative Indexing: Accesses elements relative to the end of the string, where index [-1] returns the final character.

    s = 'Python'
    print(s[1:4])  # Outputs: 'yth'
    print(s[-1])   # Outputs: 'n'
    &nbsp;&nbsp;&nbsp;&nbsp;```
    * **String Methods (`.find()` and `.split()`)**:
    * `.split(separator)`: Splits a string into a list of substrings using a designated delimiter.
    * `.find(substring)`: Returns the zero-based starting index of the first matching substring, or `-1` if not found.
    

    python text = 'data,python,ai' parts = text.split(',') # Yields ['data', 'python', 'ai'] idx = text.find('python') # Returns integer index 5 print(parts[1], idx) # Outputs: python 5     ```

  • Lists

    • Ordered, mutable sequences capable of storing elements of varying data types.

    • Adding Elements (.append()): Adds a single element to the end of a list in place.

    nums = [1, 2]
    nums.append(3)  # nums becomes [1, 2, 3]
    &nbsp;&nbsp;&nbsp;&nbsp;```
    * **Removing Elements (`list.pop()` vs. `list.remove()`)**:
    * `pop(index)`: Removes and returns the element at the specified index position.
    * `remove(value)`: Searches for and deletes the first matching instance of the specified value.
    

    python nums = [3, 8, 2] nums.pop(1) # Removes element at index 1 (value 8) nums.remove(3) # Removes first occurrence of value 3     ```

  • Dictionaries

    • Key-value stores optimized for fast lookups via keys.

    • Accessing values: Syntax dict[key] extracts the associated value. python student = {'name': 'Alice', 'score': 95} print(student['score']) # Outputs: 95 &nbsp;&nbsp;&nbsp;&nbsp;

    • Updating & Adding: Assigning a value to an existing key overwrites it (student['score'] = 98); assigning to a non-existent key creates a new key-value pair (student['grade'] = 'A').

  • File Handling Process & Modes

    • Process follows three steps: Open file handle -> Perform read/write operations -> Close file handle.

    • File Modes in open(filename, mode):

    • 'r': Read mode (default). Opens file for reading; throws error if file does not exist.

    • 'w': Write mode. Opens file for writing, overwriting existing content or creating a new file.

    • 'a': Append mode. Opens file for writing, preserving existing content and appending new data at the end.

    • File Syntax Example: python f = open('results.txt', 'a') f.write('Score: 90 ') f.close() &nbsp;&nbsp;&nbsp;&nbsp;

Relational Databases & SQL

  • Relational Database Principles

    • Organizes data into structured tables (relations) comprising rows (records) and columns (attributes).

    • Tables connect to one another through primary key and foreign key relationships.

  • Common SQL Commands

    • Data Querying: SELECT retrieves specific attributes from tables.

    • Data Filtering: WHERE applies conditions to filter rows.

    • Data Insertion: INSERT INTO adds new records.

    • Data Modification: UPDATE modifies existing records.

    • Data Deletion: DELETE removes records.

  • Database Normalization & Structure

    • Best practices mandate organizing data across tables to minimize data redundancy, protect data integrity, and prevent update/deletion anomalies.

Fast Numerical Computing with NumPy

  • NumPy Arrays vs. Python Lists

    • NumPy Arrays (ndarray): Require all contained elements to share a single, fixed data type (homogeneous structure). Optimized in C for vector operations.

    • Python Lists: Can store heterogeneous elements with dynamic sizing, incurring higher memory overhead.

  • Array Creation Functions

    • np.array(list): Converts a Python list into a NumPy array.

    • np.ones((m, n)): Creates an array of shape (m,n)(m, n) populated entirely with float 1.0 values.

    • np.zeros((m, n)): Creates an array of shape (m,n)(m, n) populated entirely with float 0.0 values.

    • np.linspace(start, stop, num): Generates num evenly spaced values over the interval [start, stop] (inclusive of endpoints by default).

    • np.arange(stop): Returns evenly spaced values from 00 up to stop.

  • Array Attributes

    • .shape: Returns a tuple containing array dimensions along each axis (e.g., rows and columns).

    • .size: Returns the total count of elements across all axes.

    • .ndim: Returns the integer count of array dimensions (axes).

    • .dtype: Identifies the specific data type of elements stored in the array.

    • Example Execution: python import numpy as np A = np.array([[1, 2, 3], [4, 5, 6]]) print(A.ndim, A.size) # Outputs: 2 6 &nbsp;&nbsp;&nbsp;&nbsp;     Explanation: .ndim yields 2 (2D matrix) and .size yields 6 (2×3=62 \times 3 = 6 total elements).

  • Automatic Array Upcasting

    • NumPy automatically casts elements to a common data type when initialized with mixed types to maintain array homogeneity. python import numpy as np arr = np.array([1, 2, 3.5]) # Integers are upcast to floats: array([1. , 2. , 3.5]) &nbsp;&nbsp;&nbsp;&nbsp;

  • Array Indexing & Slicing Syntax (start:stop:step)

    • Slicing extracts subsets where start is inclusive, stop is exclusive, and step defines stride size.

    • Reversing Arrays: A negative step [::-1] reverses array elements.

    • Strided Selection: [::2] selects every second element.

    x = np.arange(10)
    print(x[::2])   # Outputs: [0, 2, 4, 6, 8]
    print(x[::-1])  # Outputs: [9, 8, 7, 6, 5, 4, 3, 2, 1, 0]
    &nbsp;&nbsp;&nbsp;&nbsp;```
    * **2D Array Indexing (`[row, column]`)**:
    * `A[row, :]`: Selects an entire row.
    * `A[:, col]`: Selects an entire column.
    

    python A = np.array([[10, 20], [30, 40]]) print(A[1, 0]) # Outputs: 30 print(A[:, 1]) # Outputs: [20, 40]     ```

  • Vectorized Element-Wise Operations

    • Operations execute on matching elements without explicit for loops.

    • Element-wise multiplication (np.multiply() or *):

    import numpy as np
    a = np.array([2, 3])
    b = np.array([4, 5])
    print(np.multiply(a, b))  # Outputs: [ 8 15]
    &nbsp;&nbsp;&nbsp;&nbsp;```
    * Scalar Operations & In-Place Multiplication:
    

    python import numpy as np arr = np.array([1, 2, 3]) arr *= 5 print(arr) # Outputs: [ 5 10 15]     ```

  • Aggregation Functions & Axis Parameters

    • np.sum(arr): Sums all elements in array.

    • np.mean(arr): Computes arithmetic mean across all elements.

    • np.mean(arr, axis=0): Computes mean down columns.

    • np.mean(arr, axis=1): Computes mean across rows.

    • np.median(arr): Calculates median of array.

    • Summation Example: python arr = np.array([10, 20, 30]) print(np.sum(arr)) # Outputs: 60 (10 + 20 + 30) &nbsp;&nbsp;&nbsp;&nbsp;

  • Boolean Indexing & Mask Summation

    • Applying comparison operators to arrays returns a Boolean array (mask) filled with True and False values.

    • Passing a mask into an array filters for elements satisfying the condition: arr[arr > 100].

    • Calling np.sum(mask) treats True as 1 and False as 0, effectively counting matching elements. python sales = np.array([400000, 600000]) mask = sales > 500000 print(np.sum(mask)) # Outputs: 1 &nbsp;&nbsp;&nbsp;&nbsp;

  • Random Number Generation Functions

    • np.random.rand(n): Generates nn random floats drawn from uniform distribution over [0,1)[0, 1).

    • np.random.normal(mean, std, size): Generates samples drawn from normal distribution defined by mean and standard deviation.

    • np.random.randint(low, high, size): Generates random integers bounded in range [low, high).

  • NumPy Output Printing Format

    • NumPy outputs arrays formatted inside square brackets with space-separated values (e.g., [1 2 3]), unlike Python lists which use comma separation.

Data Manipulation & Analysis with Pandas

  • Pandas Library Purpose & Core Structures

    • Specialized library for tabular data manipulation, cleaning, and analysis using DataFrame structures.

  • DataFrame Attributes vs. Methods

    • Attributes: Access structural properties without parentheses (e.g., df.shape).

    • Methods: Perform computations or manipulations and require execution parentheses (e.g., df.head()). python import pandas as pd df = pd.read_csv('data.csv') print(df.shape) # Attribute: returns tuple (rows, columns) print(df.head()) # Method: returns top 5 rows &nbsp;&nbsp;&nbsp;&nbsp;

  • Data Loading & Header Configurations

    • pd.read_csv('filename.csv'): Parses tabular CSV files into a DataFrame object.

    • Handling Files without Headers: Use parameter header=None to prevent the first line from being treated as titles, and pass explicit column labels using names. python import pandas as pd df = pd.read_csv('data.csv', header=None, names=['age', 'score']) &nbsp;&nbsp;&nbsp;&nbsp;

  • Creating Calculated Columns

    • New columns are constructed through vectorized arithmetic across existing columns. python import pandas as pd df = pd.read_csv('sales.csv') df['revenue'] = df['units'] * df['price'] # Calculates row-by-row revenue &nbsp;&nbsp;&nbsp;&nbsp;

  • Pandas Boolean Filtering

    • Filters rows by passing logical condition masks inside DataFrame selection brackets. python df['ppg'] = df['points'] / df['games'] high_scorers = df[df['ppg'] > 20] # Extracts rows where ppg exceeds 20 &nbsp;&nbsp;&nbsp;&nbsp;

  • Data Selection Tools: loc vs. iloc

    • df.loc: Label-based selector using explicit row/column name labels.

    • df.iloc: Position-based selector using integer index positions. python df.loc[:, 'age'] # Selects 'age' column by label df.iloc[:, 0] # Selects first column (index 0) by integer position &nbsp;&nbsp;&nbsp;&nbsp;

Data Visualization with Matplotlib

  • Matplotlib Library Purpose

    • Data visualization library used to create static scatter plots, line graphs, and histograms.

  • Scatter Plots (plt.scatter())

    • Used to compare numerical variables against x and y axes.

    • Plotting Series vs. Scalars:

    • Passing DataFrame columns/Series plots all dataset observations at once.

    • Passing single-element lists or scalars (e.g., plt.scatter([25], [50000], color='red')) plots individual data points.

    • Scatter Plot Markers & Options: python import matplotlib.pyplot as plt plt.scatter(df['age'], df['income'], marker='x') plt.show() &nbsp;&nbsp;&nbsp;&nbsp;     Explanation: Places variable age on the x-axis and income on the y-axis, rendering datapoints as 'x' markers.

Linear Regression & Predictive Modeling

  • Linear Model Representation

    • Represents mathematical relationships between continuous predictors and continuous target variables.      

      Linear Regression Equation
    • General Linear Regression Model Formula:     Y=β0+β1x1+β2x2+⋯+βpxp+ϵY = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + \dots + \beta_p x_p + \epsilon

    • Model Components:

    • YY: Target outcome variable (dependent variable).

    • β0\beta_0: Intercept / Constant term.

    • x1,x2,…,xpx_1, x_2, \dots, x_p: Predictor variables (independent variables).

    • β1,β2,…,βp\beta_1, \beta_2, \dots, \beta_p: Beta coefficients representing effect size per predictor.

    • ϵ\epsilon: Error term / Random unobserved noise.

  • Single vs. Multiple Linear Regression

    • Simple (Single) Linear Regression: Predicts numeric outcome YY using a single predictor xx:     y=β0+β1xy = \beta_0 + \beta_1 x

    • Multiple Linear Regression: Predicts outcome YY using two or more predictors x1,x2,…,xpx_1, x_2, \dots, x_p.

  • Simple Linear Regression Mechanics & Formulas      

    Simple Linear Regression Plot
    • Observed vs. Predicted Values:

    • Observed outcome value for observation ii: yiy_i

    • Predicted outcome value for observation ii: y^i\hat{y}_i

    • Random Error (Residual): ϵi=yi−y^i\epsilon_i = y_i - \hat{y}_i

    • Prediction Formula for Unseen Data:     Y^new=β0+β1xnew\hat{Y}_{new} = \beta_0 + \beta_1 x_{new}

    • Parameter Mathematical Calculations:

    • Slope Coefficient (β1\beta_1):       β1=r×sysx\beta_1 = r \times \frac{s_y}{s_x}

    • Intercept Term (β0\beta_0):       β0=yˉ−β1xˉ\beta_0 = \bar{y} - \beta_1 \bar{x}

    • Notation Definitions:

      • xˉ\bar{x}: Mean of predictor variable xx.

      • yˉ\bar{y}: Mean of outcome variable yy.

      • sxs_x: Standard deviation of variable xx.

      • sys_y: Standard deviation of variable yy.

      • rr: Pearson correlation coefficient between xx and yy.

  • Pearson Correlation Coefficient (rr)

    • Quantifies direction and strength of linear relationship between variables xx and y$.\n  \n  ![Correlation Coefficient Formula](https://assets.knowt.com/pdf-flow-prod/a055a258-9c5d-4025-8682-f5fa090bf4fb-figures/3.png)\n\n * Exact Formula:\n    r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}\n * Range and Interpretation Scale:\n    \n    ![Correlation Visualizations](https://assets.knowt.com/pdf-flow-prod/a055a258-9c5d-4025-8682-f5fa090bf4fb-figures/4.png)\n\n * r = +1: Perfect positive linear relationship.\n * r = 0.9: Strong positive linear relationship.\n * r = 0.5: Weak positive linear relationship.\n * r = 0: No linear relationship.\n * r = -0.5: Weak negative linear relationship.\n * r = -0.9: Strong negative linear relationship.\n * r = -1: Perfect negative linear relationship.\n\n* **Model Performance Metrics**\n * **Coefficient of Determination (R^2)**:\n * Statistical metric indicating proportion of outcome variance explained by independent predictor variables.\n    \n    ![R Squared Formula](https://assets.knowt.com/pdf-flow-prod/a055a258-9c5d-4025-8682-f5fa090bf4fb-figures/6.png)\n\n * Formula:\n      R^2 = 1 - \frac{SS_{RES}}{SS_{TOT}} = 1 - \frac{\sum_i (y_i - \hat{y}i)^2}{\sum_i (y_i - \bar{y})^2}\n * Properties:\n * Ranges strictly between 0andand1.\n * Higher values indicate superior model fit on training data.\n * In Simple Linear Regression, R^2equalsthesquaredcorrelationcoefficient(equals the squared correlation coefficient (r^2).\n * **Root Mean Squared Error (RMSE)**:\n * Measures standard deviation of model prediction errors in original target units.\n    \n    ![RMSE Formula](https://assets.knowt.com/pdf-flow-prod/a055a258-9c5d-4025-8682-f5fa090bf4fb-figures/7.png)\n\n * Formula:\n      RMSE = \sqrt{\frac{\sum{i=1}^n (\hat{y}_i - y_i)^2}{n}}\n * Where \hat{y}_1, \hat{y}_2, \dots, \hat{y}_narepredictedvalues,are predicted values,y_1, y_2, \dots, y_nareobservedvalues,andare observed values, andn is the total observation count.\n\n* **Overfitting & Data Partitioning**\n * **The Overfitting Problem**: Statistical models can produce complex equations fitting noise in training data perfectly (100\% fit). When applied to unseen data, highly complex models fail.\n  \n  ![Overfitting Curve](https://assets.knowt.com/pdf-flow-prod/a055a258-9c5d-4025-8682-f5fa090bf4fb-figures/15.png)\n\n * **Data Partitioning Solution**: Datasets are partitioned into distinct subsets to evaluate performance on unseen data:\n    \n    ![Data Partitioning](https://assets.knowt.com/pdf-flow-prod/a055a258-9c5d-4025-8682-f5fa090bf4fb-figures/16.png)\n\n * **Training Partition**: Used by algorithms to learn parameters and build candidate models.\n * **Validation / Test Partition**: Unseen data partition used to evaluate performance metrics and select optimal models.\n\n# Machine Learning with Scikit-Learn\n\n* **Scikit-Learn (`sklearn`) Library Purpose**\n * Python library for machine learning models, dataset partitioning, preprocessing, and evaluation metrics.\n\n* **Dataset Partitioning (`train_test_split`)**\n * Utility from `sklearn.model_selection` partitioning features Xandtargetand targety into training and testing sets. python from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split(X, y) &nbsp;&nbsp;&nbsp;&nbsp;

  • Scikit-Learn Workflow Paradigm: .fit() vs. .predict()

    • .fit(X_train, y_train): Model learns parameters and coefficients from training subset.

    • .predict(X_test): Applies learned parameters to unseen features to generate predictions. python model.fit(X_train, y_train) # Learns parameters predictions = model.predict(X_test) # Generates predictions &nbsp;&nbsp;&nbsp;&nbsp;

  • Feature Scaling Principles & Implementation

    • Rationale: Rescales numeric features onto comparable scales so variables with large ranges (e.g., salary in tens of thousands) do not dominate distance calculations over variables with smaller ranges (e.g., age in tens).

    • Standard scaling transforms features to zero mean and unit variance. python from sklearn.preprocessing import StandardScaler scaler = StandardScaler() X_scaled = scaler.fit_transform(X) &nbsp;&nbsp;&nbsp;&nbsp;

  • k-Nearest Neighbors (k−NN)SupervisedClassification</strong></p><ul><li><p>Supervisedalgorithmclassifyingnewobservationsbasedonmajorityvoteof-NN) Supervised Classification</strong></p><ul><li><p>Supervised algorithm classifying new observations based on majority vote ofk$$ nearest known neighbors using distance metrics. python from sklearn.neighbors import KNeighborsClassifier m = KNeighborsClassifier(n_neighbors=5) m.fit(X_train, y_train) yhat = m.predict(X_test) &nbsp;&nbsp;&nbsp;&nbsp;

  • k-Means Unsupervised Clustering

    • Unsupervised algorithm grouping unlabelled observations into clusters surrounding computed cluster centroids. python from sklearn.cluster import KMeans m = KMeans(n_clusters=3) m.fit(X) labels = m.labels_ &nbsp;&nbsp;&nbsp;&nbsp;

  • Linear Regression Implementation & Parameter Extraction

    • Model fitting, parameter extraction, and evaluation scoring: python from sklearn.linear_model import LinearRegression m = LinearRegression() m.fit(X_train, y_train) print(m.intercept_, m.coef_) # Prints intercept term and array of predictor beta coefficients r2 = m.score(X_test, y_test) # Computes R^2 model-fit score on test partition print(r2) &nbsp;&nbsp;&nbsp;&nbsp;

  • Computing Performance Metrics in Code

    • Evaluating regression performance using sklearn.metrics routines: python from sklearn.metrics import root_mean_squared_error, r2_score rmse = root_mean_squared_error(y_test, y_pred) r2 = r2_score(y_test, y_pred) &nbsp;&nbsp;&nbsp;&nbsp;