Computing for Business Analytics and Machine Learning Vocabulary Flashcards
Python Core Programming Concepts
Fundamental Python Data Types
int: Represents whole numbers without fractional components (e.g.,a = 10).float: Represents floating-point real numbers containing decimal points (e.g.,b = 3.14). Standard division/always yields afloat(e.g.,10 / 4 = 2.5).str: Represents sequence of characters wrapped in single or double quotes (e.g.,c = 'AI').list: Represents ordered, mutable sequences of elements that can hold heterogenous data types (e.g.,d = [1, 2]).dict: Represents key-value mappings contained within curly braces (e.g.,e = {'a': 1}).
Assignment vs. Comparison Operators
Assignment Operator (
=): Assigns a value or evaluated expression on the right-hand side to a variable on the left-hand side.
x = 7 # Assigns integer 7 to variable x ``` * **Equality Comparison Operator (`==`)**: Evaluates whether two values or expressions are equal, returning a Boolean (`True` or `False`).python print(x == 7) # Returns True ```
Relational Comparison Operators (
<,<=,>,>=,!=):<: Strictly less than.<=: Less than or equal to.>: Strictly greater than.>=: Greater than or equal to.!=: Not equal to.Example Execution:
a = 5 b = 10 print(a < b, a <= 5, b > 12) # Outputs: True True False ``` *Explanation*: `5 < 10` is `True`, `5 <= 5` is `True`, and `10 > 12` is `False`.python x = 15 y = 20 print(x != y and x >= 10) # Outputs: True
`` *Explanation*:15 != 20evaluates toTrue, and15 >= 10evaluates toTrue.True and TrueyieldsTrue`.Arithmetic Operators & Order of Operations
Addition (
+) and Subtraction (-): Evaluated from left to right at equal precedence.
x = 15 y = 4 print(x + y - 2) # Outputs: 17 ``` *Explanation*: Evaluates left-to-right: , then . * Multiplication (`*`) and Division (`/`): Take precedence over addition and subtraction.python a = 10 b = 4 print(a / b + a * 2) # Outputs: 22.5
`` *Explanation*:10 / 4evaluates to2.5,10 * 2evaluates to20`. Addition yields .Modulo Operator (
%): Returns the integer remainder of division between two numbers.
remainder = 7 % 3 # Evaluates to 1 ``` * **Exponentiation Operator (`**`)**: Raises the base value to the power of the exponent.python result = 2 ** 3 # Evaluates to 8 ```
In-Place Operators (e.g.,
*=): Performs an arithmetic operation on the variable's current value and updates the variable in place.
Decision Structures & Logical Evaluation
Two-Way Decisions: Employs
ifandelsebranches to direct control flow based on a single Boolean condition.Multi-Branch Decisions: Uses
if,elif, andelsestructures to test multiple non-overlapping conditions sequentially.Boolean Logic & Short-Circuit Evaluation: Logical operators (
and,or,not) evaluate Boolean expressions. Inandexpressions, if the first operand isFalse, evaluation stops immediately because the result must beFalse.
Loops & Control Flow
Fixed-Count Loops (
for): Iterates over a sequence or range a predetermined number of times.Conditional Loops (
while): Repeatedly executes a block of code as long as a Boolean condition remainsTrue.Sequence Generator (
range(stop)): Generates an arithmetic progression starting at , incrementing by , and stopping right before the specifiedstopinteger.
for i in range(5): print(i) # Prints 0, 1, 2, 3, 4 ``` * **While Loop State Updates**: A state variable within a `while` loop must be updated in the loop body to avoid infinite execution loops.python x = 1 while x < 8: x *= 2 # Updates x to prevent an infinite loop ```
Nested Logic Iteration: Pattern where an inner loop completes its full cycle of iterations for every single iteration of the outer loop.
python count = 0 for i in range(3): for j in range(2): count += 1 # Executes 3 x 2 = 6 times
Accumulator Pattern
Pattern used to compute running totals or cumulative values.
Requires explicit initialization of a variable (e.g.,
total = 0) before entering the loop, followed by an update statement (e.g.,total += value) inside the loop body.
Functions & Execution Flow
Purpose: Encapsulate reusable blocks of logic, improve code modularity, and reduce redundancy.
Defining and Calling Functions: Defined using
def function_name(parameters):and invoked using function call syntax.print()vs.return:print(): Only displays output text to the console; does not produce a value that can be captured or reused by caller scripts.return: Immediately terminates function execution and sends a value back to the caller for assignment or further processing.
def add(a, b): return a + b result = add(4, 8) # Stores 12 in result ``` * **Function Program Flow & Parameter Mapping**:python def double_val(x): return x * 2 val = 5 res = double_val(val) print(res) # Outputs: 10
`` *Program Execution Flow*: Variablevalis assigned5. Callingdouble_val(val)binds parameterxto5. The expressionx * 2evaluates to10and is returned to assign variableres, which prints as10`.Hard-Coding Errors in Functions: Occurs when a function uses fixed literal values instead of passing its declared parameters, causing invalid behavior regardless of input arguments.
python def calc_total(price, tax_rate): total = 100 * 0.05 # Error: hard-codes 100 and 0.05 instead of using price and tax_rate return total print(calc_total(50, 0.10)) # Always returns 5.0 regardless of arguments passed
Random Number Generation Concepts
Used in algorithms for simulation, sampling, and initialization.
Distinguishes between uniform integer generation within ranges and continuous floating-point generation.
Module Imports and Namespaces
import modulevs.from module import function:import numpy as np: Loads the library into its own namespace under the aliasnp, requiring module prefix notation (e.g.,np.sqrt(16)).from math import sqrt: Importssqrtdirectly into the local scope, eliminating the need for a module prefix (e.g.,sqrt(16)).Import Alias (
as): Syntax assigning shorthand names to imported libraries for cleaner calls.python import numpy as np print(np.sum([1, 2, 3])) # Uses 'np' alias
Console Input & Type Conversion Errors
input()Function: Reads user input from console strictly as a string datatype (str). String inputs must be explicitly cast using functions likeint()orfloat()before executing numeric operations.python user_val = input('Enter score: ') score = int(user_val) # Converts string input to integer Concatenation
TypeError: In Python, attempting to concatenate a string and an integer using the+operator causes aTypeError(can only concatenate str (not "int") to str).python age = 20 # print('Age is ' + age) # Raises TypeError print('Age is ' + str(age)) # Fixed via explicit string conversion print('Age is', age) # Fixed using comma separation in print()
Python Data Structures & Manipulations
Strings
String Slicing (
[start:stop]): Extracts a substring starting from indexstartup to but excluding indexstop.Negative Indexing: Accesses elements relative to the end of the string, where index
[-1]returns the final character.
s = 'Python' print(s[1:4]) # Outputs: 'yth' print(s[-1]) # Outputs: 'n' ``` * **String Methods (`.find()` and `.split()`)**: * `.split(separator)`: Splits a string into a list of substrings using a designated delimiter. * `.find(substring)`: Returns the zero-based starting index of the first matching substring, or `-1` if not found.python text = 'data,python,ai' parts = text.split(',') # Yields ['data', 'python', 'ai'] idx = text.find('python') # Returns integer index 5 print(parts[1], idx) # Outputs: python 5 ```
Lists
Ordered, mutable sequences capable of storing elements of varying data types.
Adding Elements (
.append()): Adds a single element to the end of a list in place.
nums = [1, 2] nums.append(3) # nums becomes [1, 2, 3] ``` * **Removing Elements (`list.pop()` vs. `list.remove()`)**: * `pop(index)`: Removes and returns the element at the specified index position. * `remove(value)`: Searches for and deletes the first matching instance of the specified value.python nums = [3, 8, 2] nums.pop(1) # Removes element at index 1 (value 8) nums.remove(3) # Removes first occurrence of value 3 ```
Dictionaries
Key-value stores optimized for fast lookups via keys.
Accessing values: Syntax
dict[key]extracts the associated value.python student = {'name': 'Alice', 'score': 95} print(student['score']) # Outputs: 95 Updating & Adding: Assigning a value to an existing key overwrites it (
student['score'] = 98); assigning to a non-existent key creates a new key-value pair (student['grade'] = 'A').
File Handling Process & Modes
Process follows three steps: Open file handle -> Perform read/write operations -> Close file handle.
File Modes in
open(filename, mode):'r': Read mode (default). Opens file for reading; throws error if file does not exist.'w': Write mode. Opens file for writing, overwriting existing content or creating a new file.'a': Append mode. Opens file for writing, preserving existing content and appending new data at the end.File Syntax Example:
python f = open('results.txt', 'a') f.write('Score: 90 ') f.close()
Relational Databases & SQL
Relational Database Principles
Organizes data into structured tables (relations) comprising rows (records) and columns (attributes).
Tables connect to one another through primary key and foreign key relationships.
Common SQL Commands
Data Querying:
SELECTretrieves specific attributes from tables.Data Filtering:
WHEREapplies conditions to filter rows.Data Insertion:
INSERT INTOadds new records.Data Modification:
UPDATEmodifies existing records.Data Deletion:
DELETEremoves records.
Database Normalization & Structure
Best practices mandate organizing data across tables to minimize data redundancy, protect data integrity, and prevent update/deletion anomalies.
Fast Numerical Computing with NumPy
NumPy Arrays vs. Python Lists
NumPy Arrays (
ndarray): Require all contained elements to share a single, fixed data type (homogeneous structure). Optimized in C for vector operations.Python Lists: Can store heterogeneous elements with dynamic sizing, incurring higher memory overhead.
Array Creation Functions
np.array(list): Converts a Python list into a NumPy array.np.ones((m, n)): Creates an array of shape populated entirely with float1.0values.np.zeros((m, n)): Creates an array of shape populated entirely with float0.0values.np.linspace(start, stop, num): Generatesnumevenly spaced values over the interval[start, stop](inclusive of endpoints by default).np.arange(stop): Returns evenly spaced values from up tostop.
Array Attributes
.shape: Returns a tuple containing array dimensions along each axis (e.g., rows and columns)..size: Returns the total count of elements across all axes..ndim: Returns the integer count of array dimensions (axes)..dtype: Identifies the specific data type of elements stored in the array.Example Execution:
python import numpy as np A = np.array([[1, 2, 3], [4, 5, 6]]) print(A.ndim, A.size) # Outputs: 2 6 Explanation:.ndimyields2(2D matrix) and.sizeyields6( total elements).
Automatic Array Upcasting
NumPy automatically casts elements to a common data type when initialized with mixed types to maintain array homogeneity.
python import numpy as np arr = np.array([1, 2, 3.5]) # Integers are upcast to floats: array([1. , 2. , 3.5])
Array Indexing & Slicing Syntax (
start:stop:step)Slicing extracts subsets where
startis inclusive,stopis exclusive, andstepdefines stride size.Reversing Arrays: A negative step
[::-1]reverses array elements.Strided Selection:
[::2]selects every second element.
x = np.arange(10) print(x[::2]) # Outputs: [0, 2, 4, 6, 8] print(x[::-1]) # Outputs: [9, 8, 7, 6, 5, 4, 3, 2, 1, 0] ``` * **2D Array Indexing (`[row, column]`)**: * `A[row, :]`: Selects an entire row. * `A[:, col]`: Selects an entire column.python A = np.array([[10, 20], [30, 40]]) print(A[1, 0]) # Outputs: 30 print(A[:, 1]) # Outputs: [20, 40] ```
Vectorized Element-Wise Operations
Operations execute on matching elements without explicit
forloops.Element-wise multiplication (
np.multiply()or*):
import numpy as np a = np.array([2, 3]) b = np.array([4, 5]) print(np.multiply(a, b)) # Outputs: [ 8 15] ``` * Scalar Operations & In-Place Multiplication:python import numpy as np arr = np.array([1, 2, 3]) arr *= 5 print(arr) # Outputs: [ 5 10 15] ```
Aggregation Functions & Axis Parameters
np.sum(arr): Sums all elements in array.np.mean(arr): Computes arithmetic mean across all elements.np.mean(arr, axis=0): Computes mean down columns.np.mean(arr, axis=1): Computes mean across rows.np.median(arr): Calculates median of array.Summation Example:
python arr = np.array([10, 20, 30]) print(np.sum(arr)) # Outputs: 60 (10 + 20 + 30)
Boolean Indexing & Mask Summation
Applying comparison operators to arrays returns a Boolean array (
mask) filled withTrueandFalsevalues.Passing a mask into an array filters for elements satisfying the condition:
arr[arr > 100].Calling
np.sum(mask)treatsTrueas1andFalseas0, effectively counting matching elements.python sales = np.array([400000, 600000]) mask = sales > 500000 print(np.sum(mask)) # Outputs: 1
Random Number Generation Functions
np.random.rand(n): Generates random floats drawn from uniform distribution over .np.random.normal(mean, std, size): Generates samples drawn from normal distribution defined by mean and standard deviation.np.random.randint(low, high, size): Generates random integers bounded in range[low, high).
NumPy Output Printing Format
NumPy outputs arrays formatted inside square brackets with space-separated values (e.g.,
[1 2 3]), unlike Python lists which use comma separation.
Data Manipulation & Analysis with Pandas
Pandas Library Purpose & Core Structures
Specialized library for tabular data manipulation, cleaning, and analysis using
DataFramestructures.
DataFrame Attributes vs. Methods
Attributes: Access structural properties without parentheses (e.g.,
df.shape).Methods: Perform computations or manipulations and require execution parentheses (e.g.,
df.head()).python import pandas as pd df = pd.read_csv('data.csv') print(df.shape) # Attribute: returns tuple (rows, columns) print(df.head()) # Method: returns top 5 rows
Data Loading & Header Configurations
pd.read_csv('filename.csv'): Parses tabular CSV files into aDataFrameobject.Handling Files without Headers: Use parameter
header=Noneto prevent the first line from being treated as titles, and pass explicit column labels usingnames.python import pandas as pd df = pd.read_csv('data.csv', header=None, names=['age', 'score'])
Creating Calculated Columns
New columns are constructed through vectorized arithmetic across existing columns.
python import pandas as pd df = pd.read_csv('sales.csv') df['revenue'] = df['units'] * df['price'] # Calculates row-by-row revenue
Pandas Boolean Filtering
Filters rows by passing logical condition masks inside DataFrame selection brackets.
python df['ppg'] = df['points'] / df['games'] high_scorers = df[df['ppg'] > 20] # Extracts rows where ppg exceeds 20
Data Selection Tools:
locvs.ilocdf.loc: Label-based selector using explicit row/column name labels.df.iloc: Position-based selector using integer index positions.python df.loc[:, 'age'] # Selects 'age' column by label df.iloc[:, 0] # Selects first column (index 0) by integer position
Data Visualization with Matplotlib
Matplotlib Library Purpose
Data visualization library used to create static scatter plots, line graphs, and histograms.
Scatter Plots (
plt.scatter())Used to compare numerical variables against x and y axes.
Plotting Series vs. Scalars:
Passing DataFrame columns/Series plots all dataset observations at once.
Passing single-element lists or scalars (e.g.,
plt.scatter([25], [50000], color='red')) plots individual data points.Scatter Plot Markers & Options:
python import matplotlib.pyplot as plt plt.scatter(df['age'], df['income'], marker='x') plt.show() Explanation: Places variableageon the x-axis andincomeon the y-axis, rendering datapoints as'x'markers.
Linear Regression & Predictive Modeling
Linear Model Representation
Represents mathematical relationships between continuous predictors and continuous target variables.

General Linear Regression Model Formula:
Model Components:
: Target outcome variable (dependent variable).
: Intercept / Constant term.
: Predictor variables (independent variables).
: Beta coefficients representing effect size per predictor.
: Error term / Random unobserved noise.
Single vs. Multiple Linear Regression
Simple (Single) Linear Regression: Predicts numeric outcome using a single predictor :
Multiple Linear Regression: Predicts outcome using two or more predictors .
Simple Linear Regression Mechanics & Formulas

Observed vs. Predicted Values:
Observed outcome value for observation :
Predicted outcome value for observation :
Random Error (Residual):
Prediction Formula for Unseen Data:
Parameter Mathematical Calculations:
Slope Coefficient ():
Intercept Term ():
Notation Definitions:
: Mean of predictor variable .
: Mean of outcome variable .
: Standard deviation of variable .
: Standard deviation of variable .
: Pearson correlation coefficient between and .
Pearson Correlation Coefficient ()
Quantifies direction and strength of linear relationship between variables and y$.\n \n \n\n * Exact Formula:\n r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}\n * Range and Interpretation Scale:\n \n \n\n * r = +1: Perfect positive linear relationship.\n * r = 0.9: Strong positive linear relationship.\n * r = 0.5: Weak positive linear relationship.\n * r = 0: No linear relationship.\n * r = -0.5: Weak negative linear relationship.\n * r = -0.9: Strong negative linear relationship.\n * r = -1: Perfect negative linear relationship.\n\n* **Model Performance Metrics**\n * **Coefficient of Determination (R^2)**:\n * Statistical metric indicating proportion of outcome variance explained by independent predictor variables.\n \n \n\n * Formula:\n R^2 = 1 - \frac{SS_{RES}}{SS_{TOT}} = 1 - \frac{\sum_i (y_i - \hat{y}i)^2}{\sum_i (y_i - \bar{y})^2}\n * Properties:\n * Ranges strictly between 01.\n * Higher values indicate superior model fit on training data.\n * In Simple Linear Regression, R^2r^2).\n * **Root Mean Squared Error (RMSE)**:\n * Measures standard deviation of model prediction errors in original target units.\n \n \n\n * Formula:\n RMSE = \sqrt{\frac{\sum{i=1}^n (\hat{y}_i - y_i)^2}{n}}\n * Where \hat{y}_1, \hat{y}_2, \dots, \hat{y}_ny_1, y_2, \dots, y_nn is the total observation count.\n\n* **Overfitting & Data Partitioning**\n * **The Overfitting Problem**: Statistical models can produce complex equations fitting noise in training data perfectly (100\% fit). When applied to unseen data, highly complex models fail.\n \n \n\n * **Data Partitioning Solution**: Datasets are partitioned into distinct subsets to evaluate performance on unseen data:\n \n \n\n * **Training Partition**: Used by algorithms to learn parameters and build candidate models.\n * **Validation / Test Partition**: Unseen data partition used to evaluate performance metrics and select optimal models.\n\n# Machine Learning with Scikit-Learn\n\n* **Scikit-Learn (`sklearn`) Library Purpose**\n * Python library for machine learning models, dataset partitioning, preprocessing, and evaluation metrics.\n\n* **Dataset Partitioning (`train_test_split`)**\n * Utility from `sklearn.model_selection` partitioning features Xy into training and testing sets.
python from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split(X, y)
Scikit-Learn Workflow Paradigm:
.fit()vs..predict().fit(X_train, y_train): Model learns parameters and coefficients from training subset..predict(X_test): Applies learned parameters to unseen features to generate predictions.python model.fit(X_train, y_train) # Learns parameters predictions = model.predict(X_test) # Generates predictions
Feature Scaling Principles & Implementation
Rationale: Rescales numeric features onto comparable scales so variables with large ranges (e.g., salary in tens of thousands) do not dominate distance calculations over variables with smaller ranges (e.g., age in tens).
Standard scaling transforms features to zero mean and unit variance.
python from sklearn.preprocessing import StandardScaler scaler = StandardScaler() X_scaled = scaler.fit_transform(X)
k-Nearest Neighbors (kk$$ nearest known neighbors using distance metrics.
python from sklearn.neighbors import KNeighborsClassifier m = KNeighborsClassifier(n_neighbors=5) m.fit(X_train, y_train) yhat = m.predict(X_test)
k-Means Unsupervised Clustering
Unsupervised algorithm grouping unlabelled observations into clusters surrounding computed cluster centroids.
python from sklearn.cluster import KMeans m = KMeans(n_clusters=3) m.fit(X) labels = m.labels_
Linear Regression Implementation & Parameter Extraction
Model fitting, parameter extraction, and evaluation scoring:
python from sklearn.linear_model import LinearRegression m = LinearRegression() m.fit(X_train, y_train) print(m.intercept_, m.coef_) # Prints intercept term and array of predictor beta coefficients r2 = m.score(X_test, y_test) # Computes R^2 model-fit score on test partition print(r2)
Computing Performance Metrics in Code
Evaluating regression performance using
sklearn.metricsroutines:python from sklearn.metrics import root_mean_squared_error, r2_score rmse = root_mean_squared_error(y_test, y_pred) r2 = r2_score(y_test, y_pred)