Data Mining and Data Administration Exhaustive Study Notes
First Block — Week 1: Decision Support Systems
Evolution and Definition of DSS
Decision Support Systems (DSS) emerged in the 1960s out of MIT research on interactive computer systems and Herbert Simon's theory of decision-making. Through the 1970s and 1980s, key foundational texts (such as those by Steven Alter, and Ralph H. Sprague & Hugh J. Carlson) established DSS as a distinct academic and practical field. During this era, specialised DSS variants appeared, including group DSS prototypes, executive information systems (EIS), and expert systems.
In the 1990s, the field expanded with the introduction of business intelligence (BI), online analytical processing (OLAP), data warehousing, web portals, and data mining — technologies that feed into or overlap with modern DSS.
A Decision Support System (DSS) is broadly defined as an interactive, computer-based system that helps decision-makers use data, documents, knowledge, and models to solve problems and make decisions.
Specific definitions include:
Gorry and Scott-Morton: A system that helps decision-makers solve unstructured problems.
Keen and Scott-Morton: Coupling the intellectual resources of a person with computer capability to improve the quality of semi-structured decisions.
DSS are ancillary systems: they support and augment human judgement; they do not replace or automate the human decision-maker (which is the domain of automated decision-making or expert systems that handle routine tasks).
Why DSS Matter (Benefits)
Managers require timely, accurate, relevant, and complete information. Standard Transaction Processing Systems (TPS) only provide routine operational reports on current activities. DSS add value by allowing managers to:
Analyse complex situations using quantitative and qualitative models.
Retrieve and manipulate historical and external data on demand.
Run what-if, goal-seeking, and scenario analyses.
Support judgement-heavy, semi-structured, or unstructured decisions.
Key benefits include:
Improved decision effectiveness (focusing on accuracy, quality, and timeliness rather than raw operational efficiency).
Competitive advantage through superior strategic planning and rapid response to market changes.
Adaptability and flexibility to meet changing organisational needs.
Support across all organisational levels (from line managers to top executives) for individuals, groups, and virtual teams.
Characteristics and Capabilities of DSS
Core capabilities defined by Alter, Power & Sharda, and Sharda et al. include:
Support mainly semi-structured and unstructured problems that standard quantitative tools or transaction systems cannot fully solve.
Support all managerial levels, ranging from operational supervisors to executive leadership.
Support individuals as well as collaborative groups (including virtual teams using web-based tools).
Support interdependent or sequential decisions across different departments.
Support all phases of the decision-making process: intelligence, design, choice, and implementation.
High flexibility and adaptability: users can add, delete, or reconfigure components to suit related problem domains.
User-friendly graphical, natural-language, or web/mobile interfaces.
Preservation of decision-maker control: the DSS supports and advises but never overrides human judgement.
Development flexibility: simple DSS can be built by end users (e.g., spreadsheet models), whereas enterprise DSS are built by IS specialists integrating data warehouses, OLAP, and data mining.
Extensive model use to allow experimentation with different strategies under varying scenarios.
DSS vs Decision Automation vs Transaction Processing
Transaction Processing Systems (TPS): Automate routine transaction recording. Their primary focus is data integrity, consistency, and efficient record-keeping. They are used continuously by clerical and operational staff.
Decision Automation: Uses software algorithms to fully make and execute programmed decisions for well-structured, routine situations, removing the human entirely from the loop.
Decision Support Systems (DSS): Sit between TPS and decision automation. They are used on demand by managers and analysts to support (rather than replace) human judgement on flexible, semi-structured or unstructured problems by drawing upon historical, internal, and external data.
Gorry & Scott-Morton Framework
Gorry and Scott-Morton (1971) synthesized decision characteristics into a matrix combining two axes:
Degree of Structuredness (Rows):
Structured Decisions: Routine, repetitive problems with standard, known solution procedures (e.g., inventory reorder points, standard make-or-buy decisions). These can often be automated.
Unstructured Decisions: Fuzzy, complex problems with no predefined solution procedure. Objectives and solution paths are unclear, relying heavily on intuition, human judgement, and collaboration tools.
Semi-structured Decisions: Combine structured and unstructured elements (e.g., setting a marketing budget, capital acquisition analysis). Management science models handle part of the problem, while a DSS provides information to support human judgement for the unstructured remainder.
Type of Managerial Activity / Control (Columns, based on Anthony's taxonomy):
Operational Control: Efficient execution of specific tasks.
Management Control: Acquisition and efficient utilisation of resources to achieve organizational goals.
Strategic Planning: Establishing long-range goals, policies, and resource allocation strategies.
Insight of the framework: Traditional MIS and management science tools suffice for structured operational problems, but semi-structured and unstructured problems across management control and strategic planning necessitate interactive DSS.
The Decision-Making Process (Simon's Four Phases)
Herbert Simon's model represents the classic rational decision-making framework:
Intelligence: Scanning the environment to identify conditions requiring a decision. Includes problem identification, problem classification (placing the problem into a known category with standard solution approaches), problem decomposition (breaking complex problems into manageable sub-problems), and establishing problem ownership (determining who is responsible for taking action).
Design: Developing and analysing possible courses of action. Involves constructing a model (a simplified representation of reality) and establishing a principle of choice (the criterion used to accept a solution, such as risk tolerance or choosing between optimizing versus satisficing).
Choice: Selecting and committing to a specific course of action. Involves searching for, evaluating, and recommending a solution. The boundary between Design and Choice is iterative, as evaluating alternatives may trigger a return to the Design phase to generate more options.
Implementation: Putting the selected solution into effect. Involves managing organizational change, securing management support, conducting training, and monitoring outcomes. Monitoring feedback flows back into the Intelligence phase to evaluate performance and inform future decisions.
DSS Support Across Simon's Phases
Intelligence Phase: Supported by environmental scanning tools, web search engines, Business Activity Monitoring (BAM), and Business Performance Management (BPM).
Design Phase: Supported by financial and forecasting models, OLAP, and data mining to discover relationships and generate scenarios.
Choice Phase: Supported by what-if analysis, goal-seeking analysis, decision trees, and expert systems.
Implementation Phase: Supported by Group Decision Support Systems (GDSS/GSS), communication platforms, and project tracking tools.
How Businesspeople Make Decisions
Decision-making is rarely a single linear pass through Simon's four phases. Managers frequently loop back between phases as new information emerges.
Two primary philosophies of choice guide decision-making:
Normative / Rational Approach: Assumes the decision-maker is an economically rational optimizer who knows all possible alternatives and their exact consequences, selecting the objectively optimal alternative using optimization techniques.
Descriptive Approach: Recognizes bounded rationality. Managers rely on descriptive models, simulations, and scenario analyses to understand consequences, then apply judgment and experience to pick a solution. Decisions often rely on satisficing — choosing a "good enough" option that meets minimal criteria rather than searching endlessly for a single optimal choice.
DSS are valuable because they augment this bounded, judgment-driven process with richer data and a broader array of evaluated alternatives without forcing decisions into rigid, fully automated formulas.
DSS Classification / Hierarchy
The AIS SIGDSS classification framework established by D. J. Power categorises DSS by their dominant driver or structural component:


Communications-driven & Group DSS (GSS): Driven by communication and collaboration technology. Supports groups working together on tasks, meetings, design collaboration, groupware, chat, and video.
Data-driven DSS: Driven by databases or data warehouses. Focuses on accessing, querying, and manipulating large volumes of structured historical data. Includes Executive Information Systems (EIS) and Spatial DSS, featuring minimal quantitative modelling.
Document-driven DSS: Driven by document repositories. Focuses on searching, retrieving, and analysing unstructured documents (policies, records, web pages, media). Overlaps with knowledge management.
Knowledge-driven DSS: Driven by knowledge bases and inference engines. Employs AI-based systems (expert systems, data mining, artificial neural networks) to suggest or recommend actions using domain expertise.
Model-driven DSS: Driven by quantitative and optimization models. Includes financial, accounting, simulation, or optimization tools (e.g., Excel Solver) to perform what-if and scenario analyses.
Secondary Classifications
Compound / Hybrid DSS: Combines two or more dominant categories (e.g., combining model-driven analytics with data-driven data warehousing).
Institutional DSS: Handles recurring organizational decisions and is refined over years (e.g., automated portfolio management systems).
Ad Hoc DSS: Built to address one-off, unanticipated, or unique strategic problems.
Scope / User Base: Intra-organisational DSS (internal staff) vs. Inter-organisational DSS (extending to customers and suppliers).
Enabling Technology: Stand-alone PC applications, client/server systems, mainframe applications, or web-based applications.
Matching DSS to Decision Type and Decision-Maker
A DSS category must match both the structuredness of the problem and the user's primary need:
Routine, recurring, highly structured decisions (e.g., airline crew scheduling) suit model-driven or institutional DSS.
One-off strategic problems suit ad hoc, document-driven, or knowledge-driven DSS combined with managerial judgement.
Multi-departmental decisions suit communications-driven / group DSS.
Cognitive Style Alignment
Design must account for the user's psychological and cognitive style:
Analytical decision-makers prefer data-driven DSS with detailed tabular outputs, raw data access, and drill-down capabilities.
Heuristic / Intuitive decision-makers prefer visual interfaces, model-driven scenario comparisons, charts, and graphical summaries.
Individual vs Group DSS
Individual DSS: Designed to support a single user (or an analyst working for a executive) working independently on a standalone workstation.
Group DSS (GDSS / GSS): Designed for teams working collaboratively. Adds communication, scheduling, document sharing, brainstorming, voting, and consensus-building capabilities to traditional analytical software.
Synchronous GSS: Same-time, same-place or same-time, different-place decision meetings.
Asynchronous GSS: Different-time, different-place virtual team collaboration over web platforms.
DSS Components and Architecture
A traditional DSS comprises four primary structural subsystems:
Data Management Subsystem: Contains the database (often a data warehouse or data mart) holding internal, external, current, and historical structured/unstructured data.
Model Management Subsystem: Contains quantitative, statistical, financial, or optimization models, along with the Model Base Management System (MBMS) software used to create, update, and manipulate models.
User Interface Subsystem: Covers the interaction mechanism (dialogs, graphics, menus, dashboards, web interfaces). It is critical because the interface defines the usable system from the manager's perspective.
Knowledge-Based Management Subsystem / Architecture: Encompasses system hardware, network topology, and intelligence components (rules, inference engines) supporting the other subsystems.
Physical architecture options include model servers, data servers, network backbones, and client devices running either thick-client (desktop software) or thin-client (web browser) interfaces. Key engineering criteria include scalability, security, and network availability.
DSS vs Business Intelligence

Data Source: DSS utilizes any internal or external data source; BI relies primarily on central enterprise data warehouses.
Orientation: DSS directly supports specific decision-making tasks and is analyst-oriented; BI provides integrated performance reporting and visualization intended to guide strategic and executive decisions.
Origin: DSS originated largely from academic research in management science and Information Systems (IS); BI emerged primarily from software industry vendors.
Tooling: DSS uses custom-built tools tailored to unstructured or semi-structured problems; BI relies on commercially available packaged software suites (e.g., reporting platforms, dashboard tools).
The Impact of Culture on Decision-Making
National and organizational culture strongly influences how decisions are made:
Hierarchical, Risk-Averse Cultures: Prefer top-down decision rights, formal normative/rational models, extensive analysis, and consensus before taking action. Rely more heavily on model-driven analysis presented by specialists to senior leaders.
Individualistic, Risk-Tolerant Cultures: Decentralise decision-making, operate faster, and accept satisficing strategies by empowered individual managers. Adopt communications-driven DSS to facilitate rapid peer collaboration.
Comparison: DSS Categories

Category | Dominant Driver / Component | Focus / Typical Use | Example |
|---|---|---|---|
Communications-driven & Group DSS (GSS) | Communication/collaboration technology | Coordinates meetings, group problem-solving, joint decision-making | Video conferencing + shared whiteboard for a planning meeting |
Data-driven DSS | Database / data warehouse | Accesses, queries, and manipulates large volumes of structured historical data; minimal modeling | OLAP sales-reporting dashboard |
Document-driven DSS | Document base / repository | Searches and retrieves unstructured documents (policies, records, web pages, media) | Knowledge base of policies and past case files |
Knowledge-driven DSS | Knowledge base / inference engine | Recommends actions using domain expertise and AI rules | Expert system for medical diagnosis or credit approval |
Model-driven DSS | Quantitative / optimization models | What-if analysis, financial planning, simulation, optimization | Financial planning spreadsheet with Solver |
Compound / Hybrid DSS | Two or more dominant components | Complex real-world applications that span multiple functional needs | ERP-integrated planning DSS combining data warehouses and optimization models |
Key Terms — Week 1
Decision Support System (DSS): Interactive, computer-based system designed to help decision-makers use data, documents, knowledge, and models to solve semi-structured or unstructured problems.
Structured Decision: Routine, repetitive decision with standard, known solution procedures.
Unstructured Decision: Complex, fuzzy decision with no predefined solution procedure, relying heavily on human judgement.
Semi-structured Decision: Decision containing both programmatic/standardizable elements and unstructured elements requiring human judgement.
Gorry & Scott-Morton Framework: A matrix categorising decisions by structuredness (structured, semi-structured, unstructured) and managerial activity (operational control, management control, strategic planning).
Satisficing: Selecting a option that satisfies minimum acceptable criteria rather than searching endlessly for the optimal solution.
Normative Model: Model that prescribes the mathematically optimal solution assuming full rationality.
Descriptive Model: Model that describes or simulates how a system behaves without asserting an optimal choice.
Group DSS (GSS): Collaborative system designed to support multiple decision-makers working together synchronously or asynchronously.
Institutional DSS: System supporting recurring organizational decisions, maintained and refined over long periods.
Ad Hoc DSS: System built quickly to solve a specific, one-off, unexpected decision problem.
Self-Test Questions — Week 1
Using the Gorry & Scott-Morton framework, explain the difference between structured, semi-structured, and unstructured decisions, and give a business example of each.
List and briefly describe Simon's four phases of the decision-making process, and explain how DSS technologies support each phase.
A logistics company needs a system to let regional managers jointly plan quarterly delivery routes over video calls with shared documents. Which DSS category best fits this need, and why?
Explain three key differences between a DSS and a Business Intelligence system.
Why should DSS design take into account a decision-maker's psychological/cognitive style, and give an example of matching interface style to decision type?
First Block — Week 2: Data Warehousing, OLAP & Introduction to Data Mining
What Is a Data Warehouse and Why It Exists
A Data Warehouse is a subject-oriented, integrated, time-variant, and non-volatile collection of data designed to support managerial decision-making:
Subject-oriented: Organised around major business subjects (e.g., sales, customers, products) rather than around functional operational applications.
Integrated: Reconciles heterogeneous data from multiple disparate operational systems into a unified format with standardized identifiers, codes, and naming conventions.
Time-variant: Retains historical data over long horizons (e.g., 5–10 years), preserving past states to evaluate trends, whereas operational databases typically store only current operational states.
Non-volatile: Data is read-only once loaded. Operational updates, overwrites, and deletions do not occur in the warehouse; data is accessed purely through analytical queries.
Operational Transaction Processing (OLTP) systems are optimized for fast, transactional read/write operations and current data maintenance. Data warehouses provide integrated, cross-functional historical data for complex analytical querying, reporting, forecasting, and data mining without impacting operational performance.
Cloud Data Warehouses
Modern enterprise analytics has shifted toward cloud data warehouses. Unlike traditional on-premises warehouses — which suffer from rigid hardware limits, complex capacity planning, and batch update windows — cloud warehouses offer elastic scalability, separation of storage and compute, massively parallel processing (MPP), pay-per-use cost structures, and direct integration with modern data lakes and machine learning platforms.
Data Warehouse Architecture and Design
Common architectural patterns include:
Simple Data Warehouse: Single central repository feeding reports directly from raw and summarized data stores.
Simple with Staging Area: Includes an intermediate staging area where source data is extracted, cleaned, transformed, and normalized before being loaded into the central warehouse.
Hub and Spoke Architecture: A central enterprise data warehouse feeds smaller, departmental data marts tailored to specific business units (e.g., marketing, finance).
Sandbox Architecture: Private, isolated analytical environments where data scientists explore raw data without violating formal warehouse governance.
Development Approaches
Top-Down (Inmon Methodology): Starts by designing a comprehensive enterprise-wide data warehouse. Departmental data marts are created afterwards from the central store. High consistency and scalability, but requires high initial investment and long build times.
Bottom-Up (Kimball Methodology): Starts by building individual, business-unit-specific data marts based on dimensional models. Data marts are joined together over time via standardized "conformed dimensions". Delivers faster ROI and lower initial costs.
Hybrid Approach: Combines enterprise-level data planning (top-down governance) with incremental departmental implementations (bottom-up deployment).
Dimensional Modelling Steps
Select the Business Process: Define the business domain to model (e.g., retail point-of-sale, insurance claims).
Declare the Grain: Define the precise atomic level of detail for a single fact table record (e.g., one row per line item on a customer receipt).
Identify the Dimensions: Determine the contextual attributes describing each event (e.g., date/time, store location, customer, product).
Identify the Measures: Choose the numeric, quantitative facts to be aggregated and analyzed (e.g., quantity sold, total revenue, tax amount, net profit).
Dimensional schemas structure data into a central Fact Table containing numeric measures surrounded by Dimension Tables containing descriptive context (forming a Star Schema or Snowflake Schema).
How Data Gets Into a Warehouse
Data enters a warehouse via pipelines:
ETL (Extract, Transform, Load): Data is extracted from operational databases, transformed in a staging environment (cleansing, deduplication, key mapping, aggregation), and loaded into the target warehouse. Best for structured, highly governed legacy architectures.
ELT (Extract, Load, Transform): Raw data is extracted and loaded directly into high-performance cloud data warehouses or data lakes first, then transformed on demand using the target platform's computational power. Best for massive, unstructured, or rapidly changing data flows.
OLAP (Online Analytical Processing)
OLAP provides multidimensional analytical capabilities, allowing managers to query and explore warehouse data interactively across multiple dimensions:
Slicing: Extracting a single two-dimensional subset from a multidimensional data cube by fixing one dimension (e.g., filtering sales data for only the year 2023).
Dicing: Extracting a smaller sub-cube by applying selection criteria across multiple dimensions simultaneously (e.g., sales in Region A for Product Category B during Q2).
Drill-Down: Navigating from summarized, high-level data to detailed lower-level data (e.g., moving from annual sales to monthly sales, down to daily store transactions).
Roll-Up: Aggregating low-level detailed data up to higher-level concepts along a dimensional hierarchy (e.g., rolling up individual store locations into state totals).
Pivoting (Rotation): Rotating the data axes in a visual display to examine data from alternative perspectives.
Data Warehouse vs DBMS vs Data Lake


Aspect | Transactional DBMS | Data Warehouse | Data Lake |
|---|---|---|---|
Purpose | Day-to-day operations (OLTP) | Analysis & decision support (OLAP) | Flexible storage of raw, uncurated data |
Data | Current, application-specific | Historical, integrated, enterprise-wide | Any format (structured, semi-structured, unstructured) |
Schema | Fixed (Entity-Relationship model) | Structured, predefined (e.g., dimensional/star) | No predefined schema (Schema-on-read) |
Key Guarantee | ACID transaction properties | Consistency & non-volatility for reporting | None inherent — governance must be applied manually |
Cost | Lower baseline cost | Higher (structural setup, ETL pipelines, compute engines) | Low (cheap, scalable raw blob storage) |
Typical Users | Clerical staff, operational workers, applications | Business analysts, managers, executive BI platforms | Data scientists, machine learning engineers |
Risk | N/A (well-governed by design) | Expensive to build and maintain | Can degrade into an unmanageable "data swamp" without strict governance |
A Data Lakehouse represents a unified hybrid architecture that implements warehouse structure, ACID transactions, and governance directly over cheap data lake storage formats.
Data Warehouse Advantages and Disadvantages
Advantages
Enhanced decision quality through consistent, validated data.
Comprehensive cross-functional insights by unifying siloed source systems.
High-performance historical trend analysis and long-range forecasting.
Offloading analytical workloads from operational OLTP databases.
Centralised security, access control, and data governance.
Disadvantages
High capital expenditure and ongoing maintenance costs.
Significant technical complexity and long development cycles.
Rigidity of predefined schemas when business requirements change rapidly.
Potential security risk if sensitive organizational data is consolidated into a single target.
What Is Data Mining and Why It Matters
Data mining is the process of discovering interesting, non-trivial, implicit, previously unknown, and potentially useful patterns, rules, anomalies, and knowledge from large collections of data.
Although popularly termed "data mining", it is more accurately viewed as the core algorithmic step within the broader Knowledge Discovery in Databases (KDD) process.
Four defining properties of data mining:
Automated or semi-automated discovery of implicit patterns.
Prediction of likely future outcomes based on historical trends.
Generation of actionable insights for decision-making.
Operation on large-scale datasets and databases.
Data mining integrates techniques from statistics, machine learning, pattern recognition, database management, information retrieval, data visualization, and high-performance computing. Applications span financial fraud detection, credit scoring, retail market basket analysis, telecommunications churn prediction, healthcare diagnostics, and bioinformatics.
The Data Mining Process (Pipeline)
The formal Knowledge Discovery in Databases (KDD) pipeline comprises seven iterative steps:
Data Cleaning: Removing noise, erroneous data, and inconsistent records.
Data Integration: Combining multiple heterogeneous data sources into a unified structure.
Data Selection: Retrieving data relevant to the specific analytical objective.
Data Transformation: Consolidating data into forms appropriate for mining (e.g., aggregation, normalization, binning).
Data Mining: Applying computational algorithms to extract data patterns.
Pattern Evaluation: Identifying truly interesting patterns using statistical measures of interestingness.
Knowledge Presentation: Visualizing and documenting extracted patterns for decision-makers.
At a higher level, the process condenses into three main phases: Data Collection, Feature Extraction & Preprocessing, and Algorithmic Execution & Evaluation. Real-world applications require spending up to 80% of total project effort on preprocessing steps.
Types of Data Suited to Mining
Multidimensional / Record Data: Standard tabular structures containing fixed numeric and categorical attributes (e.g., database records, matrix structures).
Text and Web Data: Unstructured sparse documents represented as bags-of-words or TF-IDF vectors.
Time-Series and Sequential Data: Sequentially ordered continuous numerical measurements or event streams (e.g., stock ticker streams, sensor logs, DNA sequences).
Spatial and Spatiotemporal Data: Attributes bound to physical geographic coordinates and temporal paths (e.g., satellite imagery, weather patterns, GPS tracking).
Graph and Network Data: Complex structures of nodes and edges representing interconnected entities (e.g., social networks, biological networks, web page hyper-links).
The Four Major Data Mining Tasks (Introductory Overview)
Data mining problems decompose into four fundamental superproblems:
Association Pattern Mining: Unsupervised discovery of item sets or attributes that co-occur frequently within transactions (e.g., Market Basket Analysis).
Clustering: Unsupervised grouping of unlabelled records into clusters based on internal pairwise similarity metrics.
Classification: Supervised learning task that maps input feature vectors to discrete target class labels using models trained on labelled training data.
Outlier (Anomaly) Detection: Identifying records that deviate significantly from standard statistical distributions or normal cluster structures, indicating fraud, sensor failures, or intrusions.
Key Terms — Week 2
Data Warehouse: Subject-oriented, integrated, time-variant, non-volatile data repository designed for decision support.
ETL (Extract, Transform, Load): Data pipeline process where operational data is extracted, transformed in a staging area, and loaded into a warehouse.
ELT (Extract, Load, Transform): Modern pipeline approach where raw data is loaded directly into the target database engine prior to transformation.
OLAP (Online Analytical Processing): Interactive, multidimensional database exploration technology allowing slicing, dicing, drilling, and rolling up.
Data Mart: Sub-setting mechanism providing a departmental or subject-specific slice of an enterprise data warehouse.
Data Lake: Scalable repository storing raw, unstructured, semi-structured, and structured enterprise data without upfront schema definitions.
ACID Properties: Atomicity, Consistency, Isolation, Durability — strict transactional guarantees enforced by operational DBMS engines.
Data Mining: Computational process of extracting valid, novel, useful, and understandable patterns from large datasets.
KDD (Knowledge Discovery in Databases): The overarching multi-step process encompassing data cleaning, integration, selection, transformation, mining, evaluation, and presentation.
Fact Table: Central table in a dimensional schema containing quantitative numerical measures and foreign keys pointing to dimension tables.
Dimension Table: Companion tables in dimensional models holding descriptive, contextual attributes used to filter and group numeric facts.
Self-Test Questions — Week 2
Explain the four defining properties of a data warehouse (subject-oriented, integrated, time-variant, non-volatile) using a concrete business example for each.
Compare a data warehouse, a transactional DBMS, and a data lake in terms of purpose, schema, and typical user.
Describe, in plain language, what happens during the ETL process and why raw source data usually cannot be loaded directly into a warehouse.
List and briefly describe the seven steps of the knowledge discovery process, distinguishing "data mining" narrowly from the broader knowledge-discovery pipeline.
A retailer wants to know (a) which products are usually bought together, and (b) which customers are unusually different from the rest in their spending habits. Which data mining tasks map to each need, and why?
First Block — Week 3: Introduction to Python, Exploratory Data Analysis & Data Visualisation
Why Python for Data Mining
Python has become the dominant language for modern data science and data mining due to several features:
High-level, readable, clean syntax that maximizes developer productivity ("less code, more output").
An extensive ecosystem of open-source scientific and data analytics libraries.
Interpreted, interactive execution (e.g., through Jupyter Notebooks) enabling rapid experimentation and iterative analysis.
Native capabilities as a glue language, connecting database APIs, cloud infrastructure, C/C++ libraries, and web deployments.
Support from distributions such as Anaconda, which package the core data stack with environment management software.
The Core Python Data-Science Ecosystem
NumPy: Provides the baseline
ndarray(n-dimensional array) object, vectorised arithmetic operations, linear algebra functions, Fourier transforms, and random number generation capabilities.pandas: Built on top of NumPy, providing high-performance data structures: the single-column
Seriesand the two-dimensional tabularDataFrame. Used for loading, cleaning, merging, filtering, pivoting, and reshaping data.Matplotlib: Core plotting library providing complete control over 2D visual charts, axes, and figure formatting.
Seaborn: High-level visualization library built on top of Matplotlib. Integrated with pandas DataFrames to generate statistical plots (e.g., heatmaps, pair plots, violin plots).
scikit-learn: Primary machine learning library providing implementations of algorithms for classification, regression, clustering, dimensionality reduction, feature selection, and model evaluation.
Jupyter Notebooks: Interactive browser-based coding environment integrating live code execution, formatted text, LaTeX mathematics, and inline visual plots.
Getting to Know Your Data: Exploratory Data Analysis (EDA)
Exploratory Data Analysis (EDA) is the practice of systematically inspecting, summarizing, and visualizing a dataset's structure, distributions, anomalies, and relationships before formal statistical modeling or machine learning is performed.
EDA ensures analysts do not impose invalid assumptions onto the data. Objectives of EDA include:
Detecting data quality flaws (missing records, entry errors, extreme outliers, duplicated rows).
Identifying underlying spatial, temporal, or distributional shapes.
Assessing feature correlations to guide feature selection and engineering.
Informing algorithm selection based on distributional properties (e.g., checking for skewness or non-linearity).
The Standard EDA Process
Data Collection & Extraction: Ingesting raw data from files, databases, or APIs.
Data Inspection & Cleaning: Assessing data types, null counts, and structural integrity.
Visualization: Plotting variables univariate and bivariate to observe shapes and relationships.
Summary Statistics Computation: Calculating measures of central tendency, dispersion, and association.
Interpretation & Next Steps: Formulating hypotheses, identifying required data transformations, and selecting downstream algorithms.
Descriptive Statistics: Central Tendency
Measures of central tendency locate the center of a distribution using a single scalar:
Mean: The arithmetic average of all values in a numerical vector:
It utilizes all data points, but is sensitive to extreme outliers.
Median: The middle value when values are ordered sequentially. If the sample size is even, it is the average of the two central values. It is robust against extreme values and skewed distributions.
Mode: The most frequently occurring value in a dataset. It is the only measure of central tendency applicable to nominal categorical data.
Comparing the mean and median serves as a key diagnostic indicator for distribution skewness:
Mean Median indicates positive (right) skewness.
Mean Median indicates negative (left) skewness.
Mean Median indicates an approximately symmetric distribution.
Descriptive Statistics: Dispersion (Spread)
Dispersion measures quantify how widely spread data points are around their central value:
Range: Difference between the absolute maximum and minimum values: . Sensitive to outliers.
Variance (): Average of squared deviations from the sample mean:
Standard Deviation (): Square root of the variance, expressing spread in the original measurement units:
Percentiles and Quartiles: Values dividing ordered data into 100 equal parts (percentiles) or 4 equal parts (quartiles: , , percentile).
Interquartile Range (IQR): Distance between the third and first quartiles: . Robust measure of spread representing the middle 50% of data.
Five-Number Summary: Set containing Minimum, , Median (), , and Maximum. Forms the structural basis of the box plot.
Skewness: Measure of distribution asymmetry. Positive skew exhibits a elongated right tail; negative skew exhibits an elongated left tail.
Data Distributions
Normal (Gaussian) Distribution: Symmetric, bell-shaped distribution where mean, median, and mode coincide. Defined by mean and variance . Parametric statistical algorithms often assume approximate normality.
Skewed Distributions: Asymmetric real-world data distributions (e.g., household income, insurance claims, internet traffic). Requires non-parametric statistics or logarithmic data transformations prior to quantitative modelling.
Data Visualisation for Understanding Data
Visualization translates numerical tables into visual structures, exposing patterns, trends, and anomalies.
Comparison: Common Chart Types and When to Use Them

Chart Type | Best For | Example Use Case |
|---|---|---|
Histogram | Displaying the continuous distribution/shape of a single numeric variable (spread, skew, modality) | Checking whether "customer age" is normally distributed or skewed |
Box plot (box-and-whisker) | Summarising the five-number summary and spotting outliers at a glance; comparing spread across groups | Comparing salary spread across departments and flagging outlier salaries |
Scatter plot | Displaying the relationship/correlation between two continuous numeric variables | Checking whether advertising spend relates to sales revenue |
Bar chart | Comparing counts/values across discrete categories | Comparing number of purchases per product category |
Line chart | Displaying how a variable changes over an ordered sequence, typically time | Tracking monthly website traffic over a year |
Pie chart | Showing proportional share of a whole across a small number of categories | Displaying market share split between a few competitors |
Heatmap | Visualising a matrix of values (e.g., correlation matrices) using colour intensity | Displaying a correlation matrix between many numerical features |
Data Pre-processing and Preparation (Conceptual Overview)
Data preparation comprises four core operations:
Data Cleansing: Imputing or removing missing values, resolving inconsistencies, and removing duplicate entries.
Outlier Treatment: Identifying and evaluating extreme observations to determine whether to retain, transform, or drop them.
Scaling & Normalization: Standardizing attribute scales to prevent variables with large magnitudes from dominating distance computations.
Encoding: Converting categorical variables into numerical representations (e.g., one-hot encoding or ordinal encoding).
Key Terms — Week 3
Exploratory Data Analysis (EDA): Investigative process computing statistical summaries and visualizations to understand dataset characteristics.
Mean: Arithmetic average sensitive to extreme values.
Median: Middle value of a sorted dataset; robust to outliers.
Mode: Most frequent observation; applicable to nominal categorical data.
Standard Deviation: Square root of variance; measures dispersion in base measurement units.
Interquartile Range (IQR): Distance between the and percentiles ().
Five-Number Summary: Dataset descriptor comprising Minimum, , Median, , and Maximum.
Skewness: Degree of asymmetry of a distribution around its mean.
Outlier: Data point departing significantly from the general distribution of a dataset.
Imputation: Process of replacing missing data points with substituted statistical values.
NumPy: Python library supporting array processing and vectorised mathematical calculations.
pandas: Python library providing core
DataFrameandSeriesanalytical data structures.
Self-Test Questions — Week 3
Why is Python considered well-suited to data mining despite not being the fastest programming language?
What conceptual role does pandas play compared to NumPy in a typical data analysis workflow?
Why should EDA be performed before building a data mining model?
If a dataset's mean is much higher than its median, what does that suggest about the distribution's shape?
Why is standard deviation generally preferred over variance when communicating spread to a non-technical audience?
Why is the interquartile range considered more robust than the range as a measure of spread?
When would you choose a box plot over a histogram to explore a numeric variable, and why?
Why should an "outlier" not automatically be deleted from a dataset before it is investigated?
Which chart type would you use to explore the relationship between two continuous variables, and why?
Why is normalisation/standardisation of attributes often necessary before analysis, using age and salary as an example?
First Block — Week 4: Data Integration and Transformation
Data Preparation in Context
Data preparation encompasses feature extraction, data cleansing, data integration, data reduction, and data transformation. Raw operational data is messy, incomplete, inconsistent, and noisy. Preparing data ensures down-stream statistical models receive standardized inputs, directly affecting model accuracy.
Data Integration
Data Integration combines data from multiple heterogeneous operational sources into a unified, coherent data repository.
Formally, integration is represented as a triple: Global Schema (unified view), Source Schemas (heterogeneous source structures), and Mapping Rules connecting sources to the global view.
Architectural Integration Paradigms
ETL (Extract, Transform, Load): Batch process where source data is extracted, cleansed and restructured in a dedicated staging area, then loaded into a target data warehouse.
ELT (Extract, Load, Transform): Raw source data is extracted and loaded directly into a target processing engine (e.g., cloud data lake), using the target compute platform to execute transformations.
Change Data Capture (CDC): Technique that monitors database transaction logs to capture and stream incremental data state changes in real time.
Streaming Integration: Continuous real-time ingestion of event streams (e.g., using Apache Kafka).
Application Integration (API-driven): Synchronizing operational systems directly via API interfaces.
Data Virtualization: Creates an abstract real-time virtual schema over disparate sources without physically moving data into a central storage repository.
Key Integration Challenges
Schema Integration & Entity Matching: Resolving conflicts when attributes representing the same real-world entity have different names, data types, or structures across systems (e.g., matching
Cust_NotoClient_ID).Redundancy & Correlation: Attributes may be redundant if they can be derived from other attributes (e.g.,
Annual_Salaryderived fromMonthly_Salary). Redundancy causes storage bloat and distorts algorithms sensitive to multicollinearity.
Data Cleaning
Data cleaning resolves missing values, noisy/incorrect entries, and structural errors:
Missing Data Handling:
Deletion (Listwise/Pairwise): Dropping records containing null values. Only appropriate if missingness is minimal and random.
Imputation: Filling missing entries using statistical measures (mean, median, mode) or predictive algorithms (k-NN, regression). For sequential time-series data, linear interpolation or forward-filling is applied.
Model-Based Mitigation: Utilizing algorithms natively robust to missing data (e.g., XGBoost, Decision Trees).
Erroneous / Inconsistent Data Handling:
Cross-source validation checks (e.g., checking logic rules: a record cannot have
City = LondonandCountry = Japan).Statistical outlier detection to identify anomalies.
Manual inspection of flagged records. Outliers must be inspected before removal, as anomalies may represent critical events like financial fraud.
Scaling & Normalization: Standardizing attribute magnitudes across features.
Data Reduction
Data reduction produces a compact representation of the dataset that retains analytical integrity while accelerating processing efficiency.
Data Cube Aggregation: Summarizing facts across higher hierarchy dimensions (e.g., aggregating daily transactions to monthly totals).
Attribute Subset Selection: Removing irrelevant, weakly relevant, or redundant features using feature selection algorithms.
Dimensionality Reduction: Transforming original high-dimensional feature spaces into lower-dimensional representations via linear transformations (e.g., PCA, SVD).
Numerosity Reduction: Replacing raw data with smaller parametric representations (e.g., regression parameters) or non-parametric representations (e.g., histograms, clusters, sampling).
Discretisation & Concept Hierarchies: Grouping continuous numeric attributes into categorical intervals.
Sampling Strategies:
Simple Random Sampling Without Replacement (SRSWOR): Each record has an equal probability of selection without re-insertion.
Simple Random Sampling With Replacement (SRSWR): Selected records are returned to the pool, allowing multiple selections.
Stratified Sampling: Partitioning the dataset into mutually exclusive sub-groups (strata) based on a target attribute, then sampling proportionally from each stratum. Ensures rare sub-populations (e.g., minority class targets) are adequately represented.
Data Transformation
Data transformation rescales, reshapes, and consolidates data attributes into formats optimal for mining:
Smoothing: Removing noise via binning, regression, or moving averages.
Aggregation: Computing summary metrics across dataset slices.
Generalisation: Replacing granular values with higher-level categorical concepts (e.g., replacing exact birth dates with age groups).
Feature Construction: Generating derived features from existing variables to capture domain knowledge.
Scaling Operations: Normalisation vs Standardisation
Standardisation (-score scaling): Rescales data to have a mean of and a standard deviation of :
Standardisation is robust to extreme outliers because it scales relative to the standard deviation rather than fixed min/max boundaries.
Min-Max Normalisation: Rescales raw data into a fixed interval, typically :
Min-max scaling is sensitive to extreme outliers: if an erroneous outlier contains an extreme maximum value, the main body of valid data is compressed into a tight sub-interval near 0$.\n\n- Discretisation (Binning) Methods:\n - Equi-Width Binning: Divides the range of an attribute into k equal-sized intervals. Sensitive to skewed distributions.\n - Equi-Depth (Equi-Frequency) Binning: Divides data so each of the k intervals contains approximately the same number of samples. Adapts to data density.\n - Equi-Log Binning: Interval widths increase exponentially; useful for power-law distributions.\n\n## Principal Component Analysis (PCA) — Conceptual Explanation\n\nAs feature count increases, the volume of feature space grows exponentially, rendering data sparse — a phenomenon known as the Curse of Dimensionality. High dimensionality increases computational complexity, distorts Euclidean distance metrics, and leads to overfitting.\n\nPrincipal Component Analysis (PCA) is an unsupervised linear dimensionality reduction technique that transforms a set of correlated numerical features into a smaller set of uncorrelated features called Principal Components, while retaining maximum variance.\n\n### Conceptual Execution of PCA\n\n1. Standardise all numeric features so they share mean 01.\n2. Construct the Covariance Matrix of the features to quantify pairwise linear correlations.\n3. Compute Eigenvalues and Eigenvectors of the covariance matrix.\n - Eigenvectors define the spatial directions of the new orthogonal principal axes.\n - Eigenvalues quantify the amount of variance captured along each corresponding eigenvector axis.\n4. Sort eigenvectors by their eigenvalues in descending order.\n - The First Principal Component (PC_1) aligns with the direction of maximum variance.\n - The Second Principal Component (PC_2PC_1 and aligns with the second highest variance direction.\n5. Select the top kk eigenvectors to yield a reduced feature space.\n\nPCA is an unsupervised technique because it operates solely on feature covariance, ignoring target class labels.\n\n## Comparison: Data Transformation Techniques\n\n\n\n| Technique | What It Does | When to Use |\n| :--- | :--- | :--- |\n| Standardisation (z0 | When attributes have different units/scales and data contains outliers (e.g., prior to PCA or distance-based ML) |\n| Min-max normalisation | Rescales values into a fixed range (e.g., 01) using min/max values | When the value range is bounded and clean (few/no extreme outliers) |\n| Discretisation / binning | Converts continuous values into a smaller number of categorical ranges | When interpretable categorical inputs are required, or for algorithms requiring categorical data |\n| Aggregation | Summarises detailed data into higher-level totals or averages | When analysis is needed at a coarser granularity (e.g., monthly instead of daily) |\n| Generalisation | Replaces raw categorical values with higher-level concepts | When a concept hierarchy exists and macro-level patterns are more meaningful than micro detail |\n| PCA (dimensionality reduction) | Projects correlated numeric attributes onto a smaller set of uncorrelated components capturing maximum variance | When numerical features exhibit high multicollinearity and dimensionality must be reduced while preserving information |\n\n## Key Terms — Week 4\n\n- Schema Integration: Unifying disparate database attribute structures into a single global schema.\n- Entity Matching: Identifying records across multiple databases that refer to the exact same physical entity.\n- Data Reduction: Producing a compact dataset representation that yields near-identical analytical results.\n- Stratified Sampling: Sampling method ensuring proportional representation across mutually exclusive subgroups.\n- Standardisation (z01.\n- Min-Max Normalisation: Scaling transformation bounding data values strictly within a target interval (e.g., [0, 1]).\n- Discretisation: Converting continuous attributes into discrete interval bins.\n- Curse of Dimensionality: Exponential growth in data volume and computational complexity required as feature counts increase.\n- Principal Component Analysis (PCA): Unsupervised linear transformation projecting data onto orthogonal components that maximize variance.\n- Explained Variance: The proportion of a dataset's total variance accounted for by individual principal components.\n\n## Self-Test Questions — Week 4\n\n1. What are the two main challenges of data integration, and give an example of each?\n2. Explain the difference between an ETL and an ELT approach to data integration, and when each is preferable.\n3. Why is an outlier not automatically treated as an error during data cleaning?\n4. Name and briefly describe the three general approaches to handling missing data.\n5. Why might min-max normalisation give misleading results if the dataset contains an extreme, erroneous outlier value, and why is standardisation more robust in that scenario?\n6. What is the difference between equi-width and equi-depth discretisation, and why might equi-depth be preferred for a very unevenly distributed attribute like salary?\n7. What is stratified sampling, and why is it useful when a subgroup of interest is rare in the population?\n8. What problem does the "curse of dimensionality" describe, and how does PCA help address it?\n9. In your own words, what does the "first principal component" of a dataset represent?\n10. Why is data usually standardised before PCA is applied?\n11. Is PCA a supervised or unsupervised technique, and why does that classification matter?\n12. How is "explained variance" used to decide how many principal components to keep?\n\n# First Block — Week 5: Relational Databases, NoSQL Databases and Python Access\n\n## What is a Relational Database?\n\nA Relational Database Management System (RDBMS) organizes structured data into formal two-dimensional tables (relations) comprising rows (tuples) and columns (attributes).\n\n- Row (Record / Tuple): Represents a single, discrete data instance (e.g., a specific customer).\n- Column (Field / Attribute): Represents a specific property possessed by all records in that table (e.g., `Email_Address`).\n- Primary Key: An attribute (or composite set of attributes) that uniquely identifies every distinct row in a table. It cannot contain null values.\n- Foreign Key: An attribute in one table that references the primary key of another table, establishing a logical relational link between rows across tables.\n\nRelational databases enforce structural integrity and eliminate data redundancy by splitting data across normalized tables.\n\nSQL (Structured Query Language) is the standard declarative programming language used to define, query, update, and manage relational database systems (e.g., PostgreSQL, MySQL, Oracle, SQL Server).\n\n## Database and Data Warehouse Design\n\n- Entity-Relationship (ER) Modelling: Conceptual design methodology identifying real-world entities, their attributes, and structural relationships (1-to-1, 1-to-Many, Many-to-Many) prior to logical database construction.\n- Normalisation: Process of organizing database tables to minimize structural data redundancy and eliminate update, insertion, and deletion anomalies.\n\n### Data Warehouse Design Review\n\nData Warehouses adhere to four structural properties: Subject-Oriented, Integrated, Time-Variant, and Non-Volatile. Dimensional modeling follows four steps: (1) Business Process Selection, (2) Grain Declaration, (3) Dimension Identification, (4) Measure Identification.\n\n## OLAP: Online Analytical Processing\n\n- OLTP vs OLAP: Operational Transaction Processing (OLTP) manages high-frequency, short-lived, current-state transactions with ACID guarantees. Online Analytical Processing (OLAP) executes complex, read-heavy analytical queries over historical data to support strategic decision-making.\n\n### Core OLAP Operations\n\n- Roll-Up: Summarizing data up a dimension hierarchy.\n- Drill-Down: Uncovering detailed low-level figures.\n- Slicing: Extracting a single 2D plane.\n- Dicing: Selecting a sub-cube across multiple dimensions.\n\n### Architectural Implementations of OLAP\n\n- ROLAP (Relational OLAP): Queries multidimensional data directly stored in relational database tables. Highly flexible and capable of handling massive volumes, but query response can be slower as aggregations are calculated on the fly.\n- MOLAP (Multidimensional OLAP): Stores data in specialized, pre-calculated array structures called Data Cubes. Delivers fast query performance, but requires lengthy pre-computation build windows and is less flexible to dynamic schema changes.\n- HOLAP (Hybrid OLAP): Combines ROLAP and MOLAP by retaining low-level transactional details in relational tables while storing pre-aggregated summaries in multidimensional cubes.\n\n## NoSQL Databases\n\nNoSQL (Not Only SQL) databases emerged to address the scaling limitations of relational systems when processing high-volume, high-velocity, unstructured, or rapidly changing data.\n\n### Major NoSQL Data Models\n\n1. Document-Based (e.g., MongoDB): Stores semi-structured data as self-contained documents (typically JSON or BSON). Fields can vary across documents within the same collection. Suited for general web applications, content management, and e-commerce.\n2. Key-Value Stores (e.g., Redis, DynamoDB): Simplest data model pairing a unique lookup Key with an arbitrary Value payload. Optimized for high-speed in-memory caching and session management.\n3. Wide-Column / Column-Family Stores (e.g., Apache Cassandra, HBase): Organises data into dynamic column families. Designed for massively distributed analytical workloads over petabyte-scale datasets.\n4. Graph Databases (e.g., Neo4j): Explicitly stores data as Nodes (entities) connected by Edges (relationships with properties). Optimized for graph traversal queries such as social network mapping, recommendation engines, and fraud network analysis.\n\n### The BASE Consistency Model\n\nWhile relational databases strictly enforce ACID guarantees, distributed NoSQL systems often adopt the BASE model to achieve horizontal scale and high availability:\n\n- Basically Available: The system guarantees availability by responding to every request, even during node failures.\n- Soft State: System state may change over time without explicit user interaction due to ongoing background synchronization.\n- Eventual Consistency: The system guarantees that, given sufficient time without new updates, all distributed read replicas will converge to identical values.\n\nMongoDB implements document storage by organizing documents into Collections, supporting dynamic schema design and horizontal scaling via Sharding.\n\n## Comparison: SQL (Relational) vs NoSQL Databases\n\n\n\n| Aspect | SQL (Relational) | NoSQL |\n| :--- | :--- | :--- |\n| Data Storage Model | Fixed tables consisting of defined rows and columns | Varied: Documents, Key-Value pairs, Wide-Columns, or Graph Nodes/Edges |\n| Schema | Rigid — strict structure defined in advance | Flexible — dynamic schema allowing documents/records to vary |\n| Scaling Approach | Vertical (upgrading to a larger, more powerful server) | Horizontal (distributing data across clusters of commodity servers) |\n| Consistency Model | ACID (strict, immediate transactional consistency) | BASE (Eventual consistency); selective multi-document transactions |\n| Joins Across Data | Core capability — combines tables via foreign keys | Typically avoided — data is denormalized and nested together |\n| Best Suited For | Structured data requiring strict integrity (e.g., financial systems) | Large-scale, rapidly changing, unstructured data, dynamic web architectures |\n| Key Examples | PostgreSQL, Oracle, MySQL, SQL Server | MongoDB, Redis, Cassandra, DynamoDB, Neo4j |\n\n## Python and Database Access (Conceptual)\n\nPython accesses relational and NoSQL databases using database driver libraries (e.g., `psycopg2` for PostgreSQL, `pymongo` for MongoDB). The standard programmatic pattern involves:\n\n1. Establishing a secure connection string to the database host.\n2. Instantiating a Cursor or Client object.\n3. Formulating declarative query strings (SQL or JSON filters).\n4. Executing queries programmatically and fetching results into pandas `DataFrame` objects for data mining.\n5. Closing cursors and connections.\n\n## Key Terms — Week 5\n\n- Primary Key: Unique identifier column enforcing entity integrity in a relational table.\n- Foreign Key: Referential column linking a child table row to a parent table primary key.\n- SQL: Structured Query Language used to manage relational databases.\n- Normalisation: Systematic process eliminating data redundancy in relational structures.\n- ROLAP / MOLAP / HOLAP: Relational, Multidimensional, and Hybrid implementations of OLAP.\n- NoSQL: Non-relational database architectures designed for horizontal scalability and schema flexibility.\n- Document Database: NoSQL store organizing semi-structured data into JSON/BSON documents.\n- Graph Database: NoSQL store optimizing node-edge relationship traversals.\n- BASE Model: Basically Available, Soft state, Eventual consistency guarantees in NoSQL.\n\n## Self-Test Questions — Week 5\n\n1. Explain, in your own words, what a primary key and a foreign key are, and how they work together to create a relationship between two tables.\n2. What is the purpose of SQL in the context of a relational database?\n3. Why does good database/data warehouse design matter? Mention the idea of normalisation in your answer.\n4. Define a data warehouse using its four defining characteristics.\n5. Distinguish between OLAP and OLTP.\n6. Describe, conceptually, what roll-up, drill-down, slicing, and dicing each mean.\n7. Compare ROLAP, MOLAP, and HOLAP.\n8. Why did NoSQL databases emerge, and what problem were they designed to solve?\n9. Name the four major NoSQL data models covered and give one good use case for each.\n10. Compare SQL and NoSQL databases in terms of schema flexibility and scalability.\n11. What does the BASE model stand for, and how does it differ from ACID?\n12. At a conceptual level, how does a Python program interact with a database?\n\n# First Block — Week 6: Cluster Analysis — The k-Means Algorithm\n\n## What is Cluster Analysis?\n\nCluster Analysis is an unsupervised machine learning task that partitions an unlabelled dataset into natural groups (clusters). Objects placed within the same cluster exhibit high internal similarity (cohesion), while objects in different clusters exhibit high external dissimilarity (separation).\n\nBecause no ground-truth class labels are provided during training, clustering algorithms must discover patterns autonomously.\n\n### Primary Applications of Cluster Analysis\n\n- Customer Segmentation: Grouping consumers by demographic and purchasing behaviors to tailor marketing campaigns.\n- Document Clustering: Automatically organizing unlabelled news articles into thematic categories.\n- Image Segmentation: Grouping pixels with similar color or intensity profiles for computer vision.\n- Anomaly / Fraud Detection: Identifying isolated data points that fail to join any established cluster.\n- Data Summarization: Replacing large datasets with representative cluster centroids.\n\n## Representative-Based (Partitional) Clustering — The General Idea\n\nRepresentative-based (partitional) clustering algorithms divide nk distinct, non-overlapping clusters in a single partition, without creating a hierarchical tree structure.\n\nEach cluster is represented by a central point (a prototype). Partitional clustering relies on an iterative two-step optimization process:\n\n1. Assignment Step: Assign every data point to its nearest cluster representative based on a distance metric.\n2. Update Step: Recalculate the position of each cluster representative using the points currently assigned to that cluster.\n\nThis cycle repeats until cluster assignments stabilize (convergence).\n\n## How the k-Means Algorithm Works (Step by Step)\n\n\text{k-Means} is a centroid-based partitional algorithm that utilizes the Euclidean distance metric and represents each cluster prototype as the Mean (Centroid) of all data points belonging to that cluster.\n\nGiven a target number of clusters k:\n\n1. Step 1: Specify the number of clusters k.\n2. Step 2: Initialize k centroids. Starting positions are selected randomly or using initialization heuristics like k-Means++.\n3. Step 3: Assign points to nearest centroid. Calculate the squared Euclidean distance between every point x_i\mu_j:\n\nd(x_i, \mu_j) = \sum_{m=1}^{d} (x_{im} - \mu_{jm})^2\n\nAssign point x_ij yielding the minimum distance.\n\n4. Step 4: Recompute Centroids. Update the centroid of each cluster j by calculating the mean vector of all points assigned to it:\n\n\mu_j = \frac{1}{|C_j|} \sum_{x_i \in C_j} x_i\n\n5. Step 5: Iteration & Convergence. Repeat Steps 3 and 4 until one of three stopping criteria is met:\n - Centroid positions remain unchanged between iterations (\Delta \mu = 0).\n - No data points switch cluster assignments.\n - A pre-set maximum number of iterations is reached.\n\n### Properties of an Optimal k-Means Result\n\n- High Intracluster Homogeneity: Minimised Sum of Squared Errors (SSE/Inertia).\n- High Intercluster Separation: Clear boundaries between distinct cluster centroids.\n- Centrality of Representatives: Centroids accurately reflect the mean of their cluster.\n- Practical Interpretability: Clusters correspond to meaningful domain concepts.\n\n## Strengths and Weaknesses of k-Means\n\n### Strengths\n\n- Computational efficiency: Linear complexity \mathcal{O}(n \cdot k \cdot t \cdot d), making it fast on large numerical datasets.\n- Intuitive implementation and mathematical simplicity.\n- Fast convergence in practice.\n\n### Weaknesses\n\n- Requires setting the number of clusters k prior to execution.\n- Sensitive to initial centroid placement; random initialization can lead to poor local optima.\n- Sensitive to extreme outliers; because centroids rely on arithmetic means, outliers pull centroids away from true density peaks.\n- Struggles with non-globular shapes: assumes clusters are spherical and equal-sized. Fails on complex, non-convex, or elongated clusters, or clusters with varying densities.\n\n## Choosing the Right Number of Clusters (k)\n\nBecause clustering is unsupervised, identifying the optimal k requires validation metrics:\n\n1. The Elbow Method: Runs kkk=1 \dots 10k:\n\n\text{SSE} = \sum_{j=1}^{k} \sum_{x_i \in C_j} ||x_i - \mu_j||^2\n\nAs kk corresponds to the bend or "elbow" on the plot, beyond which adding more clusters yields diminishing returns in variance explained.\n\n2. Silhouette Analysis: Calculates the Silhouette Coefficient s(i)i:\n\ns(i) = \frac{b(i) - a(i)}{\max(a(i), b(i))}\n\nWhere a(i)ib(i)i to points in the closest neighboring cluster.\n\n- s(i) \approx +1: Point is well-clustered.\n- s(i) \approx 0: Point lies on the boundary between clusters.\n- s(i) \approx -1: Point is misclustered.\n\nThe k yielding the highest average Silhouette Score across all samples is selected.\n\n3. Dunn Index: Ratio of the minimum distance between any two clusters (intercluster separation) to the maximum diameter of any single cluster (intracluster compactness). Higher Dunn Index indicates superior clustering.\n\nThese represent Internal Validation metrics because they evaluate clustering structure without external class labels.\n\n## k-Medians and k-Medoids as Variants\n\nTo address k\text{-Means}' sensitivity to outliers and restriction to Euclidean spaces, two variants are used:\n\n- k-Medians: Replaces the mean centroid with the Median along each dimension, and uses Manhattan (L_1) distance instead of Euclidean distance. Highly robust to outliers, though the median representative may not be an actual observed data point.\n- k-Medoids (PAM - Partitioning Around Medoids): Selects actual observed data points (Medoids) from the dataset as cluster prototypes. Supports arbitrary dissimilarity metrics (e.g., categorical, text, or graph distances) and resists outlier distortion, but exhibits higher computational complexity \mathcal{O}(k(n-k)^2).\n\n## Comparison: k-Means vs k-Medians vs k-Medoids\n\n\n\n| Aspect | k-Means | k-Medians | k-Medoids |\n| :--- | :--- | :--- | :--- |\n| Representative Used | Mean vector of points in the cluster | Median vector of points along each dimension | An actual observed data point (Medoid) from the dataset |\n| Distance Measure | Squared Euclidean distance | Absolute-difference (Manhattan L_1) distance | Any arbitrary distance or similarity metric |\n| Sensitivity to Outliers | High — arithmetic mean is pulled by extreme values | Moderate — median resists extreme values | Low — actual representative data points anchor clusters |\n| Supported Data Types | Continuous numeric data where means exist | Continuous numeric data where medians exist | Arbitrary data types (categorical, text, graphs) |\n| Computational Speed | Fast — \mathcal{O}(n \cdot k \cdot t)k\mathcal{O}(k(n-k)^2) |\n\n## Key Terms — Week 6\n\n- Unsupervised Learning: Algorithmic learning on unlabelled data without target outcomes.\n- Cluster Analysis: Partitioning unlabelled instances into cohesive, separated groups.\n- Centroid: The arithmetic mean vector representing the center of a cluster in k\text{-Means}.\n- Inertia (SSE): Sum of squared Euclidean distances between data points and their assigned centroids.\n- Elbow Method: Heuristic identifying optimal k by finding the point of diminishing returns in SSE reduction.\n- Silhouette Score: Metric (-1+1) evaluating cluster cohesion versus separation.\n- k-Medians: Clustering variant using spatial median prototypes and Manhattan distance.\n- k-Medoids: Clustering variant utilizing actual data instances as prototypes.\n- Internal Validation: Evaluating cluster quality using data metrics without ground-truth labels.\n\n## Self-Test Questions — Week 6\n\n1. Why is clustering considered an unsupervised learning technique?\n2. Give three real-world applications of cluster analysis.\n3. What is the core idea behind representative-based/partitional clustering?\n4. List, in order, the steps of the k-Means algorithm from initial centroid choice to convergence.\n5. What are the three possible stopping conditions for k-Means?\n6. Name two weaknesses of k-Means and explain why each occurs.\n7. Explain conceptually what the elbow method looks at when choosing k.\n8. Explain conceptually what the silhouette score measures, and what a value near 1, near 0, and negative each mean.\n9. How does k-Medians differ from k-Means, and why is it considered more robust to outliers?\n10. How does k-Medoids differ from both k-Means and k-Medians in how it selects representatives?\n11. What is meant by "internal validation" of clustering quality?\n\n# First Block — Week 7: Association Pattern Mining — The Apriori Algorithm\n\n## What is Association Pattern Mining?\n\nAssociation Pattern Mining (Association Rule Mining) is an unsupervised data mining technique used to uncover relationships, co-occurrences, and dependencies among items in transactional databases.\n\nIts foundational application is Market Basket Analysis (MBA), which analyzes customer checkout transactions to identify products purchased together (e.g., "If a customer buys a computer, they are likely to also purchase antivirus software").\n\nBeyond retail, applications include cross-selling strategies, web clickstream analysis, catalog design, medical diagnosis co-occurrence analysis, and text keyword association.\n\n## Key Concepts: Itemsets, Support, Confidence\n\nLet I = {I_1, I_2, \dots, I_m}D = {T_1, T_2, \dots, T_n}TI identified by a unique Transaction ID (TID).\n\n- Itemset: A collection of zero or more items. A set containing kk\text{-itemset}.\n- Support (of an itemset XDX:\n\n\text{support}(X) = \frac{|{T_i \in D \mid X \subseteq T_i}|}{|D|}\n\n- Minimum Support Threshold (min\_sup): A user-defined cut-off. An itemset whose support \ge \text{min_sup} is designated a Frequent Itemset (or Large Itemset).\n- Association Rule: An implication expression of the form A \Rightarrow BA \subset IB \subset IA \cap B = \emptysetAB is the Consequent (right-hand side).\n- Confidence (of a rule A \Rightarrow BBA:\n\n\text{confidence}(A \Rightarrow B) = P(B \mid A) = \frac{\text{support}(A \cup B)}{\text{support}(A)}\n\n- Strong Association Rule: A rule that satisfies both the minimum support threshold (min\_sup) and minimum confidence threshold (min\_conf).\n\nRule discovery follows a two-step process: (1) Mine all frequent itemsets satisfying min\_sup, and (2) Generate strong rules from those frequent itemsets satisfying min\_conf. Step 1 carries the highest computational complexity.\n\n## The Apriori Principle (Downward Closure)\n\nMining frequent itemsets via brute-force is intractable: a universe of m2^m - 1 candidate itemsets.\n\nThe Apriori Algorithm (Agrawal & Srikant, 1994) prunes this search space using the Apriori Principle (Downward Closure Property of Support):\n\n\n\nEvery subset of a frequent itemset must also be frequent. Equivalently, if an itemset is infrequent, all of its supersets are guaranteed to be infrequent and can be pruned immediately without scanning the database.\n\nBecause support is monotonic (X \subseteq Y \implies \text{support}(Y) \le \text{support}(X)), if itemset \{\text{Beer}\} is infrequent, then \{\text{Beer}, \text{Diapers}\} cannot be frequent.\n\n## How the Apriori Algorithm Works (Step by Step)\n\nApriori uses a level-wise (breadth-first) search, discovering frequent 1-itemsets (F_1F_2F_k:\n\n1. Step 1: Count 1-Itemsets. Scan the database to calculate support for all single items (C_1F_1.\n2. Step 2: Candidate Generation (Join Step). Generate candidate (k+1)C_{k+1}kF_kk-1 items in common (under lexicographic ordering).\n3. Step 3: Candidate Pruning (Prune Step). For each candidate in C_{k+1}kF_kC_{k+1}.\n4. Step 4: Database Scan & Support Counting. Scan database DC_{k+1}.\n5. Step 5: Filter to Form F_{k+1}C_{k+1}F_{k+1}.\n6. Step 6: Loop & Terminate. Increment kF_{k+1} is empty.\n\nEach level k requires one full database scan. The total number of database scans equals the size of the maximum frequent itemset.\n\n## Generating Association Rules from Frequent Itemsets\n\nOnce all frequent itemsets are identified, candidate rules are generated:\n\n1. For each frequent itemset l \in Fs \subset l\n2. Formulate candidate rules s \Rightarrow (l - s).\n3. Compute confidence: \text{confidence} = \frac{\text{support}(l)}{\text{support}(s)}.\n4. Retain rules satisfying min\_conf.\n\nRule generation requires no additional database scans, as itemset supports were recorded during mining. Confidence exhibits monotonic behaviour for a fixed frequent itemset: as antecedent s shrinks, confidence cannot increase.\n\n## Beyond Confidence: Why Support/Confidence Can Mislead\n\nThe support-confidence framework can generate misleading rules when the consequent item is globally common. If 90% of all customers buy Milk regardless of other purchases, a rule \text{Tea} \Rightarrow \text{Milk} will exhibit high confidence even if Tea and Milk are independent or negatively associated.\n\nTo correct for baseline item popularity, additional metrics are evaluated:\n\n- Lift: Measures how much more often the antecedent and consequent co-occur than expected if they were statistically independent:\n\n\text{lift}(A \Rightarrow B) = \frac{\text{support}(A \cup B)}{\text{support}(A) \times \text{support}(B)}\n\n - \text{Lift} = 1AB are independent; no association.\n - \text{Lift} > 1AB\n - \text{Lift} < 1AB\n\n- Conviction: Compares the expected frequency that AB assuming independence against the observed frequency of incorrect predictions:\n\n\text{conviction}(A \Rightarrow B) = \frac{1 - \text{support}(B)}{1 - \text{confidence}(A \Rightarrow B)}\n\n - Conviction is directional and evaluates to infinity for uncontradicted rules.\n\nOther metrics include Pearson Correlation (\phi\chi^2) test of independence, Cosine, and Jaccard similarity.\n\n## Comparison: Support vs Confidence vs Lift\n\n\n\n| Measure | What It Measures | Conceptual Formula | Interpretation |\n| :--- | :--- | :--- | :--- |\n| Support | Frequency of itemset/rule occurrence in the dataset | \text{support}(X \cup Y) = \frac{|X \cup Y|}{|D|} | High support indicates a common pattern; filters statistically insignificant patterns |\n| Confidence | Conditional reliability of consequent YX\frac{\text{support}(X \cup Y)}{\text{support}(X)}P(Y \mid X)Y is globally common |\n| Lift | Co-occurrence frequency relative to random independence | \frac{\text{support}(X \cup Y)}{\text{support}(X) \times \text{support}(Y)}=1>1<1 negative association; corrects for popular consequents |\n\n## Key Terms — Week 7\n\n- Market Basket Analysis: Mining transactional purchase data to identify co-occurring items.\n- Itemset: A collection of one or more items.\n- Support: Proportion of total transactions containing a specific itemset.\n- Frequent Itemset: An itemset meeting or exceeding the minimum support threshold (min\_sup).\n- Association Rule: Implications of the form A \Rightarrow B meeting support and confidence cut-offs.\n- Confidence: Conditional probability P(B \mid A) evaluating rule reliability.\n- Downward Closure (Apriori Principle): Property asserting that all subsets of a frequent itemset are frequent.\n- Candidate Generation: Join operation forming (k+1)k\text{-itemsets}.\n- Candidate Pruning: Eliminating candidate itemsets containing any infrequent subset.\n- Lift: Ratio evaluating co-occurrence frequency against statistical independence.\n\n## Self-Test Questions — Week 7\n\n1. Define support and confidence in your own words, and explain the difference between what each one tells you about a rule.\n2. State the Apriori (downward closure) principle. Why does it guarantee that Apriori will never miss a truly frequent itemset while still pruning the search space?\n3. Walk through, in plain language, what happens at each level of the Apriori algorithm (from counting 1-itemsets through to termination).\n4. Why is the join step in Apriori restricted to itemsets that share their first (k-1) items, rather than joining any two frequent k-itemsets?\n5. Given frequent itemset {A, B, C} with support 0.3, and sup({A,B}) = 0.4, calculate the confidence of {A, B} => {C}.\n6. Explain why a rule can have very high confidence yet be practically useless from a business perspective. How does lift address this problem?\n7. List three real-world applications of association pattern mining beyond supermarket shelf placement.\n8. Why is support counting considered the most computationally expensive step of Apriori, and what design choices help reduce its cost?\n\n# First Block — Week 8: Multivariate Regression, Nonlinear/Polynomial Regression & the CART Algorithm\n\n## What is Regression Analysis?\n\nRegression Analysis is a supervised learning methodology for modeling the relationship between one or more independent predictor variables (features) and a continuous numeric dependent response variable (target).\n\nUnlike classification, which predicts discrete categorical class labels, regression predicts continuous scalar values (e.g., house prices, income levels, temperatures).\n\n## Simple vs Multivariate (Multiple) Linear Regression\n\n- Simple Linear Regression: Models the relationship between a single independent variable xy via a straight line:\n\ny = w_0 + w_1 x + \epsilon\n\nWhere w_0w_1\epsilon represents unobserved random error.\n\n- Multivariate (Multiple) Linear Regression: Extends linear modeling to dX_1, X_2, \dots, X_d simultaneously:\n\ny = w_0 + w_1 X_1 + w_2 X_2 + \dots + w_d X_d + \epsilon\n\nMultivariate regression captures joint predictor influences on the outcome variable.\n\n## How Linear Regression Finds the "Best Fit" (Conceptually)\n\nLinear regression computes parameters \mathbf{w}y_i\hat{y}i:\n\n\text{SSR} = \sum{i=1}^{n} (y_i - \hat{y}i)^2 = \sum{i=1}^{n} (y_i - (w_0 + \sum_{j=1}^{d} w_j x_{ij}))^2\n\nSquaring residuals ensures positive and negative errors do not cancel out, while penalizing larger deviations. OLS yields an analytical closed-form solution via matrix algebra (Normal Equations):\n\n\mathbf{w} = (\mathbf{X}^T \mathbf{X})^{-1} \mathbf{X}^T \mathbf{y}\n\nTo prevent overfitting when predictors exhibit collinearity, Regularisation techniques are applied: Ridge Regression (L_2L_1\text{-norm penalty promoting sparse feature weights}).\n\n## Assumptions of Linear/Nonlinear Regression\n\n1. Correct Functional Form: The model functional structure reflects the true data-generating relationship.\n2. Independence of Errors: Observations and residual errors are un-correlated.\n3. Homoscedasticity: Constant variance of residual errors across all levels of independent variables.\n4. Normality of Residuals: Error terms \epsilon follow a normal distribution centered at zero.\n5. Absence of Multicollinearity: Independent predictor variables do not exhibit strong linear correlation.\n\n## Nonlinear Regression: When a Straight Line Doesn't Fit\n\nWhen data exhibits non-linear, curved patterns (e.g., exponential decay, saturation curves), fitting a linear model produces poor fits and patterned residuals.\n\nNonlinear Regression models arbitrary non-linear relationships y = f(\mathbf{X}, \mathbf{\beta}) + \epsilon:\n\n- Parametric Nonlinear Regression: Assumes a specific functional form (e.g., exponential y = a e^{b x}y = a \ln(x) + b).\n- Non-Parametric Nonlinear Regression: Learns functional shapes directly from data without predefined formulas (e.g., kernel regression, local regression).\n\n## Polynomial Regression: Fitting Curves with Powers of Variables\n\nPolynomial Regression fits non-linear curves by extending linear regression to include higher-degree power terms of the predictor variables:\n\ny = \beta_0 + \beta_1 x + \beta_2 x^2 + \beta_3 x^3 + \dots + \beta_n x^n + \epsilon\n\nAlthough the fitted curve is non-linear in predictor x\beta_j. OLS matrix solver equations remain directly applicable.\n\nThe polynomial degree n controls model flexibility:\n\n- Too low a degree (n=1): Underfitting.\n- Too high a degree (n excessive): Overfitting; the curve oscillates to fit training noise, failing to generalize.\n\n## The CART Algorithm: Decision Trees for Classification and Regression\n\nClassification And Regression Trees (CART) (Breiman et al., 1984) is a non-parametric recursive binary splitting algorithm that builds decision trees for discrete classification and continuous regression tasks.\n\n### Binary Tree Construction Steps\n\n1. Step 1: Start with full training data at the root node.\n2. Step 2: Evaluate candidate binary splits across all predictor features. For continuous features, test numeric thresholds; for categorical features, test category sub-sets.\n3. Step 3: Select the split minimizing impurity or error.\n - Classification Trees: CART uses the Gini Index to evaluate node impurity. For a node tp_k:\n\n\text{Gini}(t) = 1 - \sum_{k=1}^{K} p_k^2\n\nCART selects the feature split yielding the lowest weighted average Gini index across the two child nodes.\n\n - Regression Trees: CART evaluates node impurity using Variance Reduction (minimizing the Sum of Squared Errors relative to the child node means):\n\n\text{Variance}(t) = \frac{1}{N_t} \sum_{i \in t} (y_i - \bar{y}t)^2\n\n4. Step 4: Recursively partition child nodes.\n5. Step 5: Stop and Prune. Tree growth stops when nodes become pure, reach minimum sample thresholds, or hit maximum depth. Fully grown trees overfit training data. Cost-Complexity Pruning cuts back weak branches based on performance against a validation dataset.\n\n### Classification vs Regression Predictions\n\n- Classification Trees: Leaf nodes assign the majority class label of training samples falling within that leaf.\n- Regression Trees: Leaf nodes output the average scalar target value \bar{y} of training samples falling within that leaf.\n\n## Comparison: Linear vs Nonlinear vs Polynomial Regression\n\n| Aspect | Linear Regression | Nonlinear Regression (General) | Polynomial Regression |\n| :--- | :--- | :--- | :--- |\n| Fitting Method | Ordinary Least Squares (OLS) closed-form matrix math | Iterative numerical optimization (e.g., Gauss-Newton) | Ordinary Least Squares (OLS) on expanded power features |\n| When to Use | Linear relationship between predictors and response | Data follows non-linear curves (growth, decay, saturation) | Response exhibits curved trends achievable via powers of x |\n| Assumed Relationship | Straight line or flat hyperplane | Arbitrary mathematical curves | Curvilinear polynomial shape |\n| Linear in Coefficients? | Yes | No (parameters appear non-linearly inside function) | Yes — linear with respect to parameters \beta_j |\n| Primary Risk | Underfitting if real trend is curved | Mis-specifying functional curve form | Overfitting if polynomial degree n is set too high |\n\n## Key Terms — Week 8\n\n- Regression Analysis: Modeling continuous quantitative target variables from predictor features.\n- Ordinary Least Squares (OLS): Parameter estimation method minimizing the sum of squared residuals.\n- Homoscedasticity: Assumption that regression residual error variance remains constant across predictors.\n- Multicollinearity: Strong correlation among independent predictor variables.\n- Polynomial Regression: Linear regression variant incorporating exponentiated feature powers.\n- CART (Classification and Regression Trees): Binary decision tree algorithm relying on recursive partitioning.\n- Gini Index: Measure of class impurity used by CART for classification tree splits (0 = \text{pure node}).\n- Variance Reduction: Split criterion used by CART regression trees to minimize squared prediction errors.\n- Cost-Complexity Pruning: Trimming tree branches based on validation error to prevent overfitting.\n\n## Self-Test Questions — Week 8\n\n1. Explain the difference between simple and multiple (multivariate) linear regression, and give an example of when you would need the multivariate version.\n2. In plain language, what is the "least squares" idea, and why do we square the errors rather than just summing them directly?\n3. Why can't ordinary linear regression capture a relationship that looks like a curve (e.g. exponential growth)? What two broad families of models can be used instead?\n4. Explain why polynomial regression is still considered "linear" in a technical sense even though it produces a curved fit.\n5. What risk grows as you increase the degree of a polynomial regression model, and how would you recognise it?\n6. Describe, step by step, how CART builds a decision tree from a training dataset.\n7. What is the Gini index measuring, and how does CART use it to choose the best split at a node?\n8. How does a CART regression tree differ from a CART classification tree in terms of its split criterion and its leaf predictions?\n9. Why is pruning necessary for decision trees, and how is a holdout/validation set typically used to decide what to prune?\n10. List one strength and one weakness of tree-based models such as CART compared to a linear regression model.\n\n# Second Block — Week 2: Logistic Regression & Classifier Evaluation\n\n## What is Machine Learning?\n\nMachine Learning (ML) is a branch of Artificial Intelligence where computer algorithms learn statistical patterns from data to make predictions or decisions on unseen data without explicit task-specific rules.\n\nThe machine learning pipeline comprises data ingestion, preprocessing, splitting (train/validation/test), algorithm selection, training, hyperparameter tuning, model evaluation, and deployment.\n\nClassification is a supervised learning task focused on predicting discrete categorical class labels based on input features.\n\n## What is Logistic Regression and How Does It Differ from Linear Regression?\n\nLogistic Regression is a supervised discriminative classification algorithm that estimates the probability that an input instance belongs to a specific binary target class.\n\nLinear regression cannot be applied to classification because:\n\n1. Linear regression outputs unbounded continuous values (-\infty+\infty[0, 1]).\n2. Linear regression is sensitive to extreme outliers, which shift the regression line and unpredictably alter decision threshold cut-offs.\n\nLogistic Regression solves this by taking a linear combination of features z = w_0 + \mathbf{w}^T \mathbf{x} and passing it through the non-linear Sigmoid (Logistic) Function:\n\n\sigma(z) = \frac{1}{1 + e^{-z}} = \frac{1}{1 + e^{-(w_0 + \mathbf{w}^T \mathbf{x})}}\n\nThe Sigmoid function squashes any real value into the interval (0, 1)P(y=1 \mid \mathbf{x}).\n\n- When z = 0\sigma(z) = 0.5 (the Decision Boundary).\n- When z > 0\sigma(z) \to 1.0\n- When z < 0\sigma(z) \to 0.0\n\nGeometrically, Logistic Regression defines a linear decision boundary hyperplane separating classes. Parameters are estimated using Maximum Likelihood Estimation (MLE) via iterative optimization (e.g., Gradient Descent) rather than OLS.\n\nLogistic regression coefficients represent the change in log-odds of the target event per unit change in the feature:\n\n\text{logit}(p) = \ln\left(\frac{p}{1-p}\right) = w_0 + w_1 X_1 + \dots + w_d X_d\n\n## Types of Logistic Regression\n\n- Binary Logistic Regression: Target has two mutually exclusive classes (e.g., 01, Fraud or Legitimate).\n- Multinomial Logistic Regression: Target has three or more unordered categorical classes (e.g., predicting transport mode: Bus, Train, Car).\n- Ordinal Logistic Regression: Target has three or more ordered categorical classes (e.g., survey ratings: Low, Medium, High).\n\n## Assumptions of Logistic Regression\n\n- Binary/Categorical Target: Target variable must be discrete.\n- Independence of Observations: Sample records must be independent.\n- Linearity of Log-Odds: Predictor variables must be linearly related to the log-odds of the outcome (verified via Box-Tidwell test).\n- Absence of Multicollinearity: Low correlation among predictors (verified via Variance Inflation Factor - VIF).\n- Sample Size: Requires sufficient sample size per feature for reliable MLE convergence.\n\n## Classifier Evaluation\n\nEvaluating a classifier requires proper data splitting and robust evaluation metrics:\n\n### Data Splitting Methodologies\n\n- Holdout Method: Single random split into Training Set (e.g., 70%) and Test Set (e.g., 30%). Simple, but accuracy estimates can vary based on split randomness.\n- k-Fold Cross-Validation: Dataset is partitioned into kk-1kk = n\n- Bootstrap: Resampling with replacement to generate training sets. Useful for small datasets, but over-represents duplicated training samples.\n\n### The Confusion Matrix\n\nFor a binary classification task, predictions are cross-tabulated against actual ground-truth labels:\n\n| | Predicted Negative (01) |\n| :--- | :--- | :--- |\n| Actual Negative (0) | True Negative (TN) | False Positive (FP) — Type I Error |\n| Actual Positive (1) | False Negative (FN) — Type II Error | True Positive (TP) |\n\n### Key Evaluation Metrics\n\n- Accuracy: Overall proportion of correct predictions:\n\n\text{Accuracy} = \frac{\text{TP} + \text{TN}}{\text{TP} + \text{TN} + \text{FP} + \text{FN}}\n\n - Warning: Accuracy fails on Imbalanced Datasets (e.g., if 99% of transactions are legitimate, a dummy model predicting all transactions as legitimate achieves 99% accuracy while missing all fraud cases).\n\n- Precision: Proportion of positive predictions that were truly positive:\n\n\text{Precision} = \frac{\text{TP}}{\text{TP} + \text{FP}}\n\n - Critical when the cost of False Positives is high (e.g., spam filtering, wrongful arrests).\n\n- Recall (Sensitivity / True Positive Rate): Proportion of actual positive cases correctly identified:\n\n\text{Recall} = \frac{\text{TP}}{\text{TP} + \text{FN}}\n\n - Critical when the cost of False Negatives is high (e.g., medical diagnostics, fraud detection).\n\n- F1-Score: Harmonic mean of precision and recall, providing a balanced single metric for imbalanced classes:\n\n\text{F1-Score} = 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}\n\n- ROC Curve & AUC: Receiver Operating Characteristic (ROC) curve plots True Positive Rate (Sensitivity) on the y-axis against False Positive Rate (1 - \text{Specificity}\text{AUC} = 1.0\text{AUC} = 0.5 random guessing).\n\n## Comparison: Accuracy vs Precision vs Recall vs F1-Score\n\n| Metric | What It Measures | Conceptual Formula | When It Matters Most |\n| :--- | :--- | :--- | :--- |\n| Accuracy | Proportion of all predictions that were correct | \frac{\text{TP}+\text{TN}}{\text{Total}} | Balanced classes where all error types carry equal costs |\n| Precision | Proportion of predicted positive calls that were truly positive | \frac{\text{TP}}{\text{TP}+\text{FP}} | When False Positives are costly (e.g., spam detection, loan default flagging) |\n| Recall | Proportion of actual positive instances correctly caught | \frac{\text{TP}}{\text{TP}+\text{FN}} | When False Negatives are costly (e.g., cancer detection, security breaches) |\n| F1-Score | Harmonic mean balancing precision and recall | 2 \cdot \frac{P \cdot R}{P + R} | Imbalanced dataset evaluation when both error types carry significance |\n\n## Key Terms — Week 2 (Second Block)\n\n- Machine Learning: Algorithms learning patterns from data to generate predictions on unseen instances.\n- Logistic Regression: Supervised linear classification algorithm mapping log-odds via a Sigmoid function.\n- Sigmoid Function: S-shaped curve mapping real values to probabilities in interval (0, 1).\n- Discriminative Model: Model learning class conditional probabilities P(y \mid \mathbf{x}) directly.\n- Holdout Method: Single dataset partition into distinct training and testing subsets.\n- Cross-Validation: Resampling method averaging model performance across k validation folds.\n- Confusion Matrix: Matrix cross-tabulating actual versus predicted classification counts.\n- Precision: Fraction of positive calls that are true positives (\frac{\text{TP}}{\text{TP}+\text{FP}}).\n- Recall: Fraction of true positive cases correctly identified (\frac{\text{TP}}{\text{TP}+\text{FN}}).\n- ROC Curve / AUC: Metric evaluating true positive vs. false positive trade-offs across all decision thresholds.\n\n## Self-Test Questions — Week 2 (Second Block)\n\n1. Why can't linear regression be used directly for a classification task?\n2. What role does the sigmoid/logistic function play in logistic regression?\n3. Give an example each of binary, multinomial, and ordinal logistic regression.\n4. Why must a model never be evaluated on the same data it was trained on?\n5. Define true positive, false positive, true negative, and false negative in your own words.\n6. A hospital wants to screen patients for a serious disease. Should it prioritise precision or recall, and why?\n7. Why can accuracy be a misleading metric on an imbalanced dataset, and what metric(s) would you use instead?\n8. What does the area under the ROC curve (AUC) represent, and what would an AUC of 0.5 imply about a classifier?\n9. Explain the difference between overfitting and underfitting, and how cross-validation helps detect/avoid overfitting.\n10. Why does the holdout method tend to give a pessimistic estimate of accuracy compared to using all the data for training?\n\n# Second Block — Week 3: Decision Trees and Naive Bayes\n\n## Decision Tree Terminology\n\nA Decision Tree is a non-parametric supervised model representing hierarchical decisions on feature variables as a tree structure.\n\n- Root Node: Topmost node containing the entire unsplit dataset.\n- Internal (Decision) Node: Intermediate node testing a feature split condition.\n- Leaf (Terminal) Node: End node holding the final predicted class label or numeric average.\n- Branch: Directional path representing the outcome of a decision condition.\n- Sub-tree: Discrete branch portion beneath a specified internal node.\n\n## How a Decision Tree Decides Where to Split\n\nDecision trees select splits using mathematical Impurity Measures:\n\n1. Gini Index (CART): Measures class mixing at node t:\n\n\text{Gini}(t) = 1 - \sum{k=1}^{K} p_k^2\n\n - \text{Gini} = 0: Pure node (single class present).\n - CART computes weighted average Gini indices across candidate child splits, selecting the split yielding the lowest weighted Gini index.\n\n2. Entropy & Information Gain (ID3 / C4.5): Entropy measures disorder at node t:\n\n\text{Entropy}(t) = -\sum_{k=1}^{K} p_k \log_2(p_k)\n\n - Information Gain measures the reduction in entropy achieved by a feature split:\n\n\text{GAIN}{split} = \text{Entropy}(\text{parent}) - \sum{c} \frac{N_c}{N} \text{Entropy}(\text{child}c)\n\n - Algorithms select splits maximizing Information Gain (or Gain Ratio in C4.5 to penalize multi-way branching).\n\n3. Error Rate: Misclassification error rate (1 - \max(p_k)). Less sensitive to class distribution changes than Gini or Entropy.\n\nSplits can be Univariate (evaluating a single feature) or Multivariate (evaluating linear feature combinations).\n\n## Tree-Building Process (Recursive Partitioning) and Stopping Criteria\n\nDecision trees are constructed top-down using Recursive Partitioning. Starting at the root node, the algorithm greedily selects the best available split feature and threshold, dividing data into child nodes, and recursively repeats the process.\n\nStopping criteria include:\n\n- All leaf samples belong to a single class.\n- Maximum tree depth limit reached.\n- Node sample size drops below a minimum threshold.\n- Information gain drops below a minimum threshold.\n\n## Overfitting and Pruning\n\nUnconstrained decision trees grow deep until leaves are pure, memorizing training noise and overfitting. Pruning restores generalisation:\n\n- Pre-Pruning (Early Stopping): Halting tree growth early based on depth or sample thresholds.\n- Post-Pruning: Growing a full tree, then cutting back sub-trees using validation set evaluations or cost-complexity metrics.\n\n## Rule-Based Classifiers (Briefly)\n\nDecision trees can be translated into Rule-Based Classifiers consisting of IF-THEN rules: each root-to-leaf path forms one rule antecedent, and the leaf label forms the consequent. Extracted rule sets can be pruned and reordered independently of tree structure.\n\n## Naive Bayes: The Intuition Behind Bayes' Theorem\n\nNaive Bayes is a probabilistic generative classifier grounded in Bayes' Theorem:\n\nP(Y \mid \mathbf{X}) = \frac{P(\mathbf{X} \mid Y) P(Y)}{P(\mathbf{X})}\n\n- P(Y \mid \mathbf{X})Y\mathbf{X}).\n- P(\mathbf{X} \mid Y)\mathbf{X}Y).\n- P(Y)Y).\n- P(\mathbf{X}): Marginal Likelihood (evidence normalization factor).\n\n## Why It's Called "Naive"\n\nEstimating joint likelihood P(X_1, X_2, \dots, X_d \mid Y) directly requires massive data. Naive Bayes simplifies this by assuming Conditional Independence among features given the class label:\n\nP(X_1, X_2, \dots, X_d \mid Y) = \prod{j=1}^{d} P(X_j \mid Y)\n\nThis "naive" assumption asserts that features do not interact given the class label. Despite violating real-world feature correlations, Naive Bayes performs well in high-dimensional text classification.\n\nTo prevent zero-probability errors when an unseen feature value yields P(X_j \mid Y) = 0, Laplacian Smoothing is applied:\n\nP(X_i \mid Y) = \frac{c(X_i, Y) + 1}{c(Y) + |V|}\n\n## Types of Naive Bayes\n\n- Gaussian Naive Bayes: Used for continuous numeric features, assuming features follow a Normal distribution within each class.\n- Multinomial Naive Bayes: Used for discrete word count features in text analytics.\n- Bernoulli Naive Bayes: Used for binary feature presence/absence indicators.\n\n## Strengths and Weaknesses of Naive Bayes\n\n### Strengths\n\n- High computational speed; linear complexity \mathcal{O}(n \cdot d).\n- Highly effective for high-dimensional text classification (spam filtering, sentiment analysis).\n- Robust to irrelevant features.\n\n### Weaknesses\n\n- Violated conditional independence assumptions distort probability estimates.\n- Cannot capture feature interactions.\n\n## Comparison: Decision Trees vs Naive Bayes\n\n| Aspect | Decision Trees | Naive Bayes |\n| :--- | :--- | :--- |\n| Model Type | Discriminative; hierarchical IF-THEN splits | Generative / Probabilistic; models feature distributions per class |\n| Key Mechanism | Recursive partitioning using impurity metrics (Gini, Entropy) | Bayes' Theorem with conditional independence assumption |\n| Handles Feature Correlation | Yes — models complex multi-feature interactions via sequential splits | No — assumes features are conditionally independent given the class |\n| Best Suited For | Tabular datasets with mixed categorical/numeric features and interpretable rules | High-dimensional data (e.g., text document classification) |\n| Interpretability | Very High — visual flowchart structure | Moderate — relies on probability scores |\n| Main Weakness | Prone to severe overfitting if grown deep without pruning | Conditional independence assumption is frequently unrealistic |\n\n## Key Terms — Week 3 (Second Block)\n\n- Root / Decision / Leaf Nodes: Structural elements of hierarchical decision trees.\n- Gini Index: Impurity measure evaluating class mixing for tree splits.\n- Entropy / Information Gain: Information-theoretic metrics measuring disorder reduction.\n- Recursive Partitioning: Greedy top-down tree building methodology.\n- Post-Pruning: Trimming tree branches using validation data to counter overfitting.\n- Bayes' Theorem: Mathematical rule updating posterior probability using prior and likelihood.\n- Conditional Independence: The Naive Bayes assumption that features are independent given the class label.\n- Laplacian Smoothing: Technique adding uniform constants to counts to eliminate zero probabilities.\n- Gaussian / Multinomial / Bernoulli Naive Bayes: Variants for continuous, count, and binary features.\n\n## Self-Test Questions — Week 3 (Second Block)\n\n1. Define root node, decision node, leaf node, and branch in a decision tree.\n2. How does the Gini index measure the "mixedness" of a group of data points?\n3. What is information gain, and how is it related to entropy?\n4. Describe the recursive partitioning process used to build a decision tree, and name one stopping criterion.\n5. Why does an unpruned decision tree tend to overfit, and how does pruning address this?\n6. How can a set of classification rules be derived from a decision tree?\n7. In your own words, explain what Bayes' theorem calculates and why it is useful when the posterior probability is hard to estimate directly.\n8. Why is naive Bayes called "naive," and what specific assumption does that name refer to?\n9. Which type of naive Bayes would you use for word-count features in a document, and which for continuous numeric features?\n10. Give one reason naive Bayes performs well in practice despite its unrealistic independence assumption, and one scenario where this assumption could seriously hurt its accuracy.\n\n# Second Block — Week 4: Neural Networks\n\n## What is a Neural Network?\n\nArtificial Neural Networks (ANNs) are computational models inspired by biological neural networks. They consist of interconnected nodes (artificial neurons) arranged in layers, connected by adjustable synaptic weights. Learning occurs by updating connection weights based on training errors.\n\n## The Perceptron\n\nThe Perceptron (Rosenblatt, 1958) is the simplest neural network, comprising an Input Layer and a single Output Node.\n\n### Computation Mechanics\n\n1. Multiply input features X_iw_i.\n2. Compute weighted sum and add bias bz = \sum_{i=1}^{d} w_i X_i + b\n3. Pass z through a Step Activation Function:\n\n\hat{y} = \begin{cases} +1 & \text{if } z \ge 0 \ -1 & \text{if } z < 0 \end{cases}\n\n### Perceptron Learning Rule\n\nWeights are updated iteratively after evaluating each training sample:\n\nw_i \leftarrow w_i + \eta (y - \hat{y}) X_i\n\nWhere \eta represents the Learning Rate parameter.\n\n## Why a Single Perceptron is Limited\n\nA single Perceptron can only construct linear decision boundaries. It can only solve Linearly Separable problems (e.g., AND, OR logic gates). It cannot solve non-linearly separable problems such as the XOR (Exclusive-OR) logic function.\n\n## Multilayer Neural Networks\n\nTo capture non-linear decision boundaries, Multilayer Perceptrons (MLPs) incorporate one or more Hidden Layers between input and output layers.\n\nAccording to the Universal Approximation Theorem, a feedforward network with a single hidden layer and non-linear activation functions can approximate any continuous function to arbitrary accuracy.\n\n## How a Multilayer Network Learns: Backpropagation (Conceptual)\n\nMultilayer networks are trained using Backpropagation combined with Gradient Descent:\n\n1. Forward Pass: Input features propagate forward through network layers, applying weights, biases, and activation functions to generate predicted outputs \hat{y}.\n2. Loss Calculation: A loss function (e.g., Mean Squared Error or Cross-Entropy) quantifies the error between prediction \hat{y}y\n3. Backward Pass (Backpropagation): Using the calculus Chain Rule, prediction errors propagate backwards through hidden layers. Derivatives of loss with respect to each weight are calculated to allocate error responsibility.\n4. Weight Update: Weights are updated opposite the gradient direction:\n\nw \leftarrow w - \eta \frac{\partial \text{Loss}}{\partial w}\n\nTraining cycles through all records over multiple Epochs until convergence.\n\n## Activation Functions (Conceptual Role)\n\nActivation functions introduce non-linearity into the network. Without non-linear activation functions, stacking multiple layers collapses mathematically into a single linear regression equation.\n\n- Sigmoid: \sigma(z) = \frac{1}{1 + e^{-z}}(0, 1).\n- Hyperbolic Tangent (tanh): \tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}}(-1, +1).\n- Rectified Linear Unit (ReLU): f(z) = \max(0, z). Solves vanishing gradient problems in deep networks.\n\n## Strengths and Weaknesses\n\n### Strengths\n\n- Universal function approximation; models highly complex, non-linear relationships.\n- Feature learning: automatically discovers hidden feature representations.\n\n### Weaknesses\n\n- "Black-box" nature: lacks interpretability.\n- High computational cost; requires long training times and large training datasets.\n- Susceptible to local minima and severe overfitting without regularisation (e.g., Dropout, weight decay).\n\n## Comparison: Perceptron vs Multilayer Neural Network\n\n| Aspect | Perceptron (Single Layer) | Multilayer Neural Network (MLP) |\n| :--- | :--- | :--- |\n| Boundary Capability | Strictly linear boundaries | Complex, non-linear, non-contiguous boundaries |\n| Ground Truth for Training | Known directly at output node | Estimated for hidden units via Backpropagation |\n| Layer Architecture | Input layer + Single output node | Input layer + One or more Hidden layers + Output layer |\n| Activation Function | Step / Threshold function | Non-linear functions (Sigmoid, tanh, ReLU) |\n| Learning Algorithm | Perceptron learning rule | Backpropagation via Gradient Descent |\n| Problem Types | Linearly separable problems only (e.g., AND, OR) | Arbitrary non-linear problems (e.g., XOR, vision, speech) |\n| Interpretability | High — direct linear weight inspection | Low — complex "black-box" hidden network dynamics |\n\n## Key Terms — Week 4 (Second Block)\n\n- Artificial Neural Network (ANN): Interconnected node networks modeling non-linear functions.\n- Perceptron: Two-layer linear neural network using step activation functions.\n- Hidden Layer: Intermediate node layers enabling non-linear functional mapping.\n- Linear Separability: Data state separable by a straight line or flat hyperplane.\n- Backpropagation: Algorithm calculating loss gradients backwards to update network weights.\n- Learning Rate (\eta): Step-size hyperparameter governing weight updates in gradient descent.\n- Epoch: One full training pass through the complete dataset.\n- Activation Function: Non-linear mathematical function applied at hidden/output nodes.\n\n## Self-Test Questions — Week 4 (Second Block)\n\n1. Explain, in plain language, what happens inside a perceptron's output node when it processes an input record.\n2. Why can a single-layer perceptron only solve linearly separable problems?\n3. What is the role of a hidden layer in a multilayer neural network?\n4. Describe the two phases of backpropagation and what each one does.\n5. Why are activation functions necessary for a network to learn non-linear patterns?\n6. Give two strengths and two weaknesses of neural networks as classifiers.\n7. Why might a smaller learning rate be preferred later in training even though it slows learning down?\n\n# Second Block — Week 5: Support Vector Machines & Feature Selection\n\n## What is a Support Vector Machine (SVM)?\n\nA Support Vector Machine (SVM) is a supervised discriminative classification methodology that identifies the Optimal Separating Hyperplane maximizing the Margin between two classes.\n\n- Hyperplane: A flat decision boundary of dimension d-1d\text{-dimensional} feature space.\n- Margin: The perpendicular distance between the decision hyperplane and the closest training data points of either class.\n- Support Vectors: The critical training data points lying directly on the margin boundaries. They anchor the mathematical position of the optimal hyperplane; removing non-support vectors leaves the decision boundary unchanged.\n\n## Why Maximising the Margin Helps Generalisation\n\nMaximizing the margin minimizes structural risk. A narrow margin leaves the boundary vulnerable to noise and small variations in new data points. A maximum margin provides maximum separation, lowering generalization error on unseen data.\n\n## Linearly Separable vs Non-linearly Separable Data, and Soft Margins\n\n- Hard Margin SVM: Assumes data is strictly linearly separable. No training points are permitted to cross the margin boundaries.\n- Soft Margin SVM (C-SVM): Real-world data contains noise, overlaps, and outliers. Soft Margin SVM introduces Slack Variables (\xi_i) allowing controlled margin violations.\n\nSlack parameter C controls the regularization trade-off:\n\n- High C: Enforces strict penalties for margin violations, prioritizing training accuracy at the risk of creating a narrower margin that overfits.\n- Low C: Permits more margin violations, widening the margin to improve generalisation at the risk of underfitting.\n\n## Non-linear Data and the Kernel Trick\n\nWhen data is non-linearly separable in its original space, SVM projects the data into a higher-dimensional feature space where a linear separating hyperplane can be found.\n\nComputing explicit higher-dimensional transformations is computationally expensive. SVM avoids explicit transformation using the Kernel Trick. A Kernel Function K(\mathbf{x}i, \mathbf{x}_j) computes pairwise dot-products directly in the implicit higher-dimensional feature space using original input space vectors:\n\nK(\mathbf{x}_i, \mathbf{x}_j) = \langle \Phi(\mathbf{x}_i), \Phi(\mathbf{x}_j) \rangle\n\nCommon Kernel Functions:\n\n- Linear Kernel: K(\mathbf{x}_i, \mathbf{x}_j) = \mathbf{x}_i^T \mathbf{x}_j\n- Polynomial Kernel: K(\mathbf{x}_i, \mathbf{x}_j) = (\mathbf{x}_i^T \mathbf{x}_j + c)^d\n- Radial Basis Function (RBF / Gaussian) Kernel: K(\mathbf{x}_i, \mathbf{x}_j) = \exp(-\gamma ||\mathbf{x}_i - \mathbf{x}_j||^2\n\n## Types/Variants of SVM\n\n- Linear SVM: Fast implementation for linearly separable or high-dimensional text data.\n- Non-Linear (Kernel) SVM: Flexible non-linear classifier utilizing RBF or Polynomial kernels.\n- Support Vector Regression (SVR): Extension of SVM principles to continuous target prediction.\n\n## Strengths and Weaknesses of SVMs\n\n### Strengths\n\n- Effective in high-dimensional spaces (e.g., text, bioinformatics).\n- Robust to noise because decision boundaries rely exclusively on support vectors.\n- Versatile modeling via custom kernel functions.\n\n### Weaknesses\n\n- High computational complexity \mathcal{O}(n^2 \dots n^3), making it slow on massive datasets.\n- Sensitive to hyperparameter choices (C\gamma).\n- Lacks direct probabilistic output and model interpretability.\n\n## Feature Selection: Why It Matters\n\nIncluding irrelevant, noisy, or redundant features degrades machine learning models by causing overfitting, increasing training times, and triggering the Curse of Dimensionality. Feature Selection isolates the optimal subset of informative features prior to training.\n\n## Filter, Wrapper, and Embedded Methods (Conceptual)\n\n1. Filter Methods: Evaluate feature relevance independently of any machine learning model using statistical or information-theoretic metrics. Fast and computationally cheap, but ignores model-specific feature interactions.\n2. Wrapper Methods: Use a specific machine learning model as an evaluation engine. Sub-sets of features are repeatedly added or removed (e.g., Forward Selection, Backward Elimination, Recursive Feature Elimination - RFE), training and evaluating the model on each iteration. Produces optimal accuracy for that model, but is computationally expensive.\n3. Embedded Methods: Perform feature selection during model training (e.g., Lasso L_1 regularization driving feature weights to zero, or decision tree feature importance scores).\n\n## Filter Measures: Gini Index, Entropy, Fisher Score\n\n- Gini Index: Evaluates class purity reduction for categorical features.\n- Entropy: Information-theoretic score measuring class disorder reduction.\n- Fisher Score: Evaluates continuous numeric features by calculating the ratio of inter-class variance to intra-class variance:\n\n\text{Fisher Score} = \frac{\sum{k=1}^{K} n_k (\mu_k - \mu)^2}{\sum_{k=1}^{K} n_k \sigma_k^2}\n\nHigh Fisher scores indicate the feature means are well-separated across classes relative to internal variance.\n\n- Fisher's Linear Discriminant Analysis (LDA): Identifies optimal linear projections maximizing inter-class separation while minimizing intra-class spread.\n\n## Comparison: Filter vs Wrapper vs Embedded Feature Selection\n\n| Aspect | Filter Methods | Wrapper Methods | Embedded Methods |\n| :--- | :--- | :--- | :--- |\n| Feature Evaluation Metric | Statistical / Information metrics (Gini, Entropy, Fisher Score, Correlation) | Predictive accuracy of a target machine learning algorithm | Model regularisation penalties or internal feature importance metrics |\n| Model Dependence | Completely independent of ML algorithms | Dependent on the specific ML model used to score subsets | Built directly into model training |\n| Computational Cost | Extremely fast / Low cost | High — requires retraining models across numerous subsets | Moderate — executed during a single training run |\n| Overfitting Risk | Low — does not optimize against model prediction errors | High — prone to overfitting selected feature subsets to training data | Moderate — managed via regularisation parameters |\n| Common Examples | Chi-Square, Information Gain, Fisher Score, Pearson Correlation | Sequential Forward Selection, Backward Elimination, RFE | Lasso (L_1) Regression, Ridge Regression, Decision Tree Impurity |\n\n## Key Terms — Week 5 (Second Block)\n\n- Support Vector Machine (SVM): Maximum-margin separating hyperplane classifier.\n- Support Vectors: Training samples anchoring the margin boundaries.\n- Margin: Distance between separating hyperplane and support vectors.\n- Soft Margin (C\text{-parameter}): Formulation permitting controlled margin violations.\n- Kernel Trick: Implicit higher-dimensional feature mapping via inner-product kernel functions.\n- Feature Selection: Isolating informative feature subsets to eliminate noise and redundancy.\n- Filter Methods: Model-agnostic feature selection using statistical scoring metrics.\n- Wrapper Methods: Model-dependent feature selection evaluating iterative subset accuracy.\n- Embedded Methods: Feature selection built directly into algorithm parameter learning.\n- Fisher Score: Feature ratio evaluating inter-class separation relative to intra-class spread.\n\n## Self-Test Questions — Week 5 (Second Block)\n\n1. In your own words, explain what a support vector is and why it matters more than other training points.\n2. Why does maximising the margin tend to produce a classifier that generalises better to new data?\n3. What is the difference between a hard margin and a soft margin, and why is the soft margin more realistic for real-world data?\n4. Explain the intuition behind the kernel trick without using any equations.\n5. Name one strength and one weakness of SVMs.\n6. Explain the conceptual difference between filter, wrapper, and embedded feature selection methods.\n7. What does a low Gini index or low entropy value tell you about a feature? What does a high Fisher score tell you?\n8. Why is feature selection often performed before training a classification model?\n\n# Second Block — Week 6: Text Analytics\n\n## What is text analytics and why does text need special handling?\n\nText Analytics (Text Mining) applies natural language processing (NLP) and machine learning techniques to extract structured insights from unstructured text data (emails, customer reviews, social media posts, legal contracts).\n\nText data presents computational challenges:\n\n- Unstructured Format: Text lacks predefined tables or numerical fields.\n- High-Dimensional Sparsity: A text corpus lexicon may contain hundreds of thousands of unique words, but individual documents contain only small subsets, resulting in sparse data matrices.\n- Non-Negativity: Word counts are non-negative (\ge 0).\n- Sensitivity to Syntax and Semantics: Meaning depends on context, negation, word order, and idioms.\n\n## Text Pre-processing (Cleaning) Steps\n\nBefore mining, raw text passes through a sequential pre-processing pipeline:\n\n1. Language Identification: Identifying document language to select appropriate dictionaries.\n2. Tokenization: Segmenting text streams into discrete units (tokens), such as words or phrases.\n3. Lowercasing: Converting text to lowercase to merge identical tokens.\n4. Stop-Word Removal: Filtering out common, non-discriminative words (e.g., "the", "is", "at").\n5. Stripping Punctuation, Digits, & Markup: Removing HTML tags, special symbols, and numbers.\n6. Stemming vs Lemmatization:\n - Stemming: Crude, rule-based algorithm chopping off word suffixes (e.g., "running", "runs" \to "runn"). Fast, but produces non-words.\n - Lemmatization: Uses vocabulary dictionaries and part-of-speech context to reduce words to true morphological roots (lemmas) (e.g., "better" \to "good").\n7. Advanced Linguistic Processing: Part-of-Speech (POS) tagging, syntactic parsing, and named entity recognition (NER).\n\n## Representing Text Numerically: Bag-of-Words, Term Frequency, and TF-IDF\n\nMachine learning algorithms require numeric inputs:\n\n- Bag-of-Words (BoW) Model: Represents a document as an unordered collection of word occurrence counts, ignoring grammar and word order.\n- Term Frequency (TF): Raw count of word td\text{TF}(t, d).\n- Term Frequency - Inverse Document Frequency (TF-IDF): Weighting scheme that down-weights words occurring across many documents while highlighting locally frequent, distinctive terms:\n\n\text{TF-IDF}(t, d, D) = \text{TF}(t, d) \times \text{IDF}(t, D)\n\nWhere Inverse Document Frequency (IDF) is defined as:\n\n\text{IDF}(t, D) = \ln\left(\frac{|D|}{|{d \in D \mid t \in d}|}\right)\n\n - High TF-IDF score: Word appears frequently in document dD\n\n## The Vector Space Model and Document Similarity\n\nThe Vector Space Model represents documents as multi-dimensional vectors within a space defined by vocabulary terms. Document similarity is computed using Cosine Similarity, which measures the cosine of the angle between two document vectors \mathbf{A}\mathbf{B}, normalizing for document length:\n\n\text{Cosine Similarity}(\mathbf{A}, \mathbf{B}) = \frac{\mathbf{A} \cdot \mathbf{B}}{||\mathbf{A}|| ||\mathbf{B}||} = \frac{\sum_{i=1}^{d} A_i B_i}{\sqrt{\sum_{i=1}^{d} A_i^2} \sqrt{\sum_{i=1}^{d} B_i^2}}\n\n- Value = 1.0: Identical term distributions.\n- Value = 0.0: Completely orthogonal; no shared terms.\n\n## Topic Modelling\n\nTopic Modelling is an unsupervised technique that discovers latent thematic topics within text collections.\n\nLatent Dirichlet Allocation (LDA) (Blei et al., 2003) is a generative probabilistic topic model:\n\n- Document-Topic Mixture: Assumes every document is composed of a mixture of latent topics.\n- Topic-Word Mixture: Assumes every topic is defined by a probability distribution over vocabulary words.\n\nLDA infers latent topic structures by examining co-occurrence patterns across documents. Dirichlet distributions serve as priors controlling document-topic and topic-word sparsities.\n\n## Sentiment Analysis\n\nSentiment Analysis classifies emotional polarity (Positive, Negative, Neutral) expressed in text:\n\n- Lexicon-Based (Rule-Based): Computes sentiment by matching document words against pre-scored sentiment dictionaries (e.g., SentiWordNet). Simple, but struggles with context, negation, and domain shifts.\n- Machine-Learning-Based: Trains supervised classifiers (Naive Bayes, SVM, Neural Networks) on labelled sentiment text corpora.\n- Hybrid Approaches: Combines lexicon features with machine learning architectures.\n- Challenges: Sarcasm, negation handling ("not good"), multipolarity, and Aspect-Based Sentiment Analysis (attributing sentiment to specific product features).\n\n## Applications of Text Analytics\n\nSocial media brand monitoring, automated customer support routing, spam detection, legal e-discovery, financial document analysis, clinical record mining, and fake news detection.\n\n## Comparison: Bag-of-Words vs TF-IDF vs Topic Modelling (LDA)\n\n| Technique | What It Captures | Primary Limitations |\n| :--- | :--- | :--- |\n| Bag-of-Words (BoW) | Simple unweighted token occurrence frequencies per document | Treats all terms equally; common non-informative words dominate; discards syntax/order |\n| TF-IDF | Relative term importance balancing local frequency against global corpus rarity | Flat term weights; ignores semantic themes, synonyms, context, and word order |\n| Topic Modelling (LDA) | Hidden thematic topic distributions running across document collections | Requires tuning topic count K; topics require manual human interpretation |\n\n## Key Terms — Week 6 (Second Block)\n\n- Tokenization: Splitting raw text streams into discrete word tokens.\n- Stemming: Heuristic suffix stripping yielding root stems.\n- Lemmatization: Morphological reduction to dictionary base forms.\n- Bag-of-Words (BoW): Document representation as unordered word frequency vectors.\n- TF-IDF: Numerical weighting scheme penalizing globally common terms.\n- Vector Space Model: Geometrical document representation as vectors in vocabulary space.\n- Cosine Similarity: Metric measuring angle similarity between text vectors.\n- Latent Dirichlet Allocation (LDA): Probabilistic topic model identifying document-topic mixtures.\n- Sentiment Analysis: Classification of emotional tone and polarity in text.\n\n## Self-Test Questions — Week 6 (Second Block)\n\n1. Why can't standard structured-data mining techniques be applied directly to raw text, and what technical properties (sparsity, non-negativity) make text data different?\n2. Put the following pre-processing steps in a sensible order and explain what each does: stop-word removal, tokenization, lowercasing, stemming/lemmatization.\n3. Explain, without formulas, why a word that appears in almost every document in a corpus should get a low TF-IDF weight even if it appears frequently in one particular document.\n4. What is the vector space model, and why does it allow documents to be compared for similarity?\n5. In LDA, what does it mean to say a document is a "mixture of topics" and a topic is a "mixture of words"? Why is the technique called "latent"?\n6. Compare the rule-based (lexicon) and machine-learning approaches to sentiment analysis in terms of setup effort, scalability, and accuracy across domains.\n7. Give an example of a sentence that would challenge a sentiment analysis system because of sarcasm, negation, or multipolarity.\n8. List three real-world business applications of text analytics and explain the value each provides.\n\n# Second Block — Week 7: Time Series Analysis — ARIMA Models\n\n## What is a time series and why is it different?\n\nA Time Series is a sequence of quantitative observations recorded sequentially over uniform time intervals (e.g., daily stock prices, monthly sales).\n\nTime series data differs from standard tabular data because observations are sequential. Successive observations exhibit Autocorrelation (dependency on past values). Data order cannot be shuffled.\n\n## Trend, Seasonality, and Stationarity\n\nTime series decomposition separates series into components:\n\n- Trend (T_t): Long-term upward or downward movement over time.\n- Seasonality (S_t): Predictable, repeating patterns occurring at fixed calendar intervals (e.g., annual holiday spikes).\n- Irregular Noise (I_t): Unpredictable random fluctuations.\n\n### Stationarity\n\nA time series is Strictly Stationary if its joint probability distribution remains invariant over time shifts. In practice, models require Weak (Mean-Absolute) Stationarity:\n\n1. Constant Mean: E[Y_t] = \mut\n2. Constant Variance: \text{Var}(Y_t) = \sigma^2t\n3. Autocovariance depends strictly on the lag distance kt\n\nARIMA models require stationarity because fitting stationary parameter relationships assumes historical statistical properties will remain stable in future forecast horizons.\n\n### Differencing\n\nNon-stationary series containing trends or changing means are transformed into stationary series via Differencing: subtracting previous values from current values:\n\n\Delta Y_t = Y_t - Y_{t-1}\n\nFirst-order differencing removes linear trends. If non-stationarity persists, second-order differencing \Delta^2 Y_t = \Delta Y_t - \Delta Y_{t-1} is applied.\n\nStationarity is formally evaluated using the Augmented Dickey-Fuller (ADF) Test:\n\n- Null Hypothesis (H_0): Series is non-stationary (contains a unit root).\n- Alternate Hypothesis (H_1): Series is stationary.\n- Rejecting H_0p < 0.05) confirms stationarity.\n\n## Correlation, Lag, and the ACF/PACF Plots\n\n- Lag (kY_tY_{t-k}.\n- Autocorrelation Function (ACF): Measures total linear correlation between Y_tY_{t-k}, incorporating both direct and indirect intermediate effects.\n- Partial Autocorrelation Function (PACF): Measures direct linear correlation between Y_tY_{t-k}Y_{t-1} \dots Y_{t-k+1}).\n\nACF and PACF plots guide parameter selection for ARIMA models.\n\n## What ARIMA Stands for and What Each Part Means\n\n\text{ARIMA}(p, d, q) combines three modeling components:\n\n- AR (AutoRegressive) — Parameter pY_tp past actual values:\n\nY_t = c + \phi_1 Y_{t-1} + \phi_2 Y_{t-2} + \dots + \phi_p Y_{t-p} + \epsilon_t\n\n- I (Integrated) — Parameter d: Specifies the degree of differencing required to achieve stationarity.\n- MA (Moving Average) — Parameter qY_tq past forecast residual errors (shocks):\n\nY_t = c + \epsilon_t + \theta_1 \epsilon_{t-1} + \theta_2 \epsilon_{t-2} + \dots + \theta_q \epsilon_{t-q}\n\nFull \text{ARIMA}(p, d, q)Y't:\n\nY'_t = c + \sum{i=1}^{p} \phi_i Y'{t-i} + \epsilon_t + \sum{j=1}^{q} \theta_j \epsilon_{t-j}\n\n### Special Case Formulations\n\n- \text{ARIMA}(p, 0, 0) = \text{Pure } \text{AR}(p)\n- \text{ARIMA}(0, 0, q) = \text{Pure } \text{MA}(q)\n- \text{ARIMA}(p, 0, q) = \text{ARMA}(p, q)\n\n## The General ARIMA Workflow\n\n1. Plot time series data to inspect for trends and seasonality.\n2. Execute ADF tests; apply differencing (d) until series achieves stationarity.\n3. Inspect ACF and PACF plots of stationary series to identify candidate orders for pq\n - PACF cuts off after lag p \to \text{AR}(p)\n - ACF cuts off after lag q \to \text{MA}(q)\n4. Estimate model parameters \phi_i\theta_j using Maximum Likelihood Estimation.\n5. Evaluate residual diagnostics (check if residuals behave as white noise).\n6. Forecast future values and invert differencing transformations to recover baseline scales.\n\n## Why Time Series Forecasting Matters\n\nForecasting enables proactive decision-making in retail demand planning, financial trading, macroeconomic analysis, energy grid provisioning, and capacity management.\n\n## Comparison: AR vs I vs MA Components of ARIMA\n\n| Component | What It Models | Parameter & Meaning | Diagnostic Identification |\n| :--- | :--- | :--- | :--- |\n| AR (AutoRegressive) | Current value predicted from past actual values | pp\n| I (Integrated) | Differencing applied to eliminate trends | d — degree of differencing needed to achieve stationarity | ADF test results / visual removal of trend |\n| MA (Moving Average) | Current value predicted from past error shocks | qq\n\n## Key Terms — Week 7 (Second Block)\n\n- Time Series: Chronologically ordered sequences of quantitative observations.\n- Stationarity: Statistical state featuring constant mean and constant variance over time.\n- Differencing: Subtracting past values to eliminate non-stationary trends.\n- Augmented Dickey-Fuller (ADF) Test: Statistical test evaluating unit-root non-stationarity.\n- Autocorrelation (ACF): Total correlation between series values across lag distances.\n- Partial Autocorrelation (PACF): Direct correlation between series values controlling for intermediate lags.\n- ARIMA(p, d, q): Integrated AutoRegressive Moving Average forecasting model.\n\n## Self-Test Questions — Week 7 (Second Block)\n\n1. Why does the order of observations matter for time series data in a way that it typically does not for standard tabular datasets?\n2. Define stationarity in plain terms, and explain why ARIMA models require it before they can be reliably fitted.\n3. What operation is used to turn a non-stationary series into a stationary one, and what does the parameter "d" represent?\n4. Explain the difference between autocorrelation and partial autocorrelation using the "Friday/Saturday/Sunday petrol price" style example.\n5. What do the ACF and PACF plots help you determine when building an ARIMA model?\n6. Spell out what each letter in ARIMA stands for and briefly describe what each component contributes to the model.\n7. Describe, as a workflow (not code), the steps you would follow to go from a raw non-stationary time series to a set of forecasted future values using ARIMA.\n8. Give two real-world examples of why organisations would want to forecast a time series, and explain what decisions the forecast could support.\n\n# Glossary of Key Terms\n\n- ACF plot (Autocorrelation Function plot): Graph of autocorrelation against increasing lag, used to help choose model order.\n- ACID properties: Atomicity, Consistency, Isolation, Durability — guarantees for reliable transactional processing.\n- ACID vs BASE: Strict transactional consistency vs. available, eventually-consistent behaviour.\n- Activation function: The function applied to the weighted sum at a node to produce its output; introduces non-linearity in multilayer networks.\n- AR (AutoRegressive) model: Forecasts using a weighted combination of past actual values.\n- ARIMA(p, d, q) model: ARMA model applied after differencing a non-stationary series d times, with p AR terms and q MA terms.\n- ARMA model: Combination of AR and MA components for an already-stationary series.\n- Aspect-based sentiment analysis: Attributing sentiment to specific features/aspects mentioned in text rather than the text as a whole.\n- Association pattern mining: Discovering frequently co-occurring items/attributes in data.\n- Association rule (A \Rightarrow B$$): An implication between two disjoint itemsets satisfying min_sup and min_conf.
Autocorrelation: Correlation between a series and a lagged (shifted) version of itself, including indirect effects.
Backpropagation: The algorithm that sends prediction error backward through the network to update weights layer by layer.
Bag-of-Words (BoW): Representation of a document as an unordered collection of word counts.
Bayes' theorem: A formula for updating the probability of a hypothesis (class) given observed evidence (features), using prior probability and likelihood.
Bias: An extra adjustable value added to the weighted sum, shifting the decision threshold.
Binary/multinomial/ordinal logistic regression: Variants for two classes, unordered 3+ classes, and ordered 3+ classes respectively.
Candidate generation (join step): Combining frequent k-itemsets to build candidate (k+1)-itemsets.
CART (Classification And Regression Trees): A decision tree algorithm that recursively splits data to build trees for both categorical and numeric outcomes.
Centroid: The central (mean) point representing a cluster in k-Means.
Classification: Supervised prediction of a class label for new records from a trained model.
Cluster analysis / clustering: Grouping similar data points without predefined labels.
Clustering: Unsupervised grouping of similar records.
Conditional independence assumption: The "naive" assumption that features do not influence one another once the class is known.
Confidence: Conditional probability that Y is present given X is present, sup(X U Y) / sup(X).
Confusion matrix: Table of TP, TN, FP, FN counts for a classifier's predictions.
Convergence: The point at which cluster assignments/centroids stop changing meaningfully.
Corpus: A collection of text documents used for analysis.
Cosine similarity: A measure of similarity between two document vectors based on the angle between them (robust to differing document lengths).
Curse of Dimensionality: The exponential increase in data/computation required as the number of features grows, degrading model performance and efficiency.
Data Cleaning: The process of detecting/correcting missing, erroneous, or inconsistent data entries.
Data Integration: Combining data from multiple heterogeneous sources into one coherent store.
Data lake: A repository storing raw data of any type/format without a predefined schema.
Data mart: A smaller, department- or subject-focused subset of a data warehouse.
Data mining: Extracting previously unknown, useful patterns/knowledge from large volumes of data.
Data Reduction: Producing a smaller representation of a dataset that preserves most of its analytical value.
Data warehouse: A subject-oriented, integrated, time-variant, non-volatile collection of data supporting decision-making.
Decision boundary: The line/surface separating predicted classes.
Decision Support System (DSS): An interactive, computer-based system that helps decision-makers use data, documents, knowledge, and models to solve semi-structured/unstructured problems.
Descriptive model: A model that describes/simulates how a situation behaves, without prescribing the "best" answer.
Dickey-Fuller / Augmented Dickey-Fuller (ADF) test: Statistical hypothesis test used to check whether a series is stationary.
Differencing: Subtracting each value from its predecessor to remove trend and help achieve stationarity.
Discretisation/Binning: Converting continuous numeric values into a smaller set of categorical ranges.
Document, key-value, wide-column, graph databases: The major NoSQL data models.
Downward closure / Apriori principle: Every subset of a frequent itemset is also frequent (equivalently, supersets of infrequent itemsets are infrequent).
Dunn Index: Ratio of minimum intercluster distance to maximum intracluster distance.
Elbow method: Choosing k at the point where the inertia curve's rate of decrease sharply flattens.
Embedded method: Feature selection built into the model's own training process.
Entity-relationship modelling: Identifying real-world entities and their relationships before designing tables.
Entropy: An information-theoretic measure of class "mixing" for a feature; lower is better.
Entropy / information gain: An impurity measure from information theory; information gain is the reduction in entropy achieved by a split.
Epoch: One complete pass through the entire training data set.
ETL: Extract, Transform, Load — the process of moving data from sources into a warehouse in a clean, consistent form.
ETL / ELT: Extract-Transform-Load vs. Extract-Load-Transform data integration pipelines, differing in when transformation happens relative to loading.
Explained Variance: The proportion of a dataset's total variance captured by a given (set of) principal component(s).
Exploratory Data Analysis (EDA): The process of analysing and visualising data to understand its structure, patterns, and quality before formal modelling, without prior assumptions about the data-generating population.
Feature selection: The process of identifying and keeping only the most informative features for a classification task.
Filter method: Feature selection based on statistical measures, independent of the model.
Fisher score: Ratio of between-class separation to within-class separation for a feature; higher is better.
Fisher's linear discriminant: A generalisation of the Fisher score that finds the best combined direction (of multiple features) for separating classes.
Forecasting: Predicting future values of a series based on its historical pattern.
Foreign key: Links a row to a row in another table, forming a relationship.
Frequent (large) itemset: An itemset meeting or exceeding min_sup.
Gaussian / Multinomial / Bernoulli Naive Bayes: Variants suited to continuous, count-based, and binary features respectively.
Gini index: CART's impurity measure for classification splits; 0 = pure node, higher = more mixed classes.
Gorry & Scott-Morton framework: 3x3 matrix crossing decision structuredness with managerial control level (operational, management, strategic).
Grain, dimension, measure: Data warehouse design concepts: level of detail, analytical angle, numeric fact.
Group DSS (GSS): A DSS that supports multiple decision-makers collaborating on a shared problem.
Hard margin: A margin that assumes classes are perfectly linearly separable, with no violations allowed.
Hidden layer: Layer(s) of nodes between input and output that let a network learn non-linear patterns.
Holdout, cross-validation, bootstrap: Methodologies for splitting labelled data to obtain a fair estimate of classifier accuracy.
Homoscedasticity: Constant variance of residuals across the range of predictors.
Hyperplane: The flat decision boundary (a line in 2D, a plane in 3D, etc.) that separates classes in an SVM.
Imputation: Estimating a plausible value to fill in a missing data entry.
Inertia: Sum of squared distances of points from their cluster centroid (compactness measure).
Institutional vs ad hoc DSS: Recurring, refined-over-time decisions vs one-off, unanticipated decisions.
Internal validation: Evaluating clustering quality using only the data's own structure, with no external labels.
Interquartile Range (IQR): The spread of the middle 50% of data (Q3 - Q1); robust to outliers.
Inverse Document Frequency (IDF): A measure that down-weights words that occur in many documents across the corpus.
Itemset: A set of one or more items; a k-itemset has exactly k items.
k-Medians: Variant using median representatives, more robust to outliers.
k-Medoids: Variant using actual data points as representatives, suited to varied data types.
Kernel trick: The technique of implicitly mapping data to a higher-dimensional space (via a similarity/kernel function) so that a non-linear problem becomes linearly separable there.
Lag: The time gap between an observation and an earlier observation being compared to it.
Laplacian smoothing: A technique to avoid zero-probability estimates for unseen feature/class combinations.
Latent Dirichlet Allocation (LDA): A topic-modelling technique where documents are mixtures of topics and topics are mixtures of words.
Learning rate: A parameter controlling how large each weight adjustment is.
Least squares: The method of choosing coefficients that minimise the sum of squared prediction errors.
Lemmatization: Dictionary/linguistics-based reduction of a word to its proper base form (lemma).
Lexicon: The full set of distinct words (vocabulary) used across a corpus.
Lexicon-based (rule-based) sentiment analysis: Scoring sentiment using dictionaries of pre-scored words.
Lift / Conviction: Interestingness measures that correct for the base popularity of items, revealing genuine (rather than coincidental) associations.
Linearly separable: Data that can be divided into classes by a straight line/flat plane.
Linearly separable data: Data that can be perfectly divided by a straight line/hyperplane.
Logistic regression: A discriminative, probabilistic classification model that maps a linear combination of features through a sigmoid function to produce a class probability.
MA (Moving Average) model: Forecasts using a weighted combination of past forecast errors.
Machine learning: Algorithms that learn patterns from data to make predictions without explicit task-specific programming.
Margin: The distance/gap between the separating hyperplane and the closest training points of each class.
Market Basket Analysis (MBA): The classic retail application of association mining to co-purchased items.
Matplotlib/Seaborn: Python libraries for creating charts and statistical visualisations.
Mean: The arithmetic average; sensitive to outliers.
Median: The middle value of sorted data; robust to outliers and skew.
Minimum confidence (min_conf): Threshold for accepting a rule as "strong."
Minimum support (min_sup): The threshold below which an itemset is discarded as infrequent.
Mode: The most frequently occurring value; the only central-tendency measure valid for nominal data.
Multilayer feed-forward network / multilayer perceptron: Network with input, hidden, and output layers, connections flowing forward.
Multivariate/multiple linear regression: More than one predictor combined linearly.
Named Entity Recognition: Identifying names of people, places, organisations, etc. within text.
Neuron/node: Basic computing unit of a neural network, inspired by biological neurons.
Nonlinear regression: Modelling curved (non-straight-line) relationships.
Normalisation: Rescaling attribute values onto a common, bounded scale.
Normalisation/Standardisation: Rescaling numeric attributes so they are comparable in scale.
Normative model: A model that identifies the demonstrably best alternative (optimisation).
NoSQL: Non-relational databases designed for flexibility and horizontal scale.
NumPy: Python library providing fast array-based numerical computation.
OLAP: Online Analytical Processing — interactive, multidimensional exploration (drill-down, slice/dice) of warehouse data.
OLTP: Online transaction processing; fast day-to-day operational recording of current data.
Outlier: A data point that deviates significantly from the rest of the dataset; may be an error or a genuine anomaly.
Outlier detection: Identifying records that deviate significantly from the rest of the data.
Overfitting: A model that fits training data (including its noise) too closely and generalises poorly to new data.
PACF plot (Partial Autocorrelation Function plot): Graph of partial autocorrelation against increasing lag, used to help choose model order (particularly useful for identifying p).
pandas: Python library providing the DataFrame structure for loading, cleaning, and manipulating tabular data.
Parametric vs non-parametric nonlinear regression: Assuming a specific functional form vs learning the shape from data.
Partial autocorrelation: Correlation between a series and a specific lag of itself, after removing the effect already explained by shorter lags (direct effect only).
Perceptron: The simplest neural network: input nodes plus one output node, learning only linear boundaries.
Polynomial degree: The highest power used; controls the trade-off between underfitting and overfitting.
Polynomial regression: Regression using powers of the predictor(s) as extra features to fit curves; linear in its coefficients.
Posterior probability: The updated probability of a class after observing the evidence.
Precision, Recall, F1-score, Accuracy: Quantitative metrics computed from the confusion matrix.
Predictor/independent variable (regressor): An input feature used to predict the outcome.
Primary key: Uniquely identifies a row within its table.
Principal Component: A new variable representing a weighted combination of original attributes, capturing one direction of maximum (remaining) variance in the data.
Principal Component Analysis (PCA): An unsupervised technique that reduces dimensionality by projecting data onto new, uncorrelated axes ordered by the amount of variance they capture.
Prior probability: The probability of a class before observing any evidence.
Pruning: Removing tree nodes/branches that do not improve (or worsen) performance on held-out validation data, to reduce overfitting.
Pruning step: Removing candidates that must be infrequent because one of their subsets is not frequent.
Recursive partitioning: The top-down, greedy process of repeatedly splitting nodes to build a tree.
Redundancy: An attribute that is derivable from other attributes, or duplicated due to inconsistent naming across sources.
Regression analysis: Predicting a continuous/numeric outcome from one or more input variables.
Regularisation (ridge/Lasso): Penalising large coefficients to reduce overfitting.
Representative-based (partitional) clustering: Clustering by assigning points to the nearest of k representative points.
Response/dependent variable (regressand): The numeric outcome being predicted.
ROC curve / AUC: A threshold-independent way to visualise and quantify a classifier's ability to distinguish classes.
ROLAP / MOLAP / HOLAP: Relational, multi-dimensional, and hybrid approaches to implementing OLAP.
Roll-up / drill-down / slice / dice: Core OLAP operations for aggregating, detailing, and viewing data.
Root node, decision/internal node, leaf node, branch: The structural components of a decision tree.
Rule-based classifier: A classifier expressed as a set of if-then rules, often derived from a decision tree.
Sampling (with/without replacement, stratified, biased): Selecting a representative subset of records for analysis.
Satisficing: Choosing a "good enough" solution rather than the mathematically optimal one.
Schema/entity matching: Determining whether differently-named attributes/records from different sources refer to the same real-world thing.
scikit-learn: Python library providing ready-made machine learning/data mining algorithms and preprocessing tools.
Seasonality: A repeating pattern occurring at fixed, regular time intervals.
Semi-structured decision: Partly standardisable, partly requiring human judgement.
Sentiment Analysis: Classifying text as expressing positive, negative, or neutral opinion (or a specific emotion).
Sigmoid/logistic function: A function that compresses any real number into the range (0,1), used to convert a linear score into a probability.
Silhouette score: Measure (-1 to 1) of how well a point fits its own cluster versus neighbouring clusters.
Simple linear regression: One predictor, straight-line relationship.
Skewness: A measure of the asymmetry of a distribution's shape.
Soft margin: A margin that tolerates some points crossing into/past the margin, penalised accordingly.
Split criterion: The rule used at each node to divide data (e.g. Age <= 50).
SQL: Standard language for querying/managing relational data.
Standardisation: Rescaling attribute values using the mean and standard deviation (z-scores).
Stationarity: Constant mean, constant variance, and no seasonality over time.
Stemming: Crude rule-based truncation of a word to an approximate root form.
Stop words: Common, low-information words (e.g. "the," "and") typically removed during preprocessing.
Structured decision: Routine, repetitive problem with a known standard solution procedure.
Support: The fraction/count of transactions containing a given itemset.
Support vectors: The training data points closest to the hyperplane that define/anchor the margin.
Table / row / column: Structural building blocks of a relational database.
Term Frequency (TF): How often a word occurs within one document.
TF-IDF: Combined weighting scheme that highlights words that are frequent locally but rare globally.
Time series: A sequence of observations recorded in time order, where the order/timing carries meaning.
Tokenization: Splitting text into individual word/sub-word units.
Topic Modelling: Unsupervised discovery of hidden themes across a document collection.
Transaction / TID: A single record (e.g. a shopping basket) identified by a unique transaction ID.
Trend: A long-term upward or downward movement in a series.
Underfitting: Model is too simple to capture the true pattern in the data.
Unstructured decision: Complex, fuzzy problem with no predefined solution procedure; relies on judgement.
Unsupervised learning: Learning patterns from data with no target/output variable.
Variance reduction: The regression-tree analogue of Gini index, used when the target is numeric.
Variance/Standard deviation: Measures of how spread out data values are around the mean.
Vector Space Model: Representing each document as a point/vector in a high-dimensional space defined by vocabulary words, enabling similarity comparisons.
Weight: A value on a connection between nodes representing the strength/importance of that connection; adjusted during learning.
Wrapper method: Feature selection based on retraining a specific model on different feature subsets.