Introductory Background to Data Analytics

Course Introduction and Personnel

  • Subject Matter Expert: Said Achmad (D6667).
  • Course Code: COMP6886001-Data Analytics.
  • Session: Session 1 - Introductory Background.
  • Learning Outcome (LO 1): Success is defined by the ability to describe the data analytics process.

Foundational Definitions of Data and Analytics

  • Analytics: The science focused on analyzing crude data to extract useful knowledge, specifically patterns. The scope of this process includes:
    • Data collection.
    • Organization.
    • Pre-processing.
    • Transformation.
    • Modeling.
    • Interpretation.
  • Data: In the context of the information age, data refers to a large set of bits encoding numbers, texts, images, sounds, videos, and other formats.
  • Knowledge: Knowledge is created when information is added to data, providing it with meaning.
  • Data Entities:
    • Instance or Object: Each individual member within a data set.
    • Attributes or Features: The specific characteristics that describe the instances.
  • Relational Data Sets: These are data sets represented by several tables where the relations between the tables are explicitly defined. For example, if Table 1.1 contains contacts and Table 1.2 represents family relationships between those contacts, the distinct tables remain connected by shared individuals. These structures are managed using relational databases.

Taxonomy of Data Analytics

  • Descriptive Analytics: The process of summarizing or condensing data to extract existing patterns.
    • Results: Can include statistics such as average, mean, and median, as well as plots or sets of groups containing similar instances.
  • Predictive Analytics: The process of extracting models from data to be used for future predictions.
    • Model Induction: Considered a predictive task.
  • Procedural Definitions:
    • Method or Technique: A systematic procedure used to achieve an intended goal. For example, a method to obtain the average age of contacts across different ages.
    • Algorithm: A self-contained, step-by-step set of instructions that are easily understandable by humans, allowing for the implementation of a specific method.
  • Model: A generalization obtained from data that serves as a prototype for generating predictions for new, previously unseen instances.

The Data Analytics Project Lifecycle

  • Project Phases:
    1. Understanding the problem to be solved.
    2. Defining the objectives of the project.
    3. Looking for the necessary data.
    4. Preparing the data so it is usable.
    5. Identifying suitable methods and choosing between them.
    6. Tuning the hyper-parameters of each chosen method.
    7. Analyzing and evaluating the results.
    8. Redoing pre-processing tasks and repeating experiments as necessary.

Methodologies for Data Analytics: KDD and CRISP-DM

The Knowledge Discovery in Databases (KDD) Process

The KDD process consists of a sequence of nine specific steps:

  1. Learning the application domain.
  2. Creating a target dataset.
  3. Data cleaning and pre-processing.
  4. Data reduction and projection.
  5. Choosing the data mining function.
  6. Choosing the data mining algorithm.
  7. Data mining.
  8. Interpretation.
  9. Using discovered knowledge.
The CRISP-DM Methodology

The CRoss Industry Standard Process for Data Mining (CRISP-DM) includes six iterative phases:

  1. Business Understanding: Understanding the business domain, defining the problem from a business perspective, and translating business problems into data analytics problems.
  2. Data Understanding: Collecting necessary data and performing initial visualization/summarization to gain insights, specifically identifying data quality issues like missing data or outliers.
  3. Data Preparation: Preparing the dataset for the modeling tool. This includes:
    • Data transformation.
    • Feature construction.
    • Outlier removal.
    • Missing data fulfillment.
    • Removal of incomplete instances.
  4. Modeling: Selecting and applying various methods. This phase may require returning to Data Preparation if a method has specific requirements. It includes tuning hyper-parameters.
  5. Evaluation: Ensuring the solution answers the business requirements and is meaningful from a business perspective.
  6. Deployment: Integrating the data analytics solution into the business process, such as decision-support tools, websites, or reporting processes.

Big Data and Data Science

  • Big Data Definition (The Three Vs):
    • Volume: The sheer scale of data.
    • Variety: The different formats and types of data.
    • Velocity: The speed at which data is generated and processed.
  • Distinction:
    • Big Data: Primarily concerned with technology and handling data that cannot be effectively processed with traditional methods.
    • Data Science: Concerned with the creation of models to extract patterns from complex data and applying those models to real-life problems.

Big Data Architectures and MapReduce

  • Technological Necessity: As volume, velocity, and variety increase, clusters of computers become necessary.
  • Hadoop and MapReduce: MapReduce was the first technique developed for processing big data using clusters.
  • Scalability Example: To calculate the average salary of 1×1091 \times 10^9 (1 billion) people using a cluster of 1,0001,000 computers:
    • Divide the population into 1,0001,000 chunks.
    • Each chunk contains data for 1,000,0001,000,000 (1×1061 \times 10^6) people.
    • Each computer processes one chunk.
  • Distributed System Requirements:
    • Fault Tolerance: Ensure no data chunk is lost. If a computer fails, its task must be assumed by another.
    • Redundancy: Repeating the same task and data chunk across more than one computer so the redundant computer can take over during failure.
    • Recovery: Computers with faults must be able to return to the cluster once fixed.
    • Elasticity: Computers can be easily removed or added to the cluster based on demand.
MapReduce Word Count Flow
  • Input: Raw text data (e.g., "Deer Bear River", "Car Car River", "Deer Car Bear").
  • Splitting: Breaking input into manageable parts.
  • Mapping: Assigning a value of 11 to each word (e.g., Bear, 1; River, 1).
  • Shuffling: Grouping identical keys together.
  • Reducing: Summing the values for each key (e.g., Car, 3; Deer, 2; Bear, 2; River, 2).
  • Final Result: The aggregated counts for each unique instance.

Small Data

  • Source: Produced continuously by individuals through daily activities like web navigation, shopping, medical exams, and mobile app usage.
  • Definition: A data set whose volume and format allow processing and analysis by a single person or a small organization.
  • Operational Focus:
    • Big Data: Focused on helping companies understand customers, products, and services.
    • Small Data: Focused on helping individuals understand themselves.

Summary of Key Concepts

  • Data: Raw facts lacking meaning.
  • Value: Data gains value when processed and analyzed for insights.
  • Analytics Process: Follows the sequence of Business understanding, Data understanding, Data preparation, Modeling, Evaluation, and Deployment.
  • Big Data: Large volumes requiring non-traditional processing methods.
  • Data Science: Scientific methods and algorithms used to extract knowledge.

Exercises and Reference Material

Self-Assessment Exercises
  1. Define the term "data" and explain why it is considered a valuable asset in various industries.
  2. Provide three examples of different types of data and explain how each type can be used in a business context.
  3. Choose a specific industry (e.g., healthcare, finance, retail) and describe how data analytics can be applied to make informed decisions.
References
  1. Moreira, João & de Carvalho, Andre & Horvath, Tomas. (2018). A General Introduction to Data Analytics. DOI: 10.1002/9781119296294.
  2. Pereira, Óscar & Capitão, Micael. (2014). Mediator Framework for Inserting Data into Hadoop.