Introduction to Data Science

Overview of Data Science

  • Data Science is a multidisciplinary field that utilizes scientific methods, processes, algorithms, and systems to extract knowledge and actionable insights from data in all formats (structured, semi-structured, and unstructured data).
  • The core focus of data science is extracting valuable knowledge from complex datasets.
  • It continues to evolve as one of the most promising, in-demand career paths for skilled professionals.
  • Data science intersects multiple academic and technical domains:
    • Computer Science / IT (Programming skills)
    • Mathematics
    • Statistics
    • Machine Learning
    • Information Science
    • Data Mining

Data Science Venn Diagram

  • Intersections within the Data Science Domain:

    • Machine Learning: Formed by the combination of Computer Science/IT and Math and Statistics.
    • Traditional Research: Formed by the combination of Math and Statistics and Domains/Business Knowledge.
    • Software Development: Formed by the combination of Computer Science/IT and Domains/Business Knowledge.
    • Data Science: Formed at the central core where Computer Science/IT, Math and Statistics, and Domains/Business Knowledge all overlap.
  • Roles and Competencies of Data Scientists:

    • Data scientists master the full spectrum of the data science life cycle to uncover useful intelligence for organizational decision-making.
    • They must be curious, detail-oriented, and result-oriented.
    • They require strong communication skills to articulate complex technical findings to non-technical stakeholders.
    • They need a solid quantitative foundation in statistics and linear algebra.
    • They must possess deep programming knowledge focused on data warehousing, data mining, and data modeling to construct and analyze algorithms.

Data versus Information

  • Definition of Data:

    • Unprocessed or raw facts and figures.
    • A collection of facts, concepts, or instructions represented in a formalized manner.
    • Must be processed or interpreted by humans or electronic systems to convey true meaning.
    • Can be represented in the form of:
      • Alphabets (e.g., A–ZA\text{--}Z, a–za\text{--}z)
      • Digits (e.g., 0–90\text{--}9)
      • Special characters (e.g., ++, −-, //, ∗*, <<, >>, ==)
  • Definition of Information:

    • Processed data used to make informed decisions and drive action.
    • Data transformed into a meaningful format that carries real or perceived value for current or prospective decisions.
    • Interpreted data created from organized, structured, and processed data in a specific context.
  • Data vs. Information Comparison:

    • Data: Described as unprocessed or raw facts and figures | Information: Described as processed data.
    • Data: Cannot directly assist in decision-making | Information: Helps directly in decision-making.
    • Data: Serves as raw material that can be organized, structured, and interpreted | Information: Created from organized, structured, and processed data within a particular context.
    • Data: A collection of text, images, and voice representing quantities, actions, and objects | Information: Processed representation of text, images, and voice conveying context and actionable output.

Data Processing Cycle

  • Data Processing Definition: The restructuring or reordering of raw data by human effort or automated machinery to increase its utility and derive specific value.
  • Three Core Steps in the Data Processing Cycle:
    1. Input Step: Input data is prepared and structured into a suitable format depending on the processing machine.
    2. Processing Step: Transformation activities convert input data into a higher-value, processed form.
    3. Output Step: Processing results are collected and made available as output (information).

Data Processing Cycle

  • Concrete Example of Data Processing:
    • Input (Data): Raw numeric measurements 3535, 3434, 3333, 3232, 3636 alongside raw categorical days: Monday, Tuesday, Wednesday, Thursday, Friday.
    • Processing: Arranging, sorting, combining, and applying mathematical operations.
    • Output (Information): Contextualized weather readings:
      • Monday: 35 °C35\,\text{°C}
      • Tuesday: 34 °C34\,\text{°C}
      • Wednesday: 33 °C33\,\text{°C}
      • Thursday: 32 °C32\,\text{°C}
      • Friday: 36 °C36\,\text{°C}

Data Processing Example

Data Types and Perspectives

  • Computer Programming Perspective:

    • A data type is an attribute assigned to data that instructs compilers or interpreters on how the programmer intends to utilize the data.
    • Integer (int): Used to store whole numbers.
    • Boolean (bool): Restricted to binary logical values: true or false.
    • Character (char): Used to store a single symbol or character.
    • Floating-Point Number (float): Used to store real numbers containing fractional/decimal values.
    • Alphanumeric String (string): Used to store sequences of letters, numbers, and symbols.
  • Data Analytics Perspective:

    • Data Analytics is the science of evaluating raw data to extract conclusions.
    • Categorized into three primary data structures:
      1. Structured Data:
        • Adheres strictly to a pre-defined data model and tabular structure with defined row-column relationships.
        • Easily indexed, queried, and analyzed.
        • Examples: SQL relational databases, Excel spreadsheets.
        • Represents approximately 20%20\% of enterprise data according to Gartner estimates.
        • Requires minimal storage capacity and is simple to manage with legacy solutions.
      2. Unstructured Data:
        • Lacks a pre-defined data model and structural organization.
        • Typically text-heavy, containing embedded dates, numbers, facts, or multimedia.
        • Difficult to process using traditional relational database software.
        • Examples: Audio files, video files, PDF documents, Word documents, NoSQL stores.
        • Represents approximately 80%80\% of enterprise data according to Gartner estimates.
        • Requires larger storage systems and specialized management tools.
      3. Semi-Structured Data:
        • Does not follow a strict relational table structure but contains semantic organizational elements.
        • Uses internal tags or markers to separate data elements (self-describing structure).
        • Examples: XML files, JSON files.

Structured Data vs Unstructured Data

  • Comparison Example Across Data Types:

Data Format Comparison

*   **Unstructured Text:** "The university has 56005600 students. John's ID is number 11, he is 1818 years old and already holds a B.Sc. degree. David's ID is number 22, he is 3131 years old and holds a Ph.D. degree. Robert's ID is number 33, he is 5151 years old and also holds the same degree as David, a Ph.D. degree."
*   **Semi-Structured XML:**
    ```xml}
    <University>
      <Student ID="1">
        <Name>John</Name>
        <Age>18</Age>
        <Degree>B.Sc.</Degree>
      </Student>
      <Student ID="2">
        <Name>David</Name>
        <Age>31</Age>
        <Degree>Ph.D.</Degree>
      </Student>
    </University>

        `` * **Structured Table:** * Columns:ID,Name,Age,Degree` * Row 11: 11, John, 1818, B.Sc. * Row 22: 22, David, 3131, Ph.D. * Row 33: 33, Robert, 5151, Ph.D. * Row 44: 44, Rick, 2626, M.Sc. * Row 55: 55, Michael, 1919, B.Sc.

  • Metadata – Data About Data:
    • Provides supplementary descriptive information about specific datasets.
    • Not a separate structural format technically, but essential for Big Data solutions and indexing.
    • Example: Metadata associated with photograph files specifies creation timestamps and precise geotagged locations.
    • Individual metadata fields (such as timestamps and geographic coordinates) are structured, allowing Big Data platforms to leverage them for rapid initial filtering and analysis.

Big Data Value Chain

  • The Big Data Value Chain (DVC) models the end-to-end information flow within big data architectures to extract actionable business value.

Data Value Chain

  • Five Key Phases of the Big Data Value Chain:

    1. Data Acquisition:

      • Process of gathering, filtering, and cleaning data before loading into data warehouses or storage platforms.
      • Imposes major infrastructure challenges requiring minimal latency, high transaction throughput, and dynamic structural handling.
      • Includes: Structured and unstructured sources, event processing, sensor networks, protocols, real-time data streams, and multimodality.
    2. Data Analysis:

      • Transforming raw acquired data into formats actionable for decision-making and business domain application.
      • Involves exploratory analysis, data transformation, modeling, synthesis, and extraction of high-potential business insights.
      • Includes: Stream mining, semantic analysis, machine learning, information extraction, linked data, data discovery, whole-world semantics, cross-sectorial analysis, and ecosystem analytics.
    3. Data Curation:

      • Active management throughout the data life cycle to guarantee quality meets application requirements.
      • Includes content creation, selection, classification, transformation, validation, and long-term preservation.
      • Executed by data curators, scientific curators, and annotators to elevate quality, trust, discoverability, accessibility, and reusability.
      • Includes: Provenance tracking, annotation, human-data interaction, community/crowd curation, human computation, incentivisation, and system interoperability.
    4. Data Storage:

      • Scalable persistence and management of data to deliver rapid access for downstream systems.
      • Transitions from traditional Relational Database Management Systems (RDBMS) to Data Lakes.
      • Data Lakes support diverse formats built on Hadoop clusters, cloud object storage, or NoSQL databases.
      • Includes: In-memory databases, NoSQL databases, NewSQL databases, cloud infrastructure, query interfaces, CAP theorem trade-offs (Consistency, Availability, Partition-tolerance), security, privacy, and standardization.
    5. Data Usage:

      • Integration of analyzed data directly into operational business workflows and decision support systems.
      • Drives competitiveness by reducing overhead, unlocking new service value, and optimizing key operational metrics.
      • Includes: Decision support systems, predictive analytics, in-use analytics, simulations, interactive visualizations, modeling, and automated control.

Concepts of Big Data

  • Definition of Big Data:

    • Refers to massive, complex datasets that surpass the processing, storage, and management capacities of traditional relational databases or single-node computing systems.
    • The threshold defining a "large dataset" is relative, shifting continually over time and varying across individual organizational capabilities.
  • The Four V's of Big Data:

Characteristics of Big Data

1.  **Volume (Data at Rest):**
    *   Massive scale of collected datasets.
    *   Ranges from Terabytes (101210^{12} bytes) to Exabytes (101810^{18} bytes) and Zettabytes (102110^{21} bytes or 1,000,000,000,000,000,000,0001,000,000,000,000,000,000,000 bytes).
2.  **Velocity (Data in Motion):**
    *   High speed at which real-time data streaming occurs.
    *   Requires response and processing latency measured in milliseconds to seconds.
3.  **Variety (Data in Many Forms):**
    *   Heterogeneity of incoming data streams.
    *   Spans structured tabular data, unstructured text, audio, video, sensor readings, and multimedia.
4.  **Veracity (Data in Doubt):**
    *   Degree of trust, reliability, and accuracy inherent to collected data.
    *   Complicated by data inconsistency, incompleteness, ambiguities, transmission latency, deception, and model approximations.
  • Big Data Solutions: Clustered Computing:

    • A computer cluster consists of connected independent computers operating collaboratively as a unified processing unit.
    • Single-node system hardware is insufficient for Big Data scale; cluster environments resolve scale constraints by distributing computing loads.
    • Clustering software coordinates individual machine resources to deliver robust infrastructure.
  • Benefits of Clustered Computing:

    • Resource Pooling: Combines collective capacity across three foundational hardware resources:
      • Storage Capacity (Hard Disk Drives)
      • Processing Power (CPUs)
      • Working Memory (RAM)
    • High Availability: Provides fault tolerance mechanisms that prevent individual hardware or software module failures from interrupting data accessibility or real-time processing.
    • Easy Scalability: Facilitates horizontal scaling (scaling out) by seamlessly adding standard computing nodes to the existing network without replacing core hardware.

Hadoop Ecosystem and Big Data Life Cycle

  • Overview of Apache Hadoop:
    • An open-source framework designed to simplify Big Data processing and distributed storage management.
    • Enables parallel execution of tasks across large computer clusters using simple, scalable programming paradigms.
    • Originally inspired by technical research papers published by Google.

Hadoop Ecosystem Architecture

  • Hadoop Ecosystem Architectural Layers:

    • Data Storage Layer:

      • HDFS (Hadoop Distributed File System): Underlying distributed storage layer distributing data files across cluster nodes.
      • HBase: Column-oriented NoSQL database system built over HDFS for low-latency read/write access.
    • Data Processing Layer:

      • MapReduce: Framework for writing distributed computing tasks using parallel map and reduce operations.
      • YARN (Yet Another Resource Negotiator): Cluster architectural manager providing resource allocation and job scheduling.
    • Data Access Layer:

      • Hive: Data warehousing engine providing an SQL-like query interface (HiveQL) over Hadoop.
      • Pig: High-level platform providing data-flow scripting capabilities (Pig Latin).
      • Mahout: Distributed library of machine learning and data mining algorithms.
      • Avro: Data serialization system supporting RPC service communication.
      • Sqoop: Specialized tool for bi-directional bulk data transfers between Hadoop and relational database management systems (RDBMS).
    • Data Management Layer:

      • Oozie: Server-based workflow scheduling engine to coordinate dependent Hadoop jobs.
      • Chukwa: Data collection framework for monitoring health across large distributed systems.
      • Flume: Distributed service engineered for streaming log data ingestion into HDFS.
      • ZooKeeper: Centralized service for distributed synchronization, cluster management, configuration management, and naming.
  • Big Data Life Cycle Steps in Hadoop:

    1. Data Ingestion: Ingesting raw external data into the cluster environment.
    2. Storage Processing: Organizing, partitioning, and storing data within cluster persistence mechanisms.
    3. Computation and Analysis: Executing distributed analytics, machine learning, and processing algorithms.
    4. Result Visualization: Rendering output metrics, models, and analytical dashboards for end users.