Introduction to Data Science
Overview of Data Science
- Data Science is a multidisciplinary field that utilizes scientific methods, processes, algorithms, and systems to extract knowledge and actionable insights from data in all formats (structured, semi-structured, and unstructured data).
- The core focus of data science is extracting valuable knowledge from complex datasets.
- It continues to evolve as one of the most promising, in-demand career paths for skilled professionals.
- Data science intersects multiple academic and technical domains:
- Computer Science / IT (Programming skills)
- Mathematics
- Statistics
- Machine Learning
- Information Science
- Data Mining

Intersections within the Data Science Domain:
- Machine Learning: Formed by the combination of Computer Science/IT and Math and Statistics.
- Traditional Research: Formed by the combination of Math and Statistics and Domains/Business Knowledge.
- Software Development: Formed by the combination of Computer Science/IT and Domains/Business Knowledge.
- Data Science: Formed at the central core where Computer Science/IT, Math and Statistics, and Domains/Business Knowledge all overlap.
Roles and Competencies of Data Scientists:
- Data scientists master the full spectrum of the data science life cycle to uncover useful intelligence for organizational decision-making.
- They must be curious, detail-oriented, and result-oriented.
- They require strong communication skills to articulate complex technical findings to non-technical stakeholders.
- They need a solid quantitative foundation in statistics and linear algebra.
- They must possess deep programming knowledge focused on data warehousing, data mining, and data modeling to construct and analyze algorithms.
Data versus Information
Definition of Data:
- Unprocessed or raw facts and figures.
- A collection of facts, concepts, or instructions represented in a formalized manner.
- Must be processed or interpreted by humans or electronic systems to convey true meaning.
- Can be represented in the form of:
- Alphabets (e.g., , )
- Digits (e.g., )
- Special characters (e.g., , , , , , , )
Definition of Information:
- Processed data used to make informed decisions and drive action.
- Data transformed into a meaningful format that carries real or perceived value for current or prospective decisions.
- Interpreted data created from organized, structured, and processed data in a specific context.
Data vs. Information Comparison:
- Data: Described as unprocessed or raw facts and figures | Information: Described as processed data.
- Data: Cannot directly assist in decision-making | Information: Helps directly in decision-making.
- Data: Serves as raw material that can be organized, structured, and interpreted | Information: Created from organized, structured, and processed data within a particular context.
- Data: A collection of text, images, and voice representing quantities, actions, and objects | Information: Processed representation of text, images, and voice conveying context and actionable output.
Data Processing Cycle
- Data Processing Definition: The restructuring or reordering of raw data by human effort or automated machinery to increase its utility and derive specific value.
- Three Core Steps in the Data Processing Cycle:
- Input Step: Input data is prepared and structured into a suitable format depending on the processing machine.
- Processing Step: Transformation activities convert input data into a higher-value, processed form.
- Output Step: Processing results are collected and made available as output (information).

- Concrete Example of Data Processing:
- Input (Data): Raw numeric measurements , , , , alongside raw categorical days: Monday, Tuesday, Wednesday, Thursday, Friday.
- Processing: Arranging, sorting, combining, and applying mathematical operations.
- Output (Information): Contextualized weather readings:
- Monday:
- Tuesday:
- Wednesday:
- Thursday:
- Friday:

Data Types and Perspectives
Computer Programming Perspective:
- A data type is an attribute assigned to data that instructs compilers or interpreters on how the programmer intends to utilize the data.
- Integer (
int): Used to store whole numbers. - Boolean (
bool): Restricted to binary logical values:trueorfalse. - Character (
char): Used to store a single symbol or character. - Floating-Point Number (
float): Used to store real numbers containing fractional/decimal values. - Alphanumeric String (
string): Used to store sequences of letters, numbers, and symbols.
Data Analytics Perspective:
- Data Analytics is the science of evaluating raw data to extract conclusions.
- Categorized into three primary data structures:
- Structured Data:
- Adheres strictly to a pre-defined data model and tabular structure with defined row-column relationships.
- Easily indexed, queried, and analyzed.
- Examples: SQL relational databases, Excel spreadsheets.
- Represents approximately of enterprise data according to Gartner estimates.
- Requires minimal storage capacity and is simple to manage with legacy solutions.
- Unstructured Data:
- Lacks a pre-defined data model and structural organization.
- Typically text-heavy, containing embedded dates, numbers, facts, or multimedia.
- Difficult to process using traditional relational database software.
- Examples: Audio files, video files, PDF documents, Word documents, NoSQL stores.
- Represents approximately of enterprise data according to Gartner estimates.
- Requires larger storage systems and specialized management tools.
- Semi-Structured Data:
- Does not follow a strict relational table structure but contains semantic organizational elements.
- Uses internal tags or markers to separate data elements (self-describing structure).
- Examples: XML files, JSON files.
- Structured Data:

- Comparison Example Across Data Types:

* **Unstructured Text:** "The university has students. John's ID is number , he is years old and already holds a B.Sc. degree. David's ID is number , he is years old and holds a Ph.D. degree. Robert's ID is number , he is years old and also holds the same degree as David, a Ph.D. degree."
* **Semi-Structured XML:**
```xml}
<University>
<Student ID="1">
<Name>John</Name>
<Age>18</Age>
<Degree>B.Sc.</Degree>
</Student>
<Student ID="2">
<Name>David</Name>
<Age>31</Age>
<Degree>Ph.D.</Degree>
</Student>
</University>
``
* **Structured Table:**
* Columns:ID,Name,Age,Degree`
* Row : , John, , B.Sc.
* Row : , David, , Ph.D.
* Row : , Robert, , Ph.D.
* Row : , Rick, , M.Sc.
* Row : , Michael, , B.Sc.
- Metadata – Data About Data:
- Provides supplementary descriptive information about specific datasets.
- Not a separate structural format technically, but essential for Big Data solutions and indexing.
- Example: Metadata associated with photograph files specifies creation timestamps and precise geotagged locations.
- Individual metadata fields (such as timestamps and geographic coordinates) are structured, allowing Big Data platforms to leverage them for rapid initial filtering and analysis.
Big Data Value Chain
- The Big Data Value Chain (DVC) models the end-to-end information flow within big data architectures to extract actionable business value.

Five Key Phases of the Big Data Value Chain:
Data Acquisition:
- Process of gathering, filtering, and cleaning data before loading into data warehouses or storage platforms.
- Imposes major infrastructure challenges requiring minimal latency, high transaction throughput, and dynamic structural handling.
- Includes: Structured and unstructured sources, event processing, sensor networks, protocols, real-time data streams, and multimodality.
Data Analysis:
- Transforming raw acquired data into formats actionable for decision-making and business domain application.
- Involves exploratory analysis, data transformation, modeling, synthesis, and extraction of high-potential business insights.
- Includes: Stream mining, semantic analysis, machine learning, information extraction, linked data, data discovery, whole-world semantics, cross-sectorial analysis, and ecosystem analytics.
Data Curation:
- Active management throughout the data life cycle to guarantee quality meets application requirements.
- Includes content creation, selection, classification, transformation, validation, and long-term preservation.
- Executed by data curators, scientific curators, and annotators to elevate quality, trust, discoverability, accessibility, and reusability.
- Includes: Provenance tracking, annotation, human-data interaction, community/crowd curation, human computation, incentivisation, and system interoperability.
Data Storage:
- Scalable persistence and management of data to deliver rapid access for downstream systems.
- Transitions from traditional Relational Database Management Systems (RDBMS) to Data Lakes.
- Data Lakes support diverse formats built on Hadoop clusters, cloud object storage, or NoSQL databases.
- Includes: In-memory databases, NoSQL databases, NewSQL databases, cloud infrastructure, query interfaces, CAP theorem trade-offs (Consistency, Availability, Partition-tolerance), security, privacy, and standardization.
Data Usage:
- Integration of analyzed data directly into operational business workflows and decision support systems.
- Drives competitiveness by reducing overhead, unlocking new service value, and optimizing key operational metrics.
- Includes: Decision support systems, predictive analytics, in-use analytics, simulations, interactive visualizations, modeling, and automated control.
Concepts of Big Data
Definition of Big Data:
- Refers to massive, complex datasets that surpass the processing, storage, and management capacities of traditional relational databases or single-node computing systems.
- The threshold defining a "large dataset" is relative, shifting continually over time and varying across individual organizational capabilities.
The Four V's of Big Data:

1. **Volume (Data at Rest):**
* Massive scale of collected datasets.
* Ranges from Terabytes ( bytes) to Exabytes ( bytes) and Zettabytes ( bytes or bytes).
2. **Velocity (Data in Motion):**
* High speed at which real-time data streaming occurs.
* Requires response and processing latency measured in milliseconds to seconds.
3. **Variety (Data in Many Forms):**
* Heterogeneity of incoming data streams.
* Spans structured tabular data, unstructured text, audio, video, sensor readings, and multimedia.
4. **Veracity (Data in Doubt):**
* Degree of trust, reliability, and accuracy inherent to collected data.
* Complicated by data inconsistency, incompleteness, ambiguities, transmission latency, deception, and model approximations.
Big Data Solutions: Clustered Computing:
- A computer cluster consists of connected independent computers operating collaboratively as a unified processing unit.
- Single-node system hardware is insufficient for Big Data scale; cluster environments resolve scale constraints by distributing computing loads.
- Clustering software coordinates individual machine resources to deliver robust infrastructure.
Benefits of Clustered Computing:
- Resource Pooling: Combines collective capacity across three foundational hardware resources:
- Storage Capacity (Hard Disk Drives)
- Processing Power (CPUs)
- Working Memory (RAM)
- High Availability: Provides fault tolerance mechanisms that prevent individual hardware or software module failures from interrupting data accessibility or real-time processing.
- Easy Scalability: Facilitates horizontal scaling (scaling out) by seamlessly adding standard computing nodes to the existing network without replacing core hardware.
- Resource Pooling: Combines collective capacity across three foundational hardware resources:
Hadoop Ecosystem and Big Data Life Cycle
- Overview of Apache Hadoop:
- An open-source framework designed to simplify Big Data processing and distributed storage management.
- Enables parallel execution of tasks across large computer clusters using simple, scalable programming paradigms.
- Originally inspired by technical research papers published by Google.

Hadoop Ecosystem Architectural Layers:
Data Storage Layer:
- HDFS (Hadoop Distributed File System): Underlying distributed storage layer distributing data files across cluster nodes.
- HBase: Column-oriented NoSQL database system built over HDFS for low-latency read/write access.
Data Processing Layer:
- MapReduce: Framework for writing distributed computing tasks using parallel map and reduce operations.
- YARN (Yet Another Resource Negotiator): Cluster architectural manager providing resource allocation and job scheduling.
Data Access Layer:
- Hive: Data warehousing engine providing an SQL-like query interface (HiveQL) over Hadoop.
- Pig: High-level platform providing data-flow scripting capabilities (Pig Latin).
- Mahout: Distributed library of machine learning and data mining algorithms.
- Avro: Data serialization system supporting RPC service communication.
- Sqoop: Specialized tool for bi-directional bulk data transfers between Hadoop and relational database management systems (RDBMS).
Data Management Layer:
- Oozie: Server-based workflow scheduling engine to coordinate dependent Hadoop jobs.
- Chukwa: Data collection framework for monitoring health across large distributed systems.
- Flume: Distributed service engineered for streaming log data ingestion into HDFS.
- ZooKeeper: Centralized service for distributed synchronization, cluster management, configuration management, and naming.
Big Data Life Cycle Steps in Hadoop:
- Data Ingestion: Ingesting raw external data into the cluster environment.
- Storage Processing: Organizing, partitioning, and storing data within cluster persistence mechanisms.
- Computation and Analysis: Executing distributed analytics, machine learning, and processing algorithms.
- Result Visualization: Rendering output metrics, models, and analytical dashboards for end users.