Data Science Part1

Page 1: Introduction to Data Science in Microsoft Fabric

Overview of Data Science in Microsoft Fabric

  • Microsoft Fabric offers Data Science experiences aimed at empowering users to complete end-to-end data science workflows.

  • Users can access a variety of activities ranging from data exploration to reporting insights.

Core Activities in Data Science Workflow

  1. Data Exploration: Analyze datasets to understand their structure and contents.

  2. Preparation and Cleansing: Clean and transform data to make it suitable for analysis.

  3. Modeling: Develop predictive models using machine learning techniques.

  4. Model Scoring: Evaluate model performance against datasets.

  5. Serving Insights: Integrate predictive insights into business intelligence reports.

Key Features of Microsoft Fabric's Data Science Home Page

  • Data Science Home Page: A dedicated portal where users can discover and access resources related to data science.

  • Capabilities include:

    • Creating experiments for machine learning.

    • Managing models and notebooks.

    • Importing existing notebooks.

Typical Data Science Process Steps

  • Common steps involved in machine learning projects include:

    • Problem formulation and ideation

    • Data discovery and pre-processing

    • Experimentation and modeling

    • Enriching and operationalizing predictions


Page 2: Comprehensive Overview of Data Science Process

Phases of Data Science Process

  • The data science process is broken down into the following phases:

    • Problem Formulation and Ideation: Identifying the problem and brainstorming solutions.

    • Data Discovery and Pre-processing: Gathering relevant data and cleaning it for analysis.

    • Exploration: Analyzing the data to uncover patterns and insights.

    • Modeling: Designing and training models based on prepared data.

    • Evaluation: Assessing the model's predictions and making improvements.

    • Operationalization: Integrating the model outputs into actionable insights for the business.

Collaboration Between Roles

  • Seamless data sharing among analysts and data science practitioners enhances collaboration.

  • Power BI integration allows easy sharing of reports and datasets.


Page 3: Data Exploration and Preparation Tools

Tools for Data Preparation

  • Various tools are available for effective data exploration and visualization in Microsoft Fabric, including:

    • Notebooks: Simplified access to data exploration tools.

    • Apache Spark and Python: Enables high-scale data preparation facilities.

    • Open-source Libraries: Leverage libraries for improved visualization capabilities.

Data Wrangler

  • A feature for seamless data cleansing and transformation that generates Python code automatically, aiding in:

    • Reducing tedious tasks in data preparation.

    • Building repeatable and automated workflows.


Page 4: Advanced Machine Learning Modelling

MLflow and Experimentation

  • Microsoft Fabric provides built-in support for managing machine learning experiments using MLflow:

    • Track and log experiments and models for better organization.

    • Use of various popular libraries (like Scikit Learn) to train models efficiently.

SynapseML Framework

  • An open-source library designed to simplify the construction of scalable machine learning pipelines. Offers:

    • Unified API to access multiple ML frameworks.

    • Features for both prediction and development of predictive models.


Page 5: Integration and Insights Generation

Insights Integration

  • Predictions generated from machine learning models can be directly written to OneLake and consumed via Power BI, facilitating:

    • Direct integration for easy sharing of results.

    • Scheduling options available for batch scoring via Notebooks.

Benefits of Insights Automated Distribution

  • Reduces manual effort in data refreshes, ensuring stakeholders have access to the latest predictions without delays.


Page 6: SynapseML Overview

SynapseML Functionality

  • SynapseML serves as a powerful tool within Microsoft Fabric for creating scalable ML pipelines and integrates:

    • Text analytics, vision, and anomaly detection functionalities.

    • Simple APIs for enhanced query generation.

Framework Compatibility

  • Works seamlessly with existing Apache Spark workflows, allowing easy integration of models built using SynapseML within broader applications.


Page 7: Enhanced API Features

Universal API Access

  • SynapseML provides unified access across various ML frameworks, improving the efficiency of developing complex ML applications.

Azure AI Services

  • The integration of Azure AI services within SynapseML offers pre-built intelligent services and enhances model training capabilities.


Page 8: Responsible AI Insights

Building Responsible AI Systems

  • Tools within SynapseML help in tracking models and explaining their predictions:

    • Ensures datasets are unbiased and aligns training datasets with responsible AI guidelines.

Enterprise Support

  • Full support available under Azure Synapse Analytics, allowing organizations to manage large-scale ML systems effectively.


Page 9: Data Science Tutorial Series

Tutorial Overview

  • A comprehensive series for implementing end-to-end scenarios in data science using Microsoft Fabric.

  • Includes steps for:

    1. Data ingestion

    2. Data preprocessing and cleansing

    3. Training models

    4. Generating and visualizing insights.


Page 10: Key Architectural Components

Architecture of Data Science Workflow

  • The tutorial series features architecture that supports:

    • Data ingestion from various external and internal sources.

    • Batch scoring and integration with Power BI for reporting purposes.

Components

  • Utilize tools like Data Lakehouse for storing and managing data.

  • Engage with Notebooks to enhance data transformation and exploration.


Page 11: Model and Experiment Management

Experimentation Framework

  • Fabric's built-in capabilities facilitate model training, evaluation, and registration through MLflow:

    • Enables recording and tracking of ML experiments systematically.

    • Provides easy access to different models and facilitates A/B testing.


Page 12: System Preparation for Tutorials

Prerequisites for Tutorials

  • A valid Microsoft Fabric subscription is necessary.

  • Set up a lakehouse before proceeding with tutorial notebooks to ensure a smooth workflow.


Page 13-14: Importing Notebooks

Notebook Management

  • Demonstrates how to import Jupyter notebooks into the Data Science workspace and manage environments for model experiments.


Page 15-16: Data Ingestion Techniques

Ingesting Data into Lakehouse

  • Instructions for using Apache Spark to read data into Fabric lakehouses and transforming it into delta table formats.


Page 17: Understanding the Dataset

Bank Customers Churn Dataset

  • Contains critical attributes impacting customer retention within the bank. Key attributes include:

    • Credit score, age, geographical location, and products purchased.


Page 18: Downloading Public Datasets

Data Retrieval Process

  • The tutorial outlines how to download datasets from public blobs and store them into a Fabric lakehouse for analysis.


Page 19-20: Exploratory Data Analysis using Notebooks

Data Analysis and Visualization

  • How to conduct exploratory data analysis and visualize data using seaborn and Data Wrangler, a tool integrated directly into notebooks.


Page 21: Pandas DataFrame Creation

Converting Data Formats

  • Procedure to convert Spark DataFrames to Pandas Dataframes for further analysis and ease of use in visualization libraries.


Page 22: Utilizing Data Wrangler for Data Cleaning

Role of Data Wrangler

  • Steps to initiate Data Wrangler within notebooks to conduct data cleaning and preparation tasks efficiently.


Page 23-26: Data Cleaning Techniques

Data Cleaning Steps using Data Wrangler

  • Detailed operational steps within Data Wrangler for tasks such as:

    • Removing duplicates

    • Dropping missing values

    • Finalizing and saving cleaned data.


Page 27: Data Exploration and Visualization

Visualization Techniques

  • Techniques to visualize cleaned data using various plotting libraries and methods in Python.


Page 28-31: Comparison of Categorical and Numerical Attributes

Analysis of Customer Data

  • Explore how various attributes relate to churn, focusing on distributions and trends observed in the dataset.


Page 32: Model Training and Registration

Training Techniques

  • Steps to train machine learning models (Random Forests and LightGBM) using the Microsoft Fabric environment and registering models using MLflow.


Page 33-37: Preparing Training Data

Data Set Preparation

  • Processes for organizing and preparing datasets for training while addressing challenges such as class imbalance through methods like SMOTE.


Page 38-42: Evaluating Model Performance

Performance Assessment

  • Evaluating trained models using metrics like accuracy, confusion matrix, and ROC curves to gauge effectiveness.


Page 43-44: Batch Scoring with PREDICT

Operationalizing Model Predictions

  • Procedure for performing batch scoring with the trained models and saving predictions directly within the Fabric framework.


Page 45-58: Visualizing Results with Power BI

Pivoting to Visualization

  • Instructions for creating Power BI reports based on generated model predictions, illustrating successful visual representation of data insights.


Page 59-62: Introduction to Copilot in Data Science

New Copilot Features

  • How Copilot enhances the data science workflow in Microsoft Fabric through interactive querying and code generation capabilities.


Page 63-74: Using Copilot for Data Insights and Queries

Enhanced Interaction with Data

  • Explore how the Copilot facilitates data exploration, generates insights, and aids in data manipulation through natural language commands.


Page 75-100: Advanced AI Skill Configuration

AI Skill Implementation Guide

  • A thorough walkthrough on creating and configuring an AI skill in Microsoft Fabric, covering the nuances of data access, question handling, and example-driven learning for enhanced accuracy in outputs.