Data Science Part1
Page 1: Introduction to Data Science in Microsoft Fabric
Overview of Data Science in Microsoft Fabric
Microsoft Fabric offers Data Science experiences aimed at empowering users to complete end-to-end data science workflows.
Users can access a variety of activities ranging from data exploration to reporting insights.
Core Activities in Data Science Workflow
Data Exploration: Analyze datasets to understand their structure and contents.
Preparation and Cleansing: Clean and transform data to make it suitable for analysis.
Modeling: Develop predictive models using machine learning techniques.
Model Scoring: Evaluate model performance against datasets.
Serving Insights: Integrate predictive insights into business intelligence reports.
Key Features of Microsoft Fabric's Data Science Home Page
Data Science Home Page: A dedicated portal where users can discover and access resources related to data science.
Capabilities include:
Creating experiments for machine learning.
Managing models and notebooks.
Importing existing notebooks.
Typical Data Science Process Steps
Common steps involved in machine learning projects include:
Problem formulation and ideation
Data discovery and pre-processing
Experimentation and modeling
Enriching and operationalizing predictions
Page 2: Comprehensive Overview of Data Science Process
Phases of Data Science Process
The data science process is broken down into the following phases:
Problem Formulation and Ideation: Identifying the problem and brainstorming solutions.
Data Discovery and Pre-processing: Gathering relevant data and cleaning it for analysis.
Exploration: Analyzing the data to uncover patterns and insights.
Modeling: Designing and training models based on prepared data.
Evaluation: Assessing the model's predictions and making improvements.
Operationalization: Integrating the model outputs into actionable insights for the business.
Collaboration Between Roles
Seamless data sharing among analysts and data science practitioners enhances collaboration.
Power BI integration allows easy sharing of reports and datasets.
Page 3: Data Exploration and Preparation Tools
Tools for Data Preparation
Various tools are available for effective data exploration and visualization in Microsoft Fabric, including:
Notebooks: Simplified access to data exploration tools.
Apache Spark and Python: Enables high-scale data preparation facilities.
Open-source Libraries: Leverage libraries for improved visualization capabilities.
Data Wrangler
A feature for seamless data cleansing and transformation that generates Python code automatically, aiding in:
Reducing tedious tasks in data preparation.
Building repeatable and automated workflows.
Page 4: Advanced Machine Learning Modelling
MLflow and Experimentation
Microsoft Fabric provides built-in support for managing machine learning experiments using MLflow:
Track and log experiments and models for better organization.
Use of various popular libraries (like Scikit Learn) to train models efficiently.
SynapseML Framework
An open-source library designed to simplify the construction of scalable machine learning pipelines. Offers:
Unified API to access multiple ML frameworks.
Features for both prediction and development of predictive models.
Page 5: Integration and Insights Generation
Insights Integration
Predictions generated from machine learning models can be directly written to OneLake and consumed via Power BI, facilitating:
Direct integration for easy sharing of results.
Scheduling options available for batch scoring via Notebooks.
Benefits of Insights Automated Distribution
Reduces manual effort in data refreshes, ensuring stakeholders have access to the latest predictions without delays.
Page 6: SynapseML Overview
SynapseML Functionality
SynapseML serves as a powerful tool within Microsoft Fabric for creating scalable ML pipelines and integrates:
Text analytics, vision, and anomaly detection functionalities.
Simple APIs for enhanced query generation.
Framework Compatibility
Works seamlessly with existing Apache Spark workflows, allowing easy integration of models built using SynapseML within broader applications.
Page 7: Enhanced API Features
Universal API Access
SynapseML provides unified access across various ML frameworks, improving the efficiency of developing complex ML applications.
Azure AI Services
The integration of Azure AI services within SynapseML offers pre-built intelligent services and enhances model training capabilities.
Page 8: Responsible AI Insights
Building Responsible AI Systems
Tools within SynapseML help in tracking models and explaining their predictions:
Ensures datasets are unbiased and aligns training datasets with responsible AI guidelines.
Enterprise Support
Full support available under Azure Synapse Analytics, allowing organizations to manage large-scale ML systems effectively.
Page 9: Data Science Tutorial Series
Tutorial Overview
A comprehensive series for implementing end-to-end scenarios in data science using Microsoft Fabric.
Includes steps for:
Data ingestion
Data preprocessing and cleansing
Training models
Generating and visualizing insights.
Page 10: Key Architectural Components
Architecture of Data Science Workflow
The tutorial series features architecture that supports:
Data ingestion from various external and internal sources.
Batch scoring and integration with Power BI for reporting purposes.
Components
Utilize tools like Data Lakehouse for storing and managing data.
Engage with Notebooks to enhance data transformation and exploration.
Page 11: Model and Experiment Management
Experimentation Framework
Fabric's built-in capabilities facilitate model training, evaluation, and registration through MLflow:
Enables recording and tracking of ML experiments systematically.
Provides easy access to different models and facilitates A/B testing.
Page 12: System Preparation for Tutorials
Prerequisites for Tutorials
A valid Microsoft Fabric subscription is necessary.
Set up a lakehouse before proceeding with tutorial notebooks to ensure a smooth workflow.
Page 13-14: Importing Notebooks
Notebook Management
Demonstrates how to import Jupyter notebooks into the Data Science workspace and manage environments for model experiments.
Page 15-16: Data Ingestion Techniques
Ingesting Data into Lakehouse
Instructions for using Apache Spark to read data into Fabric lakehouses and transforming it into delta table formats.
Page 17: Understanding the Dataset
Bank Customers Churn Dataset
Contains critical attributes impacting customer retention within the bank. Key attributes include:
Credit score, age, geographical location, and products purchased.
Page 18: Downloading Public Datasets
Data Retrieval Process
The tutorial outlines how to download datasets from public blobs and store them into a Fabric lakehouse for analysis.
Page 19-20: Exploratory Data Analysis using Notebooks
Data Analysis and Visualization
How to conduct exploratory data analysis and visualize data using seaborn and Data Wrangler, a tool integrated directly into notebooks.
Page 21: Pandas DataFrame Creation
Converting Data Formats
Procedure to convert Spark DataFrames to Pandas Dataframes for further analysis and ease of use in visualization libraries.
Page 22: Utilizing Data Wrangler for Data Cleaning
Role of Data Wrangler
Steps to initiate Data Wrangler within notebooks to conduct data cleaning and preparation tasks efficiently.
Page 23-26: Data Cleaning Techniques
Data Cleaning Steps using Data Wrangler
Detailed operational steps within Data Wrangler for tasks such as:
Removing duplicates
Dropping missing values
Finalizing and saving cleaned data.
Page 27: Data Exploration and Visualization
Visualization Techniques
Techniques to visualize cleaned data using various plotting libraries and methods in Python.
Page 28-31: Comparison of Categorical and Numerical Attributes
Analysis of Customer Data
Explore how various attributes relate to churn, focusing on distributions and trends observed in the dataset.
Page 32: Model Training and Registration
Training Techniques
Steps to train machine learning models (Random Forests and LightGBM) using the Microsoft Fabric environment and registering models using MLflow.
Page 33-37: Preparing Training Data
Data Set Preparation
Processes for organizing and preparing datasets for training while addressing challenges such as class imbalance through methods like SMOTE.
Page 38-42: Evaluating Model Performance
Performance Assessment
Evaluating trained models using metrics like accuracy, confusion matrix, and ROC curves to gauge effectiveness.
Page 43-44: Batch Scoring with PREDICT
Operationalizing Model Predictions
Procedure for performing batch scoring with the trained models and saving predictions directly within the Fabric framework.
Page 45-58: Visualizing Results with Power BI
Pivoting to Visualization
Instructions for creating Power BI reports based on generated model predictions, illustrating successful visual representation of data insights.
Page 59-62: Introduction to Copilot in Data Science
New Copilot Features
How Copilot enhances the data science workflow in Microsoft Fabric through interactive querying and code generation capabilities.
Page 63-74: Using Copilot for Data Insights and Queries
Enhanced Interaction with Data
Explore how the Copilot facilitates data exploration, generates insights, and aids in data manipulation through natural language commands.
Page 75-100: Advanced AI Skill Configuration
AI Skill Implementation Guide
A thorough walkthrough on creating and configuring an AI skill in Microsoft Fabric, covering the nuances of data access, question handling, and example-driven learning for enhanced accuracy in outputs.