CLASS 10 DATA SCIENCES-DATAFRAME AND CSV FILES

Scatter Plots

  • Definition: A scatter plot is a graph where each value in the data set is represented by a dot.

  • Purpose: Used to observe relationships between variables.

  • Type of Data: Ideal for plotting discontinuous data, which lacks a continuous flow.

  • Method: Use the scatter() method to create a scatter plot.

Plotting a Scatter Graph

  1. Import Libraries:

    • import matplotlib.pyplot as plt

  2. Data Representation:

    • X-axis values: x = [5,7,8,7,2,17,2,9,4,11,12,9,6] (represents the age of cars)

    • Y-axis values: y = [99,86,87,88,111,86,103,87,94,78,77,85,86] (represents speed of cars)

  3. Creating the Plot:

    • plt.scatter(x, y)

    • plt.show()

Visualization of Multivariable Relationships

  • Scatter Plot Features:

    • X and Y axes represent different parameters.

    • The color and size of the dots can indicate additional variables, allowing the visualization of up to four parameters at once through a single coordinate point.

Lab Exercise

  • Task: Write a program to create a scatter chart for the following points:

    • (2,5)

    • (9,10)

    • (8,3)

    • (5,7)

    • (6,18)

Introduction to Pandas

  • Definition: Pandas is a software library for Python used for data manipulation and analysis.

  • Origin of Name: Derived from "panel data."

  • Capabilities:

    • Handles a wide range of applications in finance, statistics, social science, and engineering.

    • Provides data structures for manipulating numerical tables and time series.

Data Structures in Pandas

  1. Series: 1-dimensional data structure.

  2. DataFrame: 2-dimensional, labeled data.

Installing Pandas

  • Command: pip install pandas

Creating a DataFrame

  • From a List:

    import pandas as pd
    data1 = [10,22,36,24,59]
    df1 = pd.DataFrame(data1)
    print(df1)
    • This will create a DataFrame with default indexed rows starting from 0.

  • From an Array:

    import numpy as np
    import pandas as pd
    A1 = np.array([100,80,94])
    A2 = np.array([40,50,60])
    A3 = np.array([70,35,46])
    df = pd.DataFrame([A1, A2, A3])
    print(df)

Working with CSV Files in Pandas

  • Data Storage:

    • DataFrames can hold 2D tabular data and can interact with CSV files.

    • Data can be imported from or exported to CSV files, MySQL, and Excel databases.

Understanding CSV Files

  • Definition: CSV stands for Comma Separated Values; a human-readable plain text format for storing tabular data.

  • Structure:

    • Each line represents a record.

    • Fields within each record are separated by commas.

Advantages of CSV Files

  • Benefits:

    • Easy to read, manage, and process.

    • Small file size and fast transfer.

    • Supported by nearly all database and spreadsheet software.

  • CSV vs. Excel:

    • Faster and simpler; can be edited with text editors whereas Excel files may have restrictions like password protection.

Handling CSV with Notepad or Text Editors

  • Creating a CSV File:

    1. Open a new file.

    2. Enter data separated by commas and new lines for rows.

    3. Save the file with a .csv extension.

Using Excel to Create CSV Files

  1. Open your excel workbook.

  2. Navigate to File > Save As.

  3. Choose CSV format from the list of file types.

Reading and Writing CSV Files in Python

  • Reading: Use the read_csv() method to load a CSV file into a DataFrame.

  • Example:

    import pandas as pd
    df = pd.read_csv('path_to_file.csv')
    print(df.head())  # displays the first 5 records

Practice Exercises

  • Tasks:

    • Display a scatter chart for specific data points.

    • Read a CSV file from your system and display its information.

    • Read a CSV and display 10 rows.