CLASS 10 DATA SCIENCES-DATAFRAME AND CSV FILES
Scatter Plots
Definition: A scatter plot is a graph where each value in the data set is represented by a dot.
Purpose: Used to observe relationships between variables.
Type of Data: Ideal for plotting discontinuous data, which lacks a continuous flow.
Method: Use the
scatter()method to create a scatter plot.
Plotting a Scatter Graph
Import Libraries:
import matplotlib.pyplot as plt
Data Representation:
X-axis values:
x = [5,7,8,7,2,17,2,9,4,11,12,9,6](represents the age of cars)Y-axis values:
y = [99,86,87,88,111,86,103,87,94,78,77,85,86](represents speed of cars)
Creating the Plot:
plt.scatter(x, y)plt.show()
Visualization of Multivariable Relationships
Scatter Plot Features:
X and Y axes represent different parameters.
The color and size of the dots can indicate additional variables, allowing the visualization of up to four parameters at once through a single coordinate point.
Lab Exercise
Task: Write a program to create a scatter chart for the following points:
(2,5)
(9,10)
(8,3)
(5,7)
(6,18)
Introduction to Pandas
Definition: Pandas is a software library for Python used for data manipulation and analysis.
Origin of Name: Derived from "panel data."
Capabilities:
Handles a wide range of applications in finance, statistics, social science, and engineering.
Provides data structures for manipulating numerical tables and time series.
Data Structures in Pandas
Series: 1-dimensional data structure.
DataFrame: 2-dimensional, labeled data.
Installing Pandas
Command:
pip install pandas
Creating a DataFrame
From a List:
import pandas as pd data1 = [10,22,36,24,59] df1 = pd.DataFrame(data1) print(df1)This will create a DataFrame with default indexed rows starting from 0.
From an Array:
import numpy as np import pandas as pd A1 = np.array([100,80,94]) A2 = np.array([40,50,60]) A3 = np.array([70,35,46]) df = pd.DataFrame([A1, A2, A3]) print(df)
Working with CSV Files in Pandas
Data Storage:
DataFrames can hold 2D tabular data and can interact with CSV files.
Data can be imported from or exported to CSV files, MySQL, and Excel databases.
Understanding CSV Files
Definition: CSV stands for Comma Separated Values; a human-readable plain text format for storing tabular data.
Structure:
Each line represents a record.
Fields within each record are separated by commas.
Advantages of CSV Files
Benefits:
Easy to read, manage, and process.
Small file size and fast transfer.
Supported by nearly all database and spreadsheet software.
CSV vs. Excel:
Faster and simpler; can be edited with text editors whereas Excel files may have restrictions like password protection.
Handling CSV with Notepad or Text Editors
Creating a CSV File:
Open a new file.
Enter data separated by commas and new lines for rows.
Save the file with a
.csvextension.
Using Excel to Create CSV Files
Open your excel workbook.
Navigate to File > Save As.
Choose CSV format from the list of file types.
Reading and Writing CSV Files in Python
Reading: Use the
read_csv()method to load a CSV file into a DataFrame.Example:
import pandas as pd df = pd.read_csv('path_to_file.csv') print(df.head()) # displays the first 5 records
Practice Exercises
Tasks:
Display a scatter chart for specific data points.
Read a CSV file from your system and display its information.
Read a CSV and display 10 rows.