BTCE | 5th Sem
Adv-Python SubjectUnit 4

Adv-Python Unit 4: Complete Concepts Guide

Unit IV: Pandas, NumPy, and Matplotlib -> Generated and Prepared By Thiruselvan (ThiruXD)

1. Introduction to Data Science Libraries

Python is widely used in data science because it provides simple syntax and powerful libraries for working with data. Three of the most important libraries for beginners are:

LibraryPurposePrimary Use
PandasTabular data handlingData cleaning, analysis, manipulation
NumPyNumerical computationFast array-based computing
MatplotlibData visualizationCreating charts and graphs

Key Insight: These three libraries form the foundation of Python data analysis. Pandas handles structured data, NumPy handles numerical operations, and Matplotlib communicates results visually.


2. Introduction to Pandas

2.1 What is Pandas?

Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures and functions for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs.

The name Pandas is derived from Panel Data, which refers to multidimensional structured datasets.

2.2 Why Use Pandas?

Pandas is especially useful when data contains:

  • Rows and columns
  • Missing values
  • Labels
  • Mixed data types
  • Dates
  • Categories
  • Numerical values

Pandas is widely used by data analysts, data scientists, researchers, and software developers because it makes data handling easier than using basic Python lists and dictionaries.

2.3 Important Features of Pandas

FeatureExplanationExample
Data loadingReads data from files and data sourcesread_csv(), read_json(), read_excel()
Data cleaningHandles missing values, duplicates, incorrect formatsdropna(), fillna(), drop_duplicates()
Data selectionSelects rows, columns, and subsetsdf["Name"], loc[], iloc[]
Data transformationCreates new columns and modifies existing datadf["Total"] = df["A"] + df["B"]
Data aggregationGroups and summarizes datagroupby(), mean(), sum()
Data exportWrites results back to filesto_csv(), to_json()

2.4 Installing and Importing Pandas

In most data science environments (Anaconda, Google Colab, Jupyter Notebook), Pandas is already available. In a local environment, install using pip.

# Install Pandas
pip install pandas
# Standard import convention
import pandas as pd

2.5 Pandas Workflow

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Import    │ →  │ 2. Create    │ →  │ 3. Clean &   │ →  │ 4. Combine   │ →  │ 5. Analyze   │
│    Data      │    │  Series/DF   │    │   Transform  │    │   Datasets   │    │  & Export    │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘
  1. Import Data: Load data from various sources.
  2. Create Series/DataFrame: Organize data into Series or DataFrame structures.
  3. Clean & Transform: Handle missing values, filter, sort, and transform data.
  4. Combine Datasets: Merge, join, or concatenate multiple datasets.
  5. Analyze & Export: Analyze data to extract insights and export results.

3. Pandas Series and DataFrames

3.1 Pandas Series

A Series is a one-dimensional labelled array. It can store integers, floats, strings, or other Python objects. A Series is similar to a single column in a spreadsheet.

import pandas as pd

marks = pd.Series([78, 85, 92, 66], index=["Amit", "Ravi", "Neha", "Sara"])
print(marks)
print(marks["Neha"])  # 92

Components of a Pandas Series:

ComponentMeaning
ValuesThe actual data stored in the Series
IndexLabels used to identify each value
Data typeThe type of values stored (int64, float64, object)

3.2 Pandas DataFrame

A DataFrame is a two-dimensional labelled data structure with rows and columns. It is the most commonly used data structure in Pandas. DataFrames can be created from:

  • Dictionaries
  • Lists
  • CSV files
  • JSON files
  • Excel files
  • Databases
  • APIs
import pandas as pd

data = {
    "Name": ["Amit", "Ravi", "Neha", "Sara"],
    "Age": [21, 22, 20, 23],
    "Marks": [78, 85, 92, 66]
}

df = pd.DataFrame(data)
print(df)

Output:

   Name  Age  Marks
0  Amit   21     78
1  Ravi   22     85
2  Neha   20     92
3  Sara   23     66

3.3 Common DataFrame Inspection Methods

OperationPurposeExample
head()Displays first five rows by defaultdf.head()
tail()Displays last five rows by defaultdf.tail()
shapeShows number of rows and columnsdf.shape
columnsShows column namesdf.columns
info()Displays summary of columns and data typesdf.info()
describe()Displays statistical summary of numerical columnsdf.describe()

4. Manipulating and Combining DataFrames

Data manipulation means changing, cleaning, selecting, filtering, sorting, grouping, or combining data. In real data analysis, raw data is rarely ready for direct use.

4.1 Selecting Columns and Rows

# Select one column
print(df["Name"])

# Select multiple columns
print(df[["Name", "Marks"]])

# Select rows using label-based indexing
print(df.loc[0])

# Select rows using integer position
print(df.iloc[0:2])
MethodUsed ForExample
df["col"]Selecting a single columndf["Marks"]
df[["col1", "col2"]]Selecting multiple columnsdf[["Name", "Age"]]
loc[]Selecting by labels or conditionsdf.loc[df["Marks"] > 80]
iloc[]Selecting by integer positiondf.iloc[0:3, 1:3]

4.2 Filtering and Creating Columns

# Filter rows
high_scorers = df[df["Marks"] > 80]
print(high_scorers)

# Create new column
df["Total"] = df["Marks"] + 10
print(df)

4.3 Handling Missing Values

Missing values are common in real datasets. Pandas represents missing values using NaN (Not a Number).

import pandas as pd

data = {"Name": ["Amit", "Ravi", "Neha"], "Marks": [78, None, 92]}
df = pd.DataFrame(data)

print(df.isnull())    # Check missing values
print(df.dropna())    # Remove rows with missing values
print(df.fillna(0))   # Replace missing values with 0

Data Cleaning Functions:

FunctionPurpose
isnull()Detects missing values
notnull()Detects non-missing values
dropna()Removes rows or columns with missing values
fillna(value)Replaces missing values with a specified value
drop_duplicates()Removes duplicate rows

4.4 Combining DataFrames

Combining DataFrames is required when data is stored in multiple tables. Pandas supports concatenation, merging, and joining.

Concatenation (stacking):

import pandas as pd

df1 = pd.DataFrame({"ID": [1, 2], "Name": ["Amit", "Ravi"]})
df2 = pd.DataFrame({"ID": [3, 4], "Name": ["Neha", "Sara"]})

combined = pd.concat([df1, df2], ignore_index=True)
print(combined)

Merging (common column):

marks = pd.DataFrame({"ID": [1, 2, 3], "Marks": [78, 85, 92]})
students = pd.DataFrame({"ID": [1, 2, 3], "Name": ["Amit", "Ravi", "Neha"]})

result = pd.merge(students, marks, on="ID")
print(result)

DataFrame Combining Methods:

OperationMeaningWhen to Use
concat()Stacks DataFrames row-wise or column-wiseWhen datasets have same columns or same index
merge()Combines DataFrames using a common columnWhen tables share a key such as ID
join()Combines DataFrames using indexWhen index labels are meaningful

5. Reading CSV and JSON Files Using Pandas

Pandas can read data from many file formats. Two of the most common formats are CSV and JSON.

5.1 Reading CSV Files

import pandas as pd

# Read a CSV file
students = pd.read_csv("students.csv")

print(students.head())
print(students.shape)
print(students.info())

5.2 Writing CSV Files

# Save a DataFrame to a CSV file
students.to_csv("cleaned_students.csv", index=False)

5.3 Reading JSON Files

import pandas as pd

# Read a JSON file
orders = pd.read_json("orders.json")
print(orders.head())

5.4 File Reading and Writing Functions

FunctionPurposeCommon Parameter
read_csv()Reads a CSV file into a DataFramesep, header, names, usecols
to_csv()Writes a DataFrame to a CSV fileindex=False
read_json()Reads a JSON file into a DataFrameorient
to_json()Writes a DataFrame to a JSON fileorient, indent

Practical Tip: After reading any file, always use head(), shape, info(), and describe() to understand the structure and quality of the imported data.


6. NumPy Arrays and ndarray Objects

6.1 What is NumPy?

NumPy stands for Numerical Python. It is a fundamental Python library for scientific computing and numerical operations. NumPy provides the ndarray object, which stores elements of the same data type in a compact and efficient way.

6.2 Why Use NumPy?

Compared with normal Python lists, NumPy arrays are:

  • Faster
  • More memory efficient
  • Support vectorized operations

Vectorization means operations can be performed on entire arrays without writing explicit loops.

import numpy as np

arr = np.array([10, 20, 30, 40])
print(arr)
print(type(arr))       # <class 'numpy.ndarray'>
print(arr.dtype)       # int64

6.3 Common NumPy Array Creation Functions

FunctionPurposeExample
np.array()Creates an array from a list or tuplenp.array([1, 2, 3])
np.zeros()Creates an array filled with zerosnp.zeros(5)
np.ones()Creates an array filled with onesnp.ones((2, 3))
np.arange()Creates values in a rangenp.arange(1, 10, 2)
np.linspace()Creates evenly spaced valuesnp.linspace(0, 1, 5)
np.eye()Creates an identity matrixnp.eye(3)

6.4 Attributes of ndarray

import numpy as np

a = np.array([[1, 2, 3], [4, 5, 6]])

print(a.shape)   # (2, 3) — rows and columns
print(a.ndim)    # 2 — number of dimensions
print(a.size)    # 6 — total number of elements
print(a.dtype)   # int64 — data type of elements
AttributeMeaning
shapeReturns the dimensions of the array
ndimReturns the number of dimensions
sizeReturns the total number of elements
dtypeReturns the data type of elements

6.5 NumPy Workflow

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Create    │ →  │ 2. Index &   │ →  │ 3. Vectorized│ →  │ 4. Join /    │ →  │ 5. Search /  │
│    ndarray   │    │    Slice     │    │   Operations │    │    Split     │    │  Sort/Filter │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘

7. Array Indexing and Slicing

Indexing is used to access individual elements, while slicing is used to access a range or subset of elements. NumPy indexing starts from 0.

7.1 One-Dimensional Arrays

import numpy as np

a = np.array([10, 20, 30, 40, 50])

print(a[0])      # 10 — first element
print(a[-1])     # 50 — last element
print(a[1:4])    # [20, 30, 40] — elements from index 1 to 3
print(a[:3])     # [10, 20, 30] — first three elements
print(a[::2])    # [10, 30, 50] — every second element

7.2 Two-Dimensional Arrays

import numpy as np

matrix = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])

print(matrix[0, 0])       # 1 — row 0, column 0
print(matrix[1, 2])       # 6 — row 1, column 2
print(matrix[:, 1])       # [2, 5, 8] — all rows, column 1
print(matrix[0:2, 1:3])   # Sub-array

7.3 Indexing and Slicing Examples

ExpressionMeaning
a[0]First element of a one-dimensional array
a[-1]Last element of a one-dimensional array
a[1:4]Elements from index 1 to 3
matrix[1, 2]Element at row 1 and column 2
matrix[:, 0]All rows from the first column
matrix[0:2, 1:3]Sub-array using row and column ranges

8. Array Operations: Joining and Splitting

8.1 Vectorized Arithmetic Operations

import numpy as np

a = np.array([10, 20, 30])
b = np.array([1, 2, 3])

print(a + b)   # [11, 22, 33]
print(a - b)   # [9, 18, 27]
print(a * b)   # [10, 40, 90]
print(a / b)   # [10., 10., 10.]

8.2 Joining Arrays

import numpy as np

a = np.array([1, 2, 3])
b = np.array([4, 5, 6])

joined = np.concatenate((a, b))
print(joined)   # [1, 2, 3, 4, 5, 6]

# Joining 2D arrays
x = np.array([[1, 2], [3, 4]])
y = np.array([[5, 6], [7, 8]])

print(np.vstack((x, y)))   # Vertical stacking
print(np.hstack((x, y)))   # Horizontal stacking

8.3 Splitting Arrays

import numpy as np

a = np.array([1, 2, 3, 4, 5, 6])

parts = np.array_split(a, 3)
print(parts)   # [array([1, 2]), array([3, 4]), array([5, 6])]

8.4 Joining and Splitting Functions

FunctionPurpose
concatenate()Joins arrays along an existing axis
vstack()Stacks arrays vertically
hstack()Stacks arrays horizontally
array_split()Splits an array into multiple parts
reshape()Changes the shape of an array without changing data

9. Searching, Sorting, and Filtering Arrays

9.1 Searching Arrays

import numpy as np

a = np.array([10, 20, 30, 20, 40])
result = np.where(a == 20)
print(result)   # (array([1, 3]),)

9.2 Sorting Arrays

import numpy as np

a = np.array([40, 10, 30, 20])
print(np.sort(a))   # [10, 20, 30, 40]

9.3 Filtering Arrays

import numpy as np

a = np.array([10, 25, 30, 45, 50])

filtered = a[a > 30]
print(filtered)   # [45, 50]

9.4 Searching, Sorting, and Filtering Summary

OperationExampleOutput Meaning
Searchnp.where(a == 20)Returns indexes where condition is true
Sortnp.sort(a)Returns sorted copy of the array
Filtera[a > 30]Returns values greater than 30
Boolean maska % 2 == 0Creates True/False condition for each value

Important: Filtering in NumPy is usually done using boolean conditions. The condition creates a True/False mask, and the mask is used to select matching elements.


10. Random Number Generation in NumPy

Random numbers are used in simulations, sampling, testing, data generation, machine learning, and probability experiments.

10.1 Basic Random Functions

import numpy as np

print(np.random.randint(1, 10))        # One random integer
print(np.random.rand(3))               # 3 random floats from 0 to 1
print(np.random.randint(1, 100, 5))    # 5 random integers

10.2 Random Arrays and Reproducibility

A seed is used to produce the same random results again. This is useful for teaching, testing, and experiments where reproducibility is important.

import numpy as np

np.random.seed(10)
print(np.random.randint(1, 100, 5))   # Same output every time

10.3 Random Number Generation Functions

FunctionPurpose
rand()Generates random floats between 0 and 1
randint()Generates random integers within a range
choice()Selects random values from a list or array

11. Introduction to Web Scraping

11.1 What is Web Scraping?

Web scraping is the process of extracting data from websites using programs. It is useful when information is available on web pages but not provided as a downloadable dataset or API.

11.2 Basic Web Scraping Workflow

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Send HTTP │ →  │ 2. Download  │ →  │ 3. Parse and │ →  │ 4. Clean and │ →  │ 5. Store or  │
│   Request    │    │    HTML      │    │ Extract Data │    │ Structure    │    │ Export Data  │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘

11.3 Steps in a Basic Web Scraping Workflow

StepMeaningCommon Tool
Request webpageDownload the webpage contentrequests
Read HTMLReceive page source code as textresponse.text
Parse HTMLUnderstand tags and structureBeautifulSoup
Extract elementsFind headings, links, tables, or pricesfind(), find_all(), select()
Store dataSave extracted data for analysisCSV, JSON, Pandas DataFrame

11.4 Simple Web Scraping Example

import requests
from bs4 import BeautifulSoup

url = "https://example.com"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")

print(soup.title.text)

11.5 Ethical Precautions

Web scraping should be performed responsibly. Programmers must respect:

  • Website terms of service
  • robots.txt rules
  • Copyright
  • Privacy
  • Server load

Scraping should not be used to collect personal or restricted data without permission.

Ethical Reminder: Before scraping a website, check whether the website allows automated access. Prefer official APIs whenever available.


12. Introduction to Matplotlib

12.1 What is Matplotlib?

Matplotlib is a Python library used for creating visualizations. It allows programmers to create line graphs, bar charts, histograms, scatter plots, pie charts, and many other types of plots.

The most commonly used Matplotlib module is pyplot, usually imported as plt. It provides functions similar to plotting commands used in MATLAB and other scientific tools.

12.2 Why Visualize Data?

Visualization helps convert numbers into patterns. A good chart can reveal:

  • Trends
  • Comparisons
  • Distributions
  • Relationships

These may not be clear from a table alone.

12.3 Matplotlib Visualization Workflow

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Prepare   │ →  │ 2. Create    │ →  │ 3. Plot Data │ →  │ 4. Customize │ →  │ 5. Show or   │
│    Data      │    │   Figure     │    │              │    │    Plot      │    │    Save      │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘
  1. Prepare Data: Load and organize data for visualization.
  2. Create Figure: Create a figure and one or more axes.
  3. Plot Data: Plot data using the appropriate plot type.
  4. Customize Plot: Add titles, labels, legends, colors, and other styling.
  5. Show or Save: Display the plot on screen or save it to a file.

12.4 Basic Line Plot

import matplotlib.pyplot as plt

x = [1, 2, 3, 4, 5]
y = [10, 15, 13, 18, 20]

plt.plot(x, y)
plt.title("Simple Line Plot")
plt.xlabel("X Values")
plt.ylabel("Y Values")
plt.show()

12.5 Common Matplotlib pyplot Functions

FunctionPurpose
plot()Creates a line plot
bar()Creates a bar chart
hist()Creates a histogram
scatter()Creates a scatter plot
pie()Creates a pie chart
title()Adds chart title
xlabel(), ylabel()Adds axis labels
legend()Displays chart legend
show()Displays the figure
savefig()Saves chart as an image file

13. Data Visualization Using Matplotlib

13.1 Choosing the Correct Visualization

Chart TypeUsed ForExample Question
Line plotShowing trends over timeHow did sales change month by month?
Bar chartComparing categoriesWhich department scored highest?
HistogramShowing distribution of numerical valuesWhat is the distribution of exam marks?
Scatter plotShowing relationship between two numerical variablesIs study time related to marks?
Pie chartShowing parts of a wholeWhat percentage of students selected each elective?

13.2 Bar Chart Example

import matplotlib.pyplot as plt

subjects = ["Python", "Maths", "DBMS", "OS"]
marks = [85, 78, 92, 74]

plt.bar(subjects, marks)
plt.title("Marks by Subject")
plt.xlabel("Subject")
plt.ylabel("Marks")
plt.show()

13.3 Histogram Example

import matplotlib.pyplot as plt

marks = [45, 56, 67, 78, 89, 90, 72, 63, 55, 81]

plt.hist(marks, bins=5)
plt.title("Distribution of Marks")
plt.xlabel("Marks Range")
plt.ylabel("Number of Students")
plt.show()

13.4 Scatter Plot Example

import matplotlib.pyplot as plt

hours = [1, 2, 3, 4, 5, 6]
marks = [40, 50, 55, 65, 75, 85]

plt.scatter(hours, marks)
plt.title("Study Hours vs Marks")
plt.xlabel("Study Hours")
plt.ylabel("Marks")
plt.show()

13.5 Pie Chart Example

import matplotlib.pyplot as plt

labels = ["Python", "Java", "C++", "R"]
students = [40, 25, 20, 15]

plt.pie(students, labels=labels, autopct="%1.1f%%")
plt.title("Programming Language Preference")
plt.show()

13.6 Visualization Rule

Every chart should have:

  • A clear title
  • Axis labels where applicable
  • Readable category names
  • A clear purpose

Avoid adding unnecessary decoration.


14. Integrated Practical Example

The following example shows how Pandas, NumPy, and Matplotlib can be used together.

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# Create sample data
students = pd.DataFrame({
    "Name": ["Amit", "Ravi", "Neha", "Sara", "John"],
    "Python": [78, 85, 92, 66, 74],
    "Maths": [80, 70, 88, 60, 79]
})

# Calculate total and average using Pandas/NumPy
students["Total"] = students["Python"] + students["Maths"]
students["Average"] = np.round(students["Total"] / 2, 2)

print(students)
print(students.describe())

# Plot average marks
plt.bar(students["Name"], students["Average"])
plt.title("Average Marks of Students")
plt.xlabel("Student")
plt.ylabel("Average Marks")
plt.show()

Role of Each Library:

LibraryRole in the Example
PandasCreates and manages the student DataFrame
NumPyRounds numerical average values
MatplotlibVisualizes average marks as a bar chart

15. Comparison: Pandas vs NumPy vs Matplotlib

DimensionPandasNumPyMatplotlib
PurposeTabular data analysisNumerical computationData visualization
Data StructureSeries, DataFramendarrayFigure, Axes
Common Functionsread_csv(), merge(), groupby()array(), sort(), where()plot(), bar(), hist()
Input Data TypeCSV, JSON, Excel, SQLLists, tuples, arraysLists, arrays, Series
Output TypeDataFrame, SeriesndarrayCharts, plots
SpeedModerateVery fastModerate
Use CaseData cleaning, analysisMathematical operationsVisual communication
Example Codedf.head()np.array([1,2,3])plt.plot(x, y)

How They Work Together:

  1. Pandas loads and cleans data from CSV/JSON.
  2. NumPy performs fast numerical operations on the data.
  3. Matplotlib visualizes the results for communication.

16. Unit Summary

PANDAS, NUMPY, AND MATPLOTLIB
│
├── Introduction
│   ├── Pandas → Tabular data
│   ├── NumPy → Numerical computation
│   └── Matplotlib → Visualization
│
├── Pandas
│   ├── Series (1D labelled array)
│   ├── DataFrame (2D labelled table)
│   ├── Inspection: head(), tail(), shape, info(), describe()
│   ├── Selection: df["col"], loc[], iloc[]
│   ├── Cleaning: dropna(), fillna(), drop_duplicates()
│   ├── Combining: concat(), merge(), join()
│   └── I/O: read_csv(), to_csv(), read_json(), to_json()
│
├── NumPy
│   ├── ndarray object
│   ├── Creation: array(), zeros(), ones(), arange(), linspace()
│   ├── Attributes: shape, ndim, size, dtype
│   ├── Indexing and slicing
│   ├── Operations: concatenate(), vstack(), hstack(), array_split()
│   ├── Search: where()
│   ├── Sort: sort()
│   ├── Filter: boolean masks
│   └── Random: rand(), randint(), choice(), seed()
│
├── Web Scraping
│   ├── Workflow: Request → Parse → Extract → Store
│   ├── Tools: requests, BeautifulSoup
│   └── Ethics: robots.txt, terms of service, privacy
│
└── Matplotlib
    ├── pyplot module
    ├── Chart types: line, bar, histogram, scatter, pie
    ├── Customization: title, labels, legend
    └── Workflow: Prepare → Create → Plot → Customize → Show/Save

17. Key Syntax — Quick Reference

ConceptSyntax/Example
Import Pandasimport pandas as pd
Import NumPyimport numpy as np
Import Matplotlibimport matplotlib.pyplot as plt
Create Seriespd.Series([1, 2, 3], index=["a", "b", "c"])
Create DataFramepd.DataFrame({"A": [1, 2], "B": [3, 4]})
Read CSVpd.read_csv("file.csv")
Write CSVdf.to_csv("file.csv", index=False)
Read JSONpd.read_json("file.json")
Write JSONdf.to_json("file.json")
Inspect DataFramedf.head(), df.info(), df.describe()
Select columndf["Name"]
Select rows by labeldf.loc[0]
Select rows by positiondf.iloc[0:3]
Filter rowsdf[df["Marks"] > 80]
Handle missingdf.dropna(), df.fillna(0)
Combine DataFramespd.concat([df1, df2]), pd.merge(a, b, on="ID")
Create NumPy arraynp.array([1, 2, 3])
Array attributesa.shape, a.ndim, a.size, a.dtype
Array indexinga[0], a[-1], a[1:4]
2D indexingmatrix[0, 1], matrix[:, 0]
Join arraysnp.concatenate((a, b))
Split arraysnp.array_split(a, 3)
Search arraynp.where(a == 20)
Sort arraynp.sort(a)
Filter arraya[a > 30]
Random integernp.random.randint(1, 100, 5)
Set seednp.random.seed(10)
Line plotplt.plot(x, y)
Bar chartplt.bar(categories, values)
Histogramplt.hist(data, bins=5)
Scatter plotplt.scatter(x, y)
Pie chartplt.pie(values, labels=labels, autopct="%1.1f%%")
Add titleplt.title("Title")
Add labelsplt.xlabel("X"), plt.ylabel("Y")
Show plotplt.show()
Save plotplt.savefig("plot.png")

On this page

1. Introduction to Data Science Libraries2. Introduction to Pandas2.1 What is Pandas?2.2 Why Use Pandas?2.3 Important Features of Pandas2.4 Installing and Importing Pandas2.5 Pandas Workflow3. Pandas Series and DataFrames3.1 Pandas Series3.2 Pandas DataFrame3.3 Common DataFrame Inspection Methods4. Manipulating and Combining DataFrames4.1 Selecting Columns and Rows4.2 Filtering and Creating Columns4.3 Handling Missing Values4.4 Combining DataFrames5. Reading CSV and JSON Files Using Pandas5.1 Reading CSV Files5.2 Writing CSV Files5.3 Reading JSON Files5.4 File Reading and Writing Functions6. NumPy Arrays and ndarray Objects6.1 What is NumPy?6.2 Why Use NumPy?6.3 Common NumPy Array Creation Functions6.4 Attributes of ndarray6.5 NumPy Workflow7. Array Indexing and Slicing7.1 One-Dimensional Arrays7.2 Two-Dimensional Arrays7.3 Indexing and Slicing Examples8. Array Operations: Joining and Splitting8.1 Vectorized Arithmetic Operations8.2 Joining Arrays8.3 Splitting Arrays8.4 Joining and Splitting Functions9. Searching, Sorting, and Filtering Arrays9.1 Searching Arrays9.2 Sorting Arrays9.3 Filtering Arrays9.4 Searching, Sorting, and Filtering Summary10. Random Number Generation in NumPy10.1 Basic Random Functions10.2 Random Arrays and Reproducibility10.3 Random Number Generation Functions11. Introduction to Web Scraping11.1 What is Web Scraping?11.2 Basic Web Scraping Workflow11.3 Steps in a Basic Web Scraping Workflow11.4 Simple Web Scraping Example11.5 Ethical Precautions12. Introduction to Matplotlib12.1 What is Matplotlib?12.2 Why Visualize Data?12.3 Matplotlib Visualization Workflow12.4 Basic Line Plot12.5 Common Matplotlib pyplot Functions13. Data Visualization Using Matplotlib13.1 Choosing the Correct Visualization13.2 Bar Chart Example13.3 Histogram Example13.4 Scatter Plot Example13.5 Pie Chart Example13.6 Visualization Rule14. Integrated Practical Example15. Comparison: Pandas vs NumPy vs Matplotlib16. Unit Summary17. Key Syntax — Quick Reference