Adv-Python Unit 4: Complete Concepts Guide
Unit IV: Pandas, NumPy, and Matplotlib -> Generated and Prepared By Thiruselvan (ThiruXD)
1. Introduction to Data Science Libraries
Python is widely used in data science because it provides simple syntax and powerful libraries for working with data. Three of the most important libraries for beginners are:
| Library | Purpose | Primary Use |
|---|---|---|
| Pandas | Tabular data handling | Data cleaning, analysis, manipulation |
| NumPy | Numerical computation | Fast array-based computing |
| Matplotlib | Data visualization | Creating charts and graphs |
Key Insight: These three libraries form the foundation of Python data analysis. Pandas handles structured data, NumPy handles numerical operations, and Matplotlib communicates results visually.
2. Introduction to Pandas
2.1 What is Pandas?
Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures and functions for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs.
The name Pandas is derived from Panel Data, which refers to multidimensional structured datasets.
2.2 Why Use Pandas?
Pandas is especially useful when data contains:
- Rows and columns
- Missing values
- Labels
- Mixed data types
- Dates
- Categories
- Numerical values
Pandas is widely used by data analysts, data scientists, researchers, and software developers because it makes data handling easier than using basic Python lists and dictionaries.
2.3 Important Features of Pandas
| Feature | Explanation | Example |
|---|---|---|
| Data loading | Reads data from files and data sources | read_csv(), read_json(), read_excel() |
| Data cleaning | Handles missing values, duplicates, incorrect formats | dropna(), fillna(), drop_duplicates() |
| Data selection | Selects rows, columns, and subsets | df["Name"], loc[], iloc[] |
| Data transformation | Creates new columns and modifies existing data | df["Total"] = df["A"] + df["B"] |
| Data aggregation | Groups and summarizes data | groupby(), mean(), sum() |
| Data export | Writes results back to files | to_csv(), to_json() |
2.4 Installing and Importing Pandas
In most data science environments (Anaconda, Google Colab, Jupyter Notebook), Pandas is already available. In a local environment, install using pip.
# Install Pandas
pip install pandas# Standard import convention
import pandas as pd2.5 Pandas Workflow
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Import │ → │ 2. Create │ → │ 3. Clean & │ → │ 4. Combine │ → │ 5. Analyze │
│ Data │ │ Series/DF │ │ Transform │ │ Datasets │ │ & Export │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘- Import Data: Load data from various sources.
- Create Series/DataFrame: Organize data into Series or DataFrame structures.
- Clean & Transform: Handle missing values, filter, sort, and transform data.
- Combine Datasets: Merge, join, or concatenate multiple datasets.
- Analyze & Export: Analyze data to extract insights and export results.
3. Pandas Series and DataFrames
3.1 Pandas Series
A Series is a one-dimensional labelled array. It can store integers, floats, strings, or other Python objects. A Series is similar to a single column in a spreadsheet.
import pandas as pd
marks = pd.Series([78, 85, 92, 66], index=["Amit", "Ravi", "Neha", "Sara"])
print(marks)
print(marks["Neha"]) # 92Components of a Pandas Series:
| Component | Meaning |
|---|---|
| Values | The actual data stored in the Series |
| Index | Labels used to identify each value |
| Data type | The type of values stored (int64, float64, object) |
3.2 Pandas DataFrame
A DataFrame is a two-dimensional labelled data structure with rows and columns. It is the most commonly used data structure in Pandas. DataFrames can be created from:
- Dictionaries
- Lists
- CSV files
- JSON files
- Excel files
- Databases
- APIs
import pandas as pd
data = {
"Name": ["Amit", "Ravi", "Neha", "Sara"],
"Age": [21, 22, 20, 23],
"Marks": [78, 85, 92, 66]
}
df = pd.DataFrame(data)
print(df)Output:
Name Age Marks
0 Amit 21 78
1 Ravi 22 85
2 Neha 20 92
3 Sara 23 663.3 Common DataFrame Inspection Methods
| Operation | Purpose | Example |
|---|---|---|
head() | Displays first five rows by default | df.head() |
tail() | Displays last five rows by default | df.tail() |
shape | Shows number of rows and columns | df.shape |
columns | Shows column names | df.columns |
info() | Displays summary of columns and data types | df.info() |
describe() | Displays statistical summary of numerical columns | df.describe() |
4. Manipulating and Combining DataFrames
Data manipulation means changing, cleaning, selecting, filtering, sorting, grouping, or combining data. In real data analysis, raw data is rarely ready for direct use.
4.1 Selecting Columns and Rows
# Select one column
print(df["Name"])
# Select multiple columns
print(df[["Name", "Marks"]])
# Select rows using label-based indexing
print(df.loc[0])
# Select rows using integer position
print(df.iloc[0:2])| Method | Used For | Example |
|---|---|---|
df["col"] | Selecting a single column | df["Marks"] |
df[["col1", "col2"]] | Selecting multiple columns | df[["Name", "Age"]] |
loc[] | Selecting by labels or conditions | df.loc[df["Marks"] > 80] |
iloc[] | Selecting by integer position | df.iloc[0:3, 1:3] |
4.2 Filtering and Creating Columns
# Filter rows
high_scorers = df[df["Marks"] > 80]
print(high_scorers)
# Create new column
df["Total"] = df["Marks"] + 10
print(df)4.3 Handling Missing Values
Missing values are common in real datasets. Pandas represents missing values using NaN (Not a Number).
import pandas as pd
data = {"Name": ["Amit", "Ravi", "Neha"], "Marks": [78, None, 92]}
df = pd.DataFrame(data)
print(df.isnull()) # Check missing values
print(df.dropna()) # Remove rows with missing values
print(df.fillna(0)) # Replace missing values with 0Data Cleaning Functions:
| Function | Purpose |
|---|---|
isnull() | Detects missing values |
notnull() | Detects non-missing values |
dropna() | Removes rows or columns with missing values |
fillna(value) | Replaces missing values with a specified value |
drop_duplicates() | Removes duplicate rows |
4.4 Combining DataFrames
Combining DataFrames is required when data is stored in multiple tables. Pandas supports concatenation, merging, and joining.
Concatenation (stacking):
import pandas as pd
df1 = pd.DataFrame({"ID": [1, 2], "Name": ["Amit", "Ravi"]})
df2 = pd.DataFrame({"ID": [3, 4], "Name": ["Neha", "Sara"]})
combined = pd.concat([df1, df2], ignore_index=True)
print(combined)Merging (common column):
marks = pd.DataFrame({"ID": [1, 2, 3], "Marks": [78, 85, 92]})
students = pd.DataFrame({"ID": [1, 2, 3], "Name": ["Amit", "Ravi", "Neha"]})
result = pd.merge(students, marks, on="ID")
print(result)DataFrame Combining Methods:
| Operation | Meaning | When to Use |
|---|---|---|
concat() | Stacks DataFrames row-wise or column-wise | When datasets have same columns or same index |
merge() | Combines DataFrames using a common column | When tables share a key such as ID |
join() | Combines DataFrames using index | When index labels are meaningful |
5. Reading CSV and JSON Files Using Pandas
Pandas can read data from many file formats. Two of the most common formats are CSV and JSON.
5.1 Reading CSV Files
import pandas as pd
# Read a CSV file
students = pd.read_csv("students.csv")
print(students.head())
print(students.shape)
print(students.info())5.2 Writing CSV Files
# Save a DataFrame to a CSV file
students.to_csv("cleaned_students.csv", index=False)5.3 Reading JSON Files
import pandas as pd
# Read a JSON file
orders = pd.read_json("orders.json")
print(orders.head())5.4 File Reading and Writing Functions
| Function | Purpose | Common Parameter |
|---|---|---|
read_csv() | Reads a CSV file into a DataFrame | sep, header, names, usecols |
to_csv() | Writes a DataFrame to a CSV file | index=False |
read_json() | Reads a JSON file into a DataFrame | orient |
to_json() | Writes a DataFrame to a JSON file | orient, indent |
Practical Tip: After reading any file, always use head(), shape, info(), and describe() to understand the structure and quality of the imported data.
6. NumPy Arrays and ndarray Objects
6.1 What is NumPy?
NumPy stands for Numerical Python. It is a fundamental Python library for scientific computing and numerical operations. NumPy provides the ndarray object, which stores elements of the same data type in a compact and efficient way.
6.2 Why Use NumPy?
Compared with normal Python lists, NumPy arrays are:
- Faster
- More memory efficient
- Support vectorized operations
Vectorization means operations can be performed on entire arrays without writing explicit loops.
import numpy as np
arr = np.array([10, 20, 30, 40])
print(arr)
print(type(arr)) # <class 'numpy.ndarray'>
print(arr.dtype) # int646.3 Common NumPy Array Creation Functions
| Function | Purpose | Example |
|---|---|---|
np.array() | Creates an array from a list or tuple | np.array([1, 2, 3]) |
np.zeros() | Creates an array filled with zeros | np.zeros(5) |
np.ones() | Creates an array filled with ones | np.ones((2, 3)) |
np.arange() | Creates values in a range | np.arange(1, 10, 2) |
np.linspace() | Creates evenly spaced values | np.linspace(0, 1, 5) |
np.eye() | Creates an identity matrix | np.eye(3) |
6.4 Attributes of ndarray
import numpy as np
a = np.array([[1, 2, 3], [4, 5, 6]])
print(a.shape) # (2, 3) — rows and columns
print(a.ndim) # 2 — number of dimensions
print(a.size) # 6 — total number of elements
print(a.dtype) # int64 — data type of elements| Attribute | Meaning |
|---|---|
shape | Returns the dimensions of the array |
ndim | Returns the number of dimensions |
size | Returns the total number of elements |
dtype | Returns the data type of elements |
6.5 NumPy Workflow
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Create │ → │ 2. Index & │ → │ 3. Vectorized│ → │ 4. Join / │ → │ 5. Search / │
│ ndarray │ │ Slice │ │ Operations │ │ Split │ │ Sort/Filter │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘7. Array Indexing and Slicing
Indexing is used to access individual elements, while slicing is used to access a range or subset of elements. NumPy indexing starts from 0.
7.1 One-Dimensional Arrays
import numpy as np
a = np.array([10, 20, 30, 40, 50])
print(a[0]) # 10 — first element
print(a[-1]) # 50 — last element
print(a[1:4]) # [20, 30, 40] — elements from index 1 to 3
print(a[:3]) # [10, 20, 30] — first three elements
print(a[::2]) # [10, 30, 50] — every second element7.2 Two-Dimensional Arrays
import numpy as np
matrix = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
print(matrix[0, 0]) # 1 — row 0, column 0
print(matrix[1, 2]) # 6 — row 1, column 2
print(matrix[:, 1]) # [2, 5, 8] — all rows, column 1
print(matrix[0:2, 1:3]) # Sub-array7.3 Indexing and Slicing Examples
| Expression | Meaning |
|---|---|
a[0] | First element of a one-dimensional array |
a[-1] | Last element of a one-dimensional array |
a[1:4] | Elements from index 1 to 3 |
matrix[1, 2] | Element at row 1 and column 2 |
matrix[:, 0] | All rows from the first column |
matrix[0:2, 1:3] | Sub-array using row and column ranges |
8. Array Operations: Joining and Splitting
8.1 Vectorized Arithmetic Operations
import numpy as np
a = np.array([10, 20, 30])
b = np.array([1, 2, 3])
print(a + b) # [11, 22, 33]
print(a - b) # [9, 18, 27]
print(a * b) # [10, 40, 90]
print(a / b) # [10., 10., 10.]8.2 Joining Arrays
import numpy as np
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
joined = np.concatenate((a, b))
print(joined) # [1, 2, 3, 4, 5, 6]
# Joining 2D arrays
x = np.array([[1, 2], [3, 4]])
y = np.array([[5, 6], [7, 8]])
print(np.vstack((x, y))) # Vertical stacking
print(np.hstack((x, y))) # Horizontal stacking8.3 Splitting Arrays
import numpy as np
a = np.array([1, 2, 3, 4, 5, 6])
parts = np.array_split(a, 3)
print(parts) # [array([1, 2]), array([3, 4]), array([5, 6])]8.4 Joining and Splitting Functions
| Function | Purpose |
|---|---|
concatenate() | Joins arrays along an existing axis |
vstack() | Stacks arrays vertically |
hstack() | Stacks arrays horizontally |
array_split() | Splits an array into multiple parts |
reshape() | Changes the shape of an array without changing data |
9. Searching, Sorting, and Filtering Arrays
9.1 Searching Arrays
import numpy as np
a = np.array([10, 20, 30, 20, 40])
result = np.where(a == 20)
print(result) # (array([1, 3]),)9.2 Sorting Arrays
import numpy as np
a = np.array([40, 10, 30, 20])
print(np.sort(a)) # [10, 20, 30, 40]9.3 Filtering Arrays
import numpy as np
a = np.array([10, 25, 30, 45, 50])
filtered = a[a > 30]
print(filtered) # [45, 50]9.4 Searching, Sorting, and Filtering Summary
| Operation | Example | Output Meaning |
|---|---|---|
| Search | np.where(a == 20) | Returns indexes where condition is true |
| Sort | np.sort(a) | Returns sorted copy of the array |
| Filter | a[a > 30] | Returns values greater than 30 |
| Boolean mask | a % 2 == 0 | Creates True/False condition for each value |
Important: Filtering in NumPy is usually done using boolean conditions. The condition creates a True/False mask, and the mask is used to select matching elements.
10. Random Number Generation in NumPy
Random numbers are used in simulations, sampling, testing, data generation, machine learning, and probability experiments.
10.1 Basic Random Functions
import numpy as np
print(np.random.randint(1, 10)) # One random integer
print(np.random.rand(3)) # 3 random floats from 0 to 1
print(np.random.randint(1, 100, 5)) # 5 random integers10.2 Random Arrays and Reproducibility
A seed is used to produce the same random results again. This is useful for teaching, testing, and experiments where reproducibility is important.
import numpy as np
np.random.seed(10)
print(np.random.randint(1, 100, 5)) # Same output every time10.3 Random Number Generation Functions
| Function | Purpose |
|---|---|
rand() | Generates random floats between 0 and 1 |
randint() | Generates random integers within a range |
choice() | Selects random values from a list or array |
11. Introduction to Web Scraping
11.1 What is Web Scraping?
Web scraping is the process of extracting data from websites using programs. It is useful when information is available on web pages but not provided as a downloadable dataset or API.
11.2 Basic Web Scraping Workflow
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Send HTTP │ → │ 2. Download │ → │ 3. Parse and │ → │ 4. Clean and │ → │ 5. Store or │
│ Request │ │ HTML │ │ Extract Data │ │ Structure │ │ Export Data │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘11.3 Steps in a Basic Web Scraping Workflow
| Step | Meaning | Common Tool |
|---|---|---|
| Request webpage | Download the webpage content | requests |
| Read HTML | Receive page source code as text | response.text |
| Parse HTML | Understand tags and structure | BeautifulSoup |
| Extract elements | Find headings, links, tables, or prices | find(), find_all(), select() |
| Store data | Save extracted data for analysis | CSV, JSON, Pandas DataFrame |
11.4 Simple Web Scraping Example
import requests
from bs4 import BeautifulSoup
url = "https://example.com"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")
print(soup.title.text)11.5 Ethical Precautions
Web scraping should be performed responsibly. Programmers must respect:
- Website terms of service
- robots.txt rules
- Copyright
- Privacy
- Server load
Scraping should not be used to collect personal or restricted data without permission.
Ethical Reminder: Before scraping a website, check whether the website allows automated access. Prefer official APIs whenever available.
12. Introduction to Matplotlib
12.1 What is Matplotlib?
Matplotlib is a Python library used for creating visualizations. It allows programmers to create line graphs, bar charts, histograms, scatter plots, pie charts, and many other types of plots.
The most commonly used Matplotlib module is pyplot, usually imported as plt. It provides functions similar to plotting commands used in MATLAB and other scientific tools.
12.2 Why Visualize Data?
Visualization helps convert numbers into patterns. A good chart can reveal:
- Trends
- Comparisons
- Distributions
- Relationships
These may not be clear from a table alone.
12.3 Matplotlib Visualization Workflow
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Prepare │ → │ 2. Create │ → │ 3. Plot Data │ → │ 4. Customize │ → │ 5. Show or │
│ Data │ │ Figure │ │ │ │ Plot │ │ Save │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘- Prepare Data: Load and organize data for visualization.
- Create Figure: Create a figure and one or more axes.
- Plot Data: Plot data using the appropriate plot type.
- Customize Plot: Add titles, labels, legends, colors, and other styling.
- Show or Save: Display the plot on screen or save it to a file.
12.4 Basic Line Plot
import matplotlib.pyplot as plt
x = [1, 2, 3, 4, 5]
y = [10, 15, 13, 18, 20]
plt.plot(x, y)
plt.title("Simple Line Plot")
plt.xlabel("X Values")
plt.ylabel("Y Values")
plt.show()12.5 Common Matplotlib pyplot Functions
| Function | Purpose |
|---|---|
plot() | Creates a line plot |
bar() | Creates a bar chart |
hist() | Creates a histogram |
scatter() | Creates a scatter plot |
pie() | Creates a pie chart |
title() | Adds chart title |
xlabel(), ylabel() | Adds axis labels |
legend() | Displays chart legend |
show() | Displays the figure |
savefig() | Saves chart as an image file |
13. Data Visualization Using Matplotlib
13.1 Choosing the Correct Visualization
| Chart Type | Used For | Example Question |
|---|---|---|
| Line plot | Showing trends over time | How did sales change month by month? |
| Bar chart | Comparing categories | Which department scored highest? |
| Histogram | Showing distribution of numerical values | What is the distribution of exam marks? |
| Scatter plot | Showing relationship between two numerical variables | Is study time related to marks? |
| Pie chart | Showing parts of a whole | What percentage of students selected each elective? |
13.2 Bar Chart Example
import matplotlib.pyplot as plt
subjects = ["Python", "Maths", "DBMS", "OS"]
marks = [85, 78, 92, 74]
plt.bar(subjects, marks)
plt.title("Marks by Subject")
plt.xlabel("Subject")
plt.ylabel("Marks")
plt.show()13.3 Histogram Example
import matplotlib.pyplot as plt
marks = [45, 56, 67, 78, 89, 90, 72, 63, 55, 81]
plt.hist(marks, bins=5)
plt.title("Distribution of Marks")
plt.xlabel("Marks Range")
plt.ylabel("Number of Students")
plt.show()13.4 Scatter Plot Example
import matplotlib.pyplot as plt
hours = [1, 2, 3, 4, 5, 6]
marks = [40, 50, 55, 65, 75, 85]
plt.scatter(hours, marks)
plt.title("Study Hours vs Marks")
plt.xlabel("Study Hours")
plt.ylabel("Marks")
plt.show()13.5 Pie Chart Example
import matplotlib.pyplot as plt
labels = ["Python", "Java", "C++", "R"]
students = [40, 25, 20, 15]
plt.pie(students, labels=labels, autopct="%1.1f%%")
plt.title("Programming Language Preference")
plt.show()13.6 Visualization Rule
Every chart should have:
- A clear title
- Axis labels where applicable
- Readable category names
- A clear purpose
Avoid adding unnecessary decoration.
14. Integrated Practical Example
The following example shows how Pandas, NumPy, and Matplotlib can be used together.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# Create sample data
students = pd.DataFrame({
"Name": ["Amit", "Ravi", "Neha", "Sara", "John"],
"Python": [78, 85, 92, 66, 74],
"Maths": [80, 70, 88, 60, 79]
})
# Calculate total and average using Pandas/NumPy
students["Total"] = students["Python"] + students["Maths"]
students["Average"] = np.round(students["Total"] / 2, 2)
print(students)
print(students.describe())
# Plot average marks
plt.bar(students["Name"], students["Average"])
plt.title("Average Marks of Students")
plt.xlabel("Student")
plt.ylabel("Average Marks")
plt.show()Role of Each Library:
| Library | Role in the Example |
|---|---|
| Pandas | Creates and manages the student DataFrame |
| NumPy | Rounds numerical average values |
| Matplotlib | Visualizes average marks as a bar chart |
15. Comparison: Pandas vs NumPy vs Matplotlib
| Dimension | Pandas | NumPy | Matplotlib |
|---|---|---|---|
| Purpose | Tabular data analysis | Numerical computation | Data visualization |
| Data Structure | Series, DataFrame | ndarray | Figure, Axes |
| Common Functions | read_csv(), merge(), groupby() | array(), sort(), where() | plot(), bar(), hist() |
| Input Data Type | CSV, JSON, Excel, SQL | Lists, tuples, arrays | Lists, arrays, Series |
| Output Type | DataFrame, Series | ndarray | Charts, plots |
| Speed | Moderate | Very fast | Moderate |
| Use Case | Data cleaning, analysis | Mathematical operations | Visual communication |
| Example Code | df.head() | np.array([1,2,3]) | plt.plot(x, y) |
How They Work Together:
- Pandas loads and cleans data from CSV/JSON.
- NumPy performs fast numerical operations on the data.
- Matplotlib visualizes the results for communication.
16. Unit Summary
PANDAS, NUMPY, AND MATPLOTLIB
│
├── Introduction
│ ├── Pandas → Tabular data
│ ├── NumPy → Numerical computation
│ └── Matplotlib → Visualization
│
├── Pandas
│ ├── Series (1D labelled array)
│ ├── DataFrame (2D labelled table)
│ ├── Inspection: head(), tail(), shape, info(), describe()
│ ├── Selection: df["col"], loc[], iloc[]
│ ├── Cleaning: dropna(), fillna(), drop_duplicates()
│ ├── Combining: concat(), merge(), join()
│ └── I/O: read_csv(), to_csv(), read_json(), to_json()
│
├── NumPy
│ ├── ndarray object
│ ├── Creation: array(), zeros(), ones(), arange(), linspace()
│ ├── Attributes: shape, ndim, size, dtype
│ ├── Indexing and slicing
│ ├── Operations: concatenate(), vstack(), hstack(), array_split()
│ ├── Search: where()
│ ├── Sort: sort()
│ ├── Filter: boolean masks
│ └── Random: rand(), randint(), choice(), seed()
│
├── Web Scraping
│ ├── Workflow: Request → Parse → Extract → Store
│ ├── Tools: requests, BeautifulSoup
│ └── Ethics: robots.txt, terms of service, privacy
│
└── Matplotlib
├── pyplot module
├── Chart types: line, bar, histogram, scatter, pie
├── Customization: title, labels, legend
└── Workflow: Prepare → Create → Plot → Customize → Show/Save17. Key Syntax — Quick Reference
| Concept | Syntax/Example |
|---|---|
| Import Pandas | import pandas as pd |
| Import NumPy | import numpy as np |
| Import Matplotlib | import matplotlib.pyplot as plt |
| Create Series | pd.Series([1, 2, 3], index=["a", "b", "c"]) |
| Create DataFrame | pd.DataFrame({"A": [1, 2], "B": [3, 4]}) |
| Read CSV | pd.read_csv("file.csv") |
| Write CSV | df.to_csv("file.csv", index=False) |
| Read JSON | pd.read_json("file.json") |
| Write JSON | df.to_json("file.json") |
| Inspect DataFrame | df.head(), df.info(), df.describe() |
| Select column | df["Name"] |
| Select rows by label | df.loc[0] |
| Select rows by position | df.iloc[0:3] |
| Filter rows | df[df["Marks"] > 80] |
| Handle missing | df.dropna(), df.fillna(0) |
| Combine DataFrames | pd.concat([df1, df2]), pd.merge(a, b, on="ID") |
| Create NumPy array | np.array([1, 2, 3]) |
| Array attributes | a.shape, a.ndim, a.size, a.dtype |
| Array indexing | a[0], a[-1], a[1:4] |
| 2D indexing | matrix[0, 1], matrix[:, 0] |
| Join arrays | np.concatenate((a, b)) |
| Split arrays | np.array_split(a, 3) |
| Search array | np.where(a == 20) |
| Sort array | np.sort(a) |
| Filter array | a[a > 30] |
| Random integer | np.random.randint(1, 100, 5) |
| Set seed | np.random.seed(10) |
| Line plot | plt.plot(x, y) |
| Bar chart | plt.bar(categories, values) |
| Histogram | plt.hist(data, bins=5) |
| Scatter plot | plt.scatter(x, y) |
| Pie chart | plt.pie(values, labels=labels, autopct="%1.1f%%") |
| Add title | plt.title("Title") |
| Add labels | plt.xlabel("X"), plt.ylabel("Y") |
| Show plot | plt.show() |
| Save plot | plt.savefig("plot.png") |