Adv-Py: Unit 4 (Book-back Questions)
Generated and Prepared By Thiruselvan (ThiruXD)
PART A — MULTIPLE CHOICE QUESTIONS (1 Mark Each)
Q1. Which library is mainly used for tabular data analysis in Python?
- NumPy
- Pandas
- Matplotlib
- Requests
Answer: (b) Pandas
Explanation: Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures (Series and DataFrame) for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs. NumPy is for numerical computation; Matplotlib is for visualization; Requests is for HTTP requests.
Q2. A Pandas Series is best described as:
- A two-dimensional table
- A one-dimensional labelled array
- A plotting function
- A web scraping tool
Answer: (b) A one-dimensional labelled array
Explanation: A Series is a one-dimensional labelled array in Pandas. It can store integers, floats, strings, or other Python objects, and is similar to a single column in a spreadsheet. A DataFrame is the two-dimensional structure.
Q3. Which function reads a CSV file into a DataFrame?
pd.read_csv()pd.open_csv()np.read_csv()plt.csv()
Answer: (a) pd.read_csv()
Explanation: pd.read_csv() is the Pandas function for reading CSV files into a DataFrame. pd.open_csv() and plt.csv() do not exist; np.read_csv() is not a NumPy function (NumPy uses np.loadtxt() or np.genfromtxt()).
Q4. What is the main array object in NumPy called?
- DataFrame
- Series
- ndarray
- pyplot
Answer: (c) ndarray
Explanation: The main array object in NumPy is called ndarray (N-dimensional array). It stores elements of the same data type in a compact and efficient way. DataFrame and Series are Pandas objects; pyplot is a Matplotlib module.
Q5. Which NumPy function is used to sort an array?
np.order()np.sort()np.filter()np.arange()
Answer: (b) np.sort()
Explanation: np.sort() returns a sorted copy of the array. np.order() does not exist; np.filter() does not exist; np.arange() creates values in a range.
Q6. Which Matplotlib function is used to display a chart?
plt.display()plt.view()plt.show()plt.open()
Answer: (c) plt.show()
Explanation: plt.show() displays the figure on screen. plt.display(), plt.view(), and plt.open() are not Matplotlib functions.
Q7. Which chart is most suitable for showing trend over time?
- Pie chart
- Line plot
- Histogram
- Box plot
Answer: (b) Line plot
Explanation: Line plots connect data points chronologically to show trends over time. Pie charts show parts of a whole; histograms show distributions; box plots show five-number summaries.
Q8. Which library is commonly used for parsing HTML in basic web scraping?
- BeautifulSoup
- NumPy
- Matplotlib
- json
Answer: (a) BeautifulSoup
Explanation: BeautifulSoup is a Python library for parsing HTML and XML documents. It is commonly used with requests for web scraping. NumPy is for numerical computation; Matplotlib is for visualization; json is for JSON data handling.
PART B — SHORT ANSWER QUESTIONS (5 Marks Each)
Q9. Define Pandas. Explain any five advantages of using Pandas for data analysis.
Answer:
Definition:Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures (Series and DataFrame) and functions for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs. The name Pandas is derived from Panel Data, which refers to multidimensional structured datasets.
Five Advantages of Pandas:
| # | Advantage | Explanation | Example |
|---|---|---|---|
| 1 | Data Loading | Reads data from many file formats and sources | read_csv(), read_json(), read_excel() |
| 2 | Data Cleaning | Handles missing values, duplicates, and incorrect formats | dropna(), fillna(), drop_duplicates() |
| 3 | Data Selection | Selects rows, columns, and subsets efficiently | df["Name"], loc[], iloc[] |
| 4 | Data Transformation | Creates new columns and modifies existing data | df["Total"] = df["A"] + df["B"] |
| 5 | Data Aggregation | Groups and summarizes data easily | groupby(), mean(), sum() |
| 6 | Data Export | Writes results back to files | to_csv(), to_json() |
Additional advantages: Handles mixed data types, supports labelled indexing, integrates well with NumPy and Matplotlib, and is faster than using basic Python lists and dictionaries for tabular data.
Q10. Differentiate between Pandas Series and DataFrame with suitable examples.
Answer:
| Aspect | Pandas Series | Pandas DataFrame |
|---|---|---|
| Definition | One-dimensional labelled array | Two-dimensional labelled data structure |
| Dimensions | 1D | 2D (rows × columns) |
| Similar to | A single column in a spreadsheet | A complete table with rows and columns |
| Index | Single index | Row index + column index |
| Data Types | Homogeneous (single dtype) | Heterogeneous (mixed dtypes) |
| Creation | pd.Series([1, 2, 3]) | pd.DataFrame({"A": [1, 2], "B": [3, 4]}) |
| Access | series[0], series["label"] | df["col"], df.loc[], df.iloc[] |
| Use Case | Single variable analysis | Multi-variable tabular analysis |
Example — Series:
import pandas as pd
marks = pd.Series([78, 85, 92, 66], index=["Amit", "Ravi", "Neha", "Sara"])
print(marks)
print(marks["Neha"]) # 92Output:
Amit 78
Ravi 85
Neha 92
Sara 66
dtype: int64
92Example — DataFrame:
import pandas as pd
data = {
"Name": ["Amit", "Ravi", "Neha", "Sara"],
"Age": [21, 22, 20, 23],
"Marks": [78, 85, 92, 66]
}
df = pd.DataFrame(data)
print(df)Output:
Name Age Marks
0 Amit 21 78
1 Ravi 22 85
2 Neha 20 92
3 Sara 23 66Key Difference: A Series is a single column; a DataFrame is a collection of Series sharing the same row index.
Q11. Explain how to select rows and columns from a DataFrame using loc[] and iloc[].
Answer:
Pandas provides two primary methods for selecting rows and columns:
loc[]— label-based indexing (uses row/column labels)iloc[]— integer-position-based indexing (uses numeric positions)
Syntax:
df.loc[row_label, column_label]df.iloc[row_position, column_position]
Example DataFrame:
import pandas as pd
data = {
"Name": ["Amit", "Ravi", "Neha", "Sara"],
"Age": [21, 22, 20, 23],
"Marks": [78, 85, 92, 66]
}
df = pd.DataFrame(data)Using loc[] (Label-Based):
# Select row with label 0
print(df.loc[0])
# Select rows 0 to 2 (inclusive)
print(df.loc[0:2])
# Select specific rows and columns
print(df.loc[0:2, ["Name", "Marks"]])
# Conditional selection
print(df.loc[df["Marks"] > 80])Using iloc[] (Position-Based):
# Select first row
print(df.iloc[0])
# Select rows 0 to 2 (exclusive of 3)
print(df.iloc[0:3])
# Select rows 0-2, columns 0-1
print(df.iloc[0:3, 0:2])
# Select last row
print(df.iloc[-1])Comparison Table:
| Method | Based On | Example | Result |
|---|---|---|---|
df.loc[0] | Label | df.loc[0] | Row with index label 0 |
df.iloc[0] | Position | df.iloc[0] | First row |
df.loc[0:2] | Label (inclusive) | df.loc[0:2] | Rows 0, 1, 2 |
df.iloc[0:3] | Position (exclusive) | df.iloc[0:3] | Rows 0, 1, 2 |
df.loc[:, "Name"] | Label | df.loc[:, "Name"] | All rows, “Name” column |
df.iloc[:, 0] | Position | df.iloc[:, 0] | All rows, first column |
Key Difference: loc[] uses labels (inclusive of endpoint); iloc[] uses integer positions (exclusive of endpoint, like Python slicing).
Q12. Write short notes on handling missing values in Pandas using dropna() and fillna().
Answer:
Missing values are common in real datasets. Pandas represents missing values using NaN (Not a Number). Two primary functions handle missing values:
1. dropna() — Remove Missing Values
Purpose: Removes rows or columns containing missing values.
Syntax:
df.dropna(axis=0, how='any', thresh=None, subset=None, inplace=False)Parameters:
| Parameter | Meaning | Default |
|---|---|---|
axis | 0 = drop rows; 1 = drop columns | 0 |
how | ‘any’ = drop if any NaN; ‘all’ = drop if all NaN | ‘any’ |
thresh | Minimum non-NaN values required to keep | None |
subset | Columns to consider | None |
inplace | Modify DataFrame in place | False |
Example:
import pandas as pd
data = {"Name": ["Amit", "Ravi", "Neha"], "Marks": [78, None, 92]}
df = pd.DataFrame(data)
print(df.dropna()) # Remove rows with any NaN
print(df.dropna(axis=1)) # Remove columns with any NaN2. fillna() — Replace Missing Values
Purpose: Replaces missing values with a specified value or method.
Syntax:
df.fillna(value=None, method=None, axis=None, inplace=False)Parameters:
| Parameter | Meaning |
|---|---|
value | Value to replace NaN with (scalar, dict, Series) |
method | ‘ffill’ (forward fill) or ‘bfill’ (backward fill) |
axis | 0 = rows; 1 = columns |
inplace | Modify DataFrame in place |
Example:
print(df.fillna(0)) # Replace NaN with 0
print(df.fillna(df.mean())) # Replace with column mean
print(df.fillna(method='ffill')) # Forward fill3. Other Missing Value Functions
| Function | Purpose |
|---|---|
isnull() | Detects missing values (returns True/False) |
notnull() | Detects non-missing values |
drop_duplicates() | Removes duplicate rows |
Complete Example:
import pandas as pd
data = {"Name": ["Amit", "Ravi", "Neha"], "Marks": [78, None, 92]}
df = pd.DataFrame(data)
print(df.isnull()) # Check missing values
print(df.dropna()) # Remove rows with missing values
print(df.fillna(0)) # Replace missing values with 0
print(df.fillna(df["Marks"].mean())) # Replace with meanBest Practice: Choose the method based on the data and analysis goal. Dropping is suitable when missing data is random and small in proportion; filling is better when missing data is substantial or when preserving rows is important.
Q13. Explain how CSV and JSON files can be read and written using Pandas.
Answer:
Pandas can read data from many file formats. Two of the most common formats are CSV (Comma-Separated Values) and JSON (JavaScript Object Notation).
1. Reading CSV Files
Syntax: pd.read_csv(filepath, sep=',', header='infer', names=None, usecols=None)
import pandas as pd
# Read a CSV file
students = pd.read_csv("students.csv")
# Inspect the data
print(students.head()) # First 5 rows
print(students.shape) # (rows, columns)
print(students.info()) # Column types and non-null counts
print(students.describe()) # Statistical summaryCommon Parameters:
| Parameter | Purpose |
|---|---|
sep | Delimiter (default ‘,’) |
header | Row to use as column names |
names | Custom column names |
usecols | Subset of columns to read |
dtype | Data types for columns |
2. Writing CSV Files
Syntax: df.to_csv(filepath, index=False)
# Save a DataFrame to a CSV file
students.to_csv("cleaned_students.csv", index=False)Common Parameters:
| Parameter | Purpose |
|---|---|
index | Include row index (default True) |
columns | Subset of columns to write |
sep | Delimiter |
header | Include column names |
3. Reading JSON Files
Syntax: pd.read_json(filepath, orient=None)
import pandas as pd
# Read a JSON file
orders = pd.read_json("orders.json")
print(orders.head())Common Parameters:
| Parameter | Purpose |
|---|---|
orient | JSON orientation: ‘records’, ‘index’, ‘columns’, ‘values’, ‘split’, ‘table’ |
lines | Read line-delimited JSON |
4. Writing JSON Files
Syntax: df.to_json(filepath, orient='records', indent=4)
# Save a DataFrame to a JSON file
orders.to_json("cleaned_orders.json", orient="records", indent=4)Common Parameters:
| Parameter | Purpose |
|---|---|
orient | JSON orientation |
indent | Pretty-print indentation |
date_format | Format for datetime values |
5. File Reading and Writing Functions Summary
| Function | Purpose | Common Parameter |
|---|---|---|
read_csv() | Reads a CSV file into a DataFrame | sep, header, names, usecols |
to_csv() | Writes a DataFrame to a CSV file | index=False |
read_json() | Reads a JSON file into a DataFrame | orient |
to_json() | Writes a DataFrame to a JSON file | orient, indent |
Practical Tip: After reading any file, always use head(), shape, info(), and describe() to understand the structure and quality of the imported data.
Q14. Define NumPy ndarray. Explain shape, ndim, size, and dtype attributes.
Answer:
Definition:NumPy (Numerical Python) is a fundamental Python library for scientific computing and numerical operations. It provides the ndarray object (N-dimensional array), which stores elements of the same data type in a compact and efficient way.
Compared with normal Python lists, NumPy arrays are:
- Faster — optimized C implementation
- More memory efficient — contiguous storage
- Support vectorized operations — operations on entire arrays without explicit loops
Creating an ndarray:
import numpy as np
arr = np.array([10, 20, 30, 40])
print(arr) # [10 20 30 40]
print(type(arr)) # <class 'numpy.ndarray'>Attributes of ndarray:
1. shape
Returns the dimensions of the array as a tuple (rows, columns, …).
a = np.array([[1, 2, 3], [4, 5, 6]])
print(a.shape) # (2, 3) — 2 rows, 3 columns2. ndim
Returns the number of dimensions (axes) of the array.
print(a.ndim) # 2 — two-dimensional array3. size
Returns the total number of elements in the array.
print(a.size) # 6 — 2 × 3 = 6 elements4. dtype
Returns the data type of the elements in the array.
print(a.dtype) # int64 (or int32 depending on platform)Complete Example:
import numpy as np
a = np.array([[1, 2, 3], [4, 5, 6]])
print(a.shape) # (2, 3)
print(a.ndim) # 2
print(a.size) # 6
print(a.dtype) # int64Summary Table:
| Attribute | Meaning | Example Output |
|---|---|---|
shape | Dimensions of the array | (2, 3) |
ndim | Number of dimensions | 2 |
size | Total number of elements | 6 |
dtype | Data type of elements | int64 |
Other Useful Attributes:
| Attribute | Meaning |
|---|---|
itemsize | Size in bytes of each element |
nbytes | Total bytes consumed by the array |
T | Transposed view of the array |
Why These Attributes Matter: They help you understand the structure and memory layout of your data before performing operations. For example, shape is essential for reshaping, and dtype determines which operations are valid.
Q15. Explain indexing and slicing in one-dimensional and two-dimensional NumPy arrays.
Answer:
Indexing is used to access individual elements, while slicing is used to access a range or subset of elements. NumPy indexing starts from 0, just like normal Python lists.
1. One-Dimensional Arrays
Syntax: array[start:stop:step]
import numpy as np
a = np.array([10, 20, 30, 40, 50])
print(a[0]) # 10 — first element
print(a[-1]) # 50 — last element
print(a[1:4]) # [20, 30, 40] — elements from index 1 to 3
print(a[:3]) # [10, 20, 30] — first three elements
print(a[::2]) # [10, 30, 50] — every second element
print(a[::-1]) # [50, 40, 30, 20, 10] — reversedIndexing and Slicing Examples:
| Expression | Meaning | Result |
|---|---|---|
a[0] | First element | 10 |
a[-1] | Last element | 50 |
a[1:4] | Elements from index 1 to 3 | [20, 30, 40] |
a[:3] | First three elements | [10, 20, 30] |
a[::2] | Every second element | [10, 30, 50] |
2. Two-Dimensional Arrays
Syntax: array[row, column] or array[row_start:row_stop, col_start:col_stop]
import numpy as np
matrix = np.array([[1, 2, 3],
[4, 5, 6],
[7, 8, 9]])
print(matrix[0, 0]) # 1 — row 0, column 0
print(matrix[1, 2]) # 6 — row 1, column 2
print(matrix[:, 1]) # [2, 5, 8] — all rows, column 1
print(matrix[0:2, 1:3]) # Sub-array [[2, 3], [5, 6]]
print(matrix[1, :]) # [4, 5, 6] — row 1, all columns
print(matrix[:, 0]) # [1, 4, 7] — all rows, column 02D Indexing and Slicing Examples:
| Expression | Meaning | Result |
|---|---|---|
matrix[0, 0] | Element at row 0, column 0 | 1 |
matrix[1, 2] | Element at row 1, column 2 | 6 |
matrix[:, 1] | All rows, column 1 | [2, 5, 8] |
matrix[0:2, 1:3] | Rows 0-1, columns 1-2 | [[2, 3], [5, 6]] |
matrix[1, :] | Row 1, all columns | [4, 5, 6] |
matrix[:, 0] | All rows, column 0 | [1, 4, 7] |
3. Key Points
- Indexing accesses a single element:
a[0],matrix[1, 2]. - Slicing accesses a range:
a[1:4],matrix[0:2, 1:3]. - Colon (
:) means “all” in that dimension. - Negative indices count from the end:
a[-1]is the last element. - Step can be used:
a[::2]takes every second element. - Slicing returns a view, not a copy — modifying the slice modifies the original array.
Example of View vs Copy:
a = np.array([1, 2, 3, 4, 5])
b = a[1:4] # b is a view
b[0] = 99
print(a) # [1, 99, 3, 4, 5] — original modifiedTo create a copy, use .copy():
b = a[1:4].copy()Q16. Discuss joining and splitting operations in NumPy with examples.
Answer:
NumPy provides functions to join (combine) and split (divide) arrays. These operations are essential for data manipulation and reshaping.
1. Joining Arrays
(a) concatenate() — Join Along an Existing Axis
import numpy as np
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
joined = np.concatenate((a, b))
print(joined) # [1 2 3 4 5 6]For 2D arrays:
x = np.array([[1, 2], [3, 4]])
y = np.array([[5, 6], [7, 8]])
print(np.concatenate((x, y), axis=0)) # Vertical (rows)
print(np.concatenate((x, y), axis=1)) # Horizontal (columns)(b) vstack() — Vertical Stacking
Stacks arrays vertically (row-wise).
x = np.array([[1, 2], [3, 4]])
y = np.array([[5, 6], [7, 8]])
print(np.vstack((x, y)))
# [[1 2]
# [3 4]
# [5 6]
# [7 8]](c) hstack() — Horizontal Stacking
Stacks arrays horizontally (column-wise).
print(np.hstack((x, y)))
# [[1 2 5 6]
# [3 4 7 8]]2. Splitting Arrays
(a) array_split() — Split into Multiple Parts
import numpy as np
a = np.array([1, 2, 3, 4, 5, 6])
parts = np.array_split(a, 3)
print(parts)
# [array([1, 2]), array([3, 4]), array([5, 6])](b) split() — Equal Division
parts = np.split(a, 3)
print(parts) # Same as above but requires equal division(c) hsplit() and vsplit() — Split 2D Arrays
matrix = np.array([[1, 2, 3, 4], [5, 6, 7, 8]])
print(np.hsplit(matrix, 2)) # Split horizontally
print(np.vsplit(matrix, 2)) # Split vertically3. Reshaping
reshape() — Change Shape Without Changing Data
a = np.array([1, 2, 3, 4, 5, 6])
reshaped = a.reshape(2, 3)
print(reshaped)
# [[1 2 3]
# [4 5 6]]4. Summary Table
| Function | Purpose | Example |
|---|---|---|
concatenate() | Joins arrays along an existing axis | np.concatenate((a, b)) |
vstack() | Stacks arrays vertically | np.vstack((x, y)) |
hstack() | Stacks arrays horizontally | np.hstack((x, y)) |
array_split() | Splits an array into multiple parts | np.array_split(a, 3) |
split() | Splits into equal parts | np.split(a, 3) |
hsplit() | Splits horizontally | np.hsplit(matrix, 2) |
vsplit() | Splits vertically | np.vsplit(matrix, 2) |
reshape() | Changes shape without changing data | a.reshape(2, 3) |
5. Practical Example
import numpy as np
# Joining
sem1 = np.array([85, 78, 92])
sem2 = np.array([88, 82, 95])
all_marks = np.concatenate((sem1, sem2))
print("Joined:", all_marks)
# Splitting
data = np.array([10, 20, 30, 40, 50, 60])
parts = np.array_split(data, 3)
print("Split:", parts)
# Reshaping
matrix = np.arange(1, 13).reshape(3, 4)
print("Reshaped:\n", matrix)Output:
Joined: [85 78 92 88 82 95]
Split: [array([10, 20]), array([30, 40]), array([50, 60])]
Reshaped:
[[ 1 2 3 4]
[ 5 6 7 8]
[ 9 10 11 12]]Q17. Explain searching, sorting, and filtering arrays in NumPy.
Answer:
NumPy provides efficient functions for searching, sorting, and filtering arrays. These operations are essential for extracting useful information from data.
1. Searching Arrays
np.where() — Find Indices Where Condition is True
import numpy as np
a = np.array([10, 20, 30, 20, 40])
result = np.where(a == 20)
print(result) # (array([1, 3]),)With conditions:
print(np.where(a > 25)) # (array([2, 4]),)np.searchsorted() — Find Insertion Position
sorted_arr = np.array([10, 20, 30, 40, 50])
print(np.searchsorted(sorted_arr, 35)) # 3 (insert before index 3)np.argmax() and np.argmin() — Index of Max/Min
a = np.array([10, 50, 30, 20])
print(np.argmax(a)) # 1 (index of 50)
print(np.argmin(a)) # 0 (index of 10)2. Sorting Arrays
np.sort() — Returns Sorted Copy
import numpy as np
a = np.array([40, 10, 30, 20])
print(np.sort(a)) # [10 20 30 40]array.sort() — Sorts In-Place
a = np.array([40, 10, 30, 20])
a.sort()
print(a) # [10 20 30 40]Sorting 2D Arrays
matrix = np.array([[3, 1, 2], [6, 5, 4]])
print(np.sort(matrix, axis=0)) # Sort each column
print(np.sort(matrix, axis=1)) # Sort each rownp.argsort() — Indices That Would Sort the Array
a = np.array([40, 10, 30, 20])
print(np.argsort(a)) # [1 3 2 0]3. Filtering Arrays
Filtering uses boolean masks — conditions that produce True/False for each element.
import numpy as np
a = np.array([10, 25, 30, 45, 50])
# Filter values greater than 30
filtered = a[a > 30]
print(filtered) # [45 50]
# Filter with multiple conditions
filtered = a[(a > 20) & (a < 50)]
print(filtered) # [25 30 45]
# Filter even numbers
evens = a[a % 2 == 0]
print(evens) # [10 30 50]Boolean Mask:
mask = a > 30
print(mask) # [False False False True True]
print(a[mask]) # [45 50]4. Summary Table
| Operation | Example | Output Meaning |
|---|---|---|
| Search | np.where(a == 20) | Returns indexes where condition is true |
| Sort | np.sort(a) | Returns sorted copy of the array |
| Filter | a[a > 30] | Returns values greater than 30 |
| Boolean mask | a % 2 == 0 | Creates True/False condition for each value |
| Argmax | np.argmax(a) | Index of maximum value |
| Argmin | np.argmin(a) | Index of minimum value |
| Argsort | np.argsort(a) | Indices that would sort the array |
5. Complete Example
import numpy as np
marks = np.array([78, 85, 92, 66, 74, 88, 95, 60])
# Searching
print("Students scoring 85:", np.where(marks == 85))
# Sorting
print("Sorted marks:", np.sort(marks))
print("Ranking indices:", np.argsort(marks))
# Filtering
print("Marks above 80:", marks[marks > 80])
print("Marks between 70 and 90:", marks[(marks >= 70) & (marks <= 90)])
# Top scorer
print("Top scorer marks:", marks.max())
print("Top scorer index:", marks.argmax())Output:
Students scoring 85: (array([1]),)
Sorted marks: [60 66 74 78 85 88 92 95]
Ranking indices: [7 3 4 0 1 5 2 6]
Marks above 80: [85 92 88 95]
Marks between 70 and 90: [78 85 74 88]
Top scorer marks: 95
Top scorer index: 6Important: Filtering in NumPy is usually done using boolean conditions. The condition creates a True/False mask, and the mask is used to select matching elements.
Q18. What is web scraping? Explain the basic steps and ethical precautions involved.
Answer:
Definition:Web scraping is the process of extracting data from websites using programs. It is useful when information is available on web pages but not provided as a downloadable dataset or API.
1. Basic Web Scraping Workflow
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Send HTTP │ → │ 2. Download │ → │ 3. Parse and │ → │ 4. Clean and │ → │ 5. Store or │
│ Request │ │ HTML │ │ Extract Data │ │ Structure │ │ Export Data │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘2. Steps in a Basic Web Scraping Workflow
| Step | Meaning | Common Tool |
|---|---|---|
| 1. Send HTTP Request | Send a request to the target website | requests |
| 2. Download HTML | Retrieve the HTML content of the page | response.text |
| 3. Parse and Extract Data | Parse the HTML and extract the required data | BeautifulSoup |
| 4. Clean and Structure Data | Clean the extracted data and structure it for use | Pandas, manual cleaning |
| 5. Store or Export Data | Save the data to a file, database, or export it for analysis | CSV, JSON, Pandas DataFrame |
3. Simple Web Scraping Example
import requests
from bs4 import BeautifulSoup
url = "https://example.com"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")
print(soup.title.text)Explanation:
requests.get(url)sends an HTTP GET request.response.textcontains the HTML source.BeautifulSoup(response.text, "html.parser")parses the HTML.soup.title.textextracts the page title.
4. Common Tools
| Tool | Purpose |
|---|---|
requests | Send HTTP requests |
BeautifulSoup | Parse HTML/XML |
lxml | Faster HTML/XML parser |
Selenium | Browser automation for dynamic pages |
Scrapy | Full-featured scraping framework |
5. Ethical Precautions
Web scraping should be performed responsibly. Programmers must respect:
| Precaution | Explanation |
|---|---|
| Terms of Service | Check if the website allows automated access |
| robots.txt | Follow the rules specified in the site’s robots.txt file |
| Copyright | Do not reproduce copyrighted content without permission |
| Privacy | Do not collect personal or restricted data without consent |
| Server Load | Avoid sending too many requests that could overload the server |
| Official APIs | Prefer official APIs whenever available |
| Rate Limiting | Add delays between requests |
| User-Agent | Identify your bot transparently |
Ethical Reminder: Before scraping a website, check whether the website allows automated access. Prefer official APIs whenever available.
6. Applications of Web Scraping
- Price comparison: Track product prices across e-commerce sites.
- News aggregation: Collect headlines from multiple news sources.
- Sentiment analysis: Gather social media posts for analysis.
- Research: Collect data for academic studies.
- Lead generation: Gather business contact information.
- Real estate: Track property listings and prices.
Q19. Explain the need for data visualization and list common Matplotlib chart types.
Answer:
Need for Data Visualization: Data visualization converts numbers into patterns. A good chart can reveal:
- Trends — how values change over time
- Comparisons — differences between categories
- Distributions — how data is spread
- Relationships — connections between variables
- Outliers — unusual data points
These may not be clear from a table alone. Visualization also helps communicate insights to non-technical audiences.
Common Matplotlib Chart Types:
| Chart Type | Used For | Example Question |
|---|---|---|
| Line plot | Showing trends over time | How did sales change month by month? |
| Bar chart | Comparing categories | Which department scored highest? |
| Histogram | Showing distribution of numerical values | What is the distribution of exam marks? |
| Scatter plot | Showing relationship between two numerical variables | Is study time related to marks? |
| Pie chart | Showing parts of a whole | What percentage of students selected each elective? |
Common Matplotlib pyplot Functions:
| Function | Purpose |
|---|---|
plot() | Creates a line plot |
bar() | Creates a bar chart |
hist() | Creates a histogram |
scatter() | Creates a scatter plot |
pie() | Creates a pie chart |
title() | Adds chart title |
xlabel(), ylabel() | Adds axis labels |
legend() | Displays chart legend |
show() | Displays the figure |
savefig() | Saves chart as an image file |
Visualization Rule: Every chart should have a clear title, axis labels where applicable, readable category names, and a purpose. Avoid adding unnecessary decoration.
Matplotlib Visualization Workflow
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Prepare │ → │ 2. Create │ → │ 3. Plot Data │ → │ 4. Customize │ → │ 5. Show or │
│ Data │ │ Figure │ │ │ │ Plot │ │ Save │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘| Step | Action | Code Example |
|---|---|---|
| 1. Prepare Data | Load and organize data for visualization | x = [1, 2, 3], y = [10, 15, 13] |
| 2. Create Figure | Create a figure and axes | plt.figure(figsize=(8, 5)) |
| 3. Plot Data | Plot data using the appropriate plot type | plt.plot(x, y) |
| 4. Customize Plot | Add titles, labels, legends, colors | plt.title("Title"), plt.xlabel("X") |
| 5. Show or Save | Display or save the plot | plt.show(), plt.savefig("plot.png") |
Complete Example:
import matplotlib.pyplot as plt
# Step 1: Prepare Data
x = [1, 2, 3, 4, 5]
y = [10, 15, 13, 18, 20]
# Step 2: Create Figure (optional)
plt.figure(figsize=(8, 5))
# Step 3: Plot Data
plt.plot(x, y)
# Step 4: Customize Plot
plt.title("Simple Line Plot")
plt.xlabel("X Values")
plt.ylabel("Y Values")
# Step 5: Show or Save
plt.show()Chart Examples:
Bar Chart:
subjects = ["Python", "Maths", "DBMS", "OS"]
marks = [85, 78, 92, 74]
plt.bar(subjects, marks)
plt.title("Marks by Subject")
plt.show()Histogram:
marks = [45, 56, 67, 78, 89, 90, 72, 63, 55, 81]
plt.hist(marks, bins=5)
plt.title("Distribution of Marks")
plt.show()Scatter Plot:
hours = [1, 2, 3, 4, 5, 6]
marks = [40, 50, 55, 65, 75, 85]
plt.scatter(hours, marks)
plt.title("Study Hours vs Marks")
plt.show()Pie Chart:
labels = ["Python", "Java", "C++", "R"]
students = [40, 25, 20, 15]
plt.pie(students, labels=labels, autopct="%1.1f%%")
plt.title("Programming Language Preference")
plt.show()PART C — LONG ANSWER / ESSAY QUESTIONS (10 Marks Each)
Q20. Explain the complete Pandas workflow for loading, inspecting, cleaning, transforming, combining, and exporting tabular data. Support your answer with suitable code examples.
Answer:
1. Introduction to Pandas
Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures (Series and DataFrame) and functions for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs.
2. Pandas Workflow Overview
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Import │ → │ 2. Create │ → │ 3. Clean & │ → │ 4. Combine │ → │ 5. Analyze │
│ Data │ │ Series/DF │ │ Transform │ │ Datasets │ │ & Export │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘3. Step 1: Import Data
Pandas can read data from many file formats.
import pandas as pd
# Read CSV
df = pd.read_csv("students.csv")
# Read JSON
df = pd.read_json("students.json")
# Read Excel
df = pd.read_excel("students.xlsx")Common Parameters:
| Parameter | Purpose |
|---|---|
sep | Delimiter |
header | Row to use as column names |
names | Custom column names |
usecols | Subset of columns |
dtype | Data types |
4. Step 2: Create and Inspect DataFrame
Creating a DataFrame:
import pandas as pd
data = {
"Name": ["Amit", "Ravi", "Neha", "Sara"],
"Age": [21, 22, 20, 23],
"Marks": [78, 85, 92, 66]
}
df = pd.DataFrame(data)Inspection Methods:
| Method | Purpose |
|---|---|
df.head() | First 5 rows |
df.tail() | Last 5 rows |
df.shape | (rows, columns) |
df.columns | Column names |
df.info() | Column types and non-null counts |
df.describe() | Statistical summary |
df.dtypes | Data types of columns |
print(df.head())
print(df.shape)
print(df.info())
print(df.describe())5. Step 3: Clean and Transform Data
Handling Missing Values:
# Detect
print(df.isnull())
# Remove
df = df.dropna()
# Fill
df = df.fillna(0)
df = df.fillna(df.mean())Removing Duplicates:
df = df.drop_duplicates()Selecting Rows and Columns:
# Select columns
df["Name"]
df[["Name", "Marks"]]
# Select rows by label
df.loc[0]
df.loc[df["Marks"] > 80]
# Select rows by position
df.iloc[0:3, 1:3]Creating New Columns:
df["Total"] = df["Python"] + df["Maths"]
df["Average"] = df["Total"] / 2
df["Grade"] = df["Average"].apply(lambda x: "A" if x > 85 else "B")Filtering:
high_scorers = df[df["Marks"] > 80]
low_attendance = df[df["Attendance"] < 75]Sorting:
df = df.sort_values("Marks", ascending=False)Grouping and Aggregation:
dept_avg = df.groupby("Department")["Marks"].mean()
dept_summary = df.groupby("Department").agg({
"Marks": ["mean", "max", "min"],
"Attendance": "mean"
})6. Step 4: Combine DataFrames
Concatenation (stacking):
df1 = pd.DataFrame({"ID": [1, 2], "Name": ["Amit", "Ravi"]})
df2 = pd.DataFrame({"ID": [3, 4], "Name": ["Neha", "Sara"]})
combined = pd.concat([df1, df2], ignore_index=True)Merging (common column):
marks = pd.DataFrame({"ID": [1, 2, 3], "Marks": [78, 85, 92]})
students = pd.DataFrame({"ID": [1, 2, 3], "Name": ["Amit", "Ravi", "Neha"]})
result = pd.merge(students, marks, on="ID")Joining (index-based):
a = pd.DataFrame({"Marks": [78, 85]}, index=[1, 2])
b = pd.DataFrame({"Name": ["Amit", "Ravi"]}, index=[1, 2])
joined = a.join(b)Combining Methods Summary:
| Operation | Meaning | When to Use |
|---|---|---|
concat() | Stacks DataFrames | Same columns or same index |
merge() | Combines using common column | Tables share a key such as ID |
join() | Combines using index | Index labels are meaningful |
7. Step 5: Analyze and Export
Export to CSV:
df.to_csv("cleaned_students.csv", index=False)Export to JSON:
df.to_json("cleaned_students.json", orient="records", indent=4)Export to Excel:
df.to_excel("cleaned_students.xlsx", index=False)8. Complete Integrated Example
import pandas as pd
# Step 1: Load
df = pd.read_csv("students.csv")
# Step 2: Inspect
print(df.head())
print(df.info())
# Step 3: Clean
df = df.dropna()
df = df.drop_duplicates()
# Step 4: Transform
df["Total"] = df["Internal"] + df["External"]
df["Average"] = df["Total"] / 2
# Step 5: Filter
top = df[df["Average"] > 80]
# Step 6: Group
dept_avg = df.groupby("Department")["Average"].mean()
# Step 7: Combine
extra = pd.read_csv("extra_students.csv")
combined = pd.concat([df, extra], ignore_index=True)
# Step 8: Export
combined.to_csv("final_results.csv", index=False)
print("Analysis complete.")9. Benefits of the Pandas Workflow
- Systematic: Each stage has clear inputs and outputs.
- Reproducible: Code can be rerun on new data.
- Maintainable: Easy to modify individual steps.
- Integrated: Works with NumPy, Matplotlib, and scikit-learn.
- Efficient: Handles large datasets faster than pure Python.
Q21. Discuss NumPy arrays in detail. Explain array creation, attributes, indexing, slicing, vectorized operations, joining, splitting, searching, sorting, filtering, and random number generation.
Answer:
1. Introduction to NumPy
NumPy (Numerical Python) is a fundamental Python library for scientific computing and numerical operations. It provides the ndarray object (N-dimensional array), which stores elements of the same data type in a compact and efficient way.
Advantages over Python lists:
- Faster — optimized C implementation
- More memory efficient — contiguous storage
- Support vectorized operations — no explicit loops
- Rich function library — linear algebra, statistics, random
2. Array Creation
| Function | Purpose | Example |
|---|---|---|
np.array() | Creates array from list/tuple | np.array([1, 2, 3]) |
np.zeros() | Array filled with zeros | np.zeros(5) |
np.ones() | Array filled with ones | np.ones((2, 3)) |
np.arange() | Values in a range | np.arange(1, 10, 2) |
np.linspace() | Evenly spaced values | np.linspace(0, 1, 5) |
np.eye() | Identity matrix | np.eye(3) |
import numpy as np
a = np.array([10, 20, 30, 40])
b = np.zeros(5)
c = np.ones((2, 3))
d = np.arange(1, 10, 2)
e = np.linspace(0, 1, 5)
f = np.eye(3)3. ndarray Attributes
| Attribute | Meaning |
|---|---|
shape | Dimensions (rows, columns) |
ndim | Number of dimensions |
size | Total number of elements |
dtype | Data type of elements |
itemsize | Size in bytes of each element |
nbytes | Total bytes consumed |
a = np.array([[1, 2, 3], [4, 5, 6]])
print(a.shape) # (2, 3)
print(a.ndim) # 2
print(a.size) # 6
print(a.dtype) # int644. Indexing and Slicing
1D Arrays:
a = np.array([10, 20, 30, 40, 50])
print(a[0]) # 10
print(a[-1]) # 50
print(a[1:4]) # [20 30 40]
print(a[::2]) # [10 30 50]2D Arrays:
matrix = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
print(matrix[0, 0]) # 1
print(matrix[1, 2]) # 6
print(matrix[:, 1]) # [2 5 8]
print(matrix[0:2, 1:3]) # [[2 3], [5 6]]5. Vectorized Operations
Operations applied to entire arrays without explicit loops.
a = np.array([10, 20, 30])
b = np.array([1, 2, 3])
print(a + b) # [11 22 33]
print(a - b) # [9 18 27]
print(a * b) # [10 40 90]
print(a / b) # [10. 10. 10.]
print(a ** 2) # [100 400 900]6. Joining Arrays
| Function | Purpose |
|---|---|
concatenate() | Joins along an existing axis |
vstack() | Stacks vertically |
hstack() | Stacks horizontally |
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
print(np.concatenate((a, b))) # [1 2 3 4 5 6]
x = np.array([[1, 2], [3, 4]])
y = np.array([[5, 6], [7, 8]])
print(np.vstack((x, y)))
print(np.hstack((x, y)))7. Splitting Arrays
| Function | Purpose |
|---|---|
array_split() | Split into multiple parts |
split() | Split into equal parts |
hsplit() | Split horizontally |
vsplit() | Split vertically |
a = np.array([1, 2, 3, 4, 5, 6])
print(np.array_split(a, 3))
# [array([1, 2]), array([3, 4]), array([5, 6])]8. Searching Arrays
| Function | Purpose |
|---|---|
np.where() | Indices where condition is true |
np.argmax() | Index of maximum value |
np.argmin() | Index of minimum value |
np.searchsorted() | Insertion position |
a = np.array([10, 20, 30, 20, 40])
print(np.where(a == 20)) # (array([1, 3]),)
print(np.argmax(a)) # 4
print(np.argmin(a)) # 09. Sorting Arrays
| Function | Purpose |
|---|---|
np.sort() | Returns sorted copy |
array.sort() | Sorts in-place |
np.argsort() | Indices that would sort |
a = np.array([40, 10, 30, 20])
print(np.sort(a)) # [10 20 30 40]
print(np.argsort(a)) # [1 3 2 0]10. Filtering Arrays
Boolean masks select matching elements.
a = np.array([10, 25, 30, 45, 50])
print(a[a > 30]) # [45 50]
print(a[(a > 20) & (a < 50)]) # [25 30 45]
print(a[a % 2 == 0]) # [10 30 50]11. Random Number Generation
| Function | Purpose |
|---|---|
np.random.rand() | Random floats 0 to 1 |
np.random.randint() | Random integers |
np.random.choice() | Random selection |
np.random.seed() | Reproducibility |
np.random.seed(10)
print(np.random.randint(1, 100, 5)) # Same every time
print(np.random.rand(3)) # 3 random floats
print(np.random.choice([1, 2, 3, 4])) # Random selection12. Summary
NumPy provides:
- Fast array operations
- Flexible indexing and slicing
- Powerful joining and splitting
- Efficient searching, sorting, and filtering
- Reproducible random number generation
- Foundation for Pandas, SciPy, and machine learning
Q22. Explain data visualization using Matplotlib. Compare line plot, bar chart, histogram, scatter plot, and pie chart with suitable examples and use cases.
Answer:
1. Introduction to Matplotlib
Matplotlib is a Python library used for creating visualizations. It allows programmers to create line graphs, bar charts, histograms, scatter plots, pie charts, and many other types of plots.
The most commonly used Matplotlib module is pyplot, usually imported as plt.
import matplotlib.pyplot as plt2. Why Visualize Data?
Visualization helps convert numbers into patterns. A good chart can reveal:
- Trends — how values change over time
- Comparisons — differences between categories
- Distributions — how data is spread
- Relationships — connections between variables
- Outliers — unusual data points
3. Matplotlib Visualization Workflow
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Prepare │ → │ 2. Create │ → │ 3. Plot Data │ → │ 4. Customize │ → │ 5. Show or │
│ Data │ │ Figure │ │ │ │ Plot │ │ Save │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘4. Comparison of Chart Types
| Chart Type | Used For | Example Question |
|---|---|---|
| Line plot | Showing trends over time | How did sales change month by month? |
| Bar chart | Comparing categories | Which department scored highest? |
| Histogram | Showing distribution of numerical values | What is the distribution of exam marks? |
| Scatter plot | Showing relationship between two numerical variables | Is study time related to marks? |
| Pie chart | Showing parts of a whole | What percentage of students selected each elective? |
5. Line Plot
Purpose: Show trends over time.
import matplotlib.pyplot as plt
x = [1, 2, 3, 4, 5]
y = [10, 15, 13, 18, 20]
plt.plot(x, y, marker='o', color='blue')
plt.title("Simple Line Plot")
plt.xlabel("X Values")
plt.ylabel("Y Values")
plt.grid(True, alpha=0.3)
plt.show()Use Cases:
- Stock prices over time
- Monthly sales trends
- Temperature changes
- Website traffic
6. Bar Chart
Purpose: Compare categories.
subjects = ["Python", "Maths", "DBMS", "OS"]
marks = [85, 78, 92, 74]
plt.bar(subjects, marks, color='steelblue')
plt.title("Marks by Subject")
plt.xlabel("Subject")
plt.ylabel("Marks")
plt.show()Use Cases:
- Sales by product
- Population by country
- Survey responses
- Department comparisons
7. Histogram
Purpose: Show distribution of numerical values.
marks = [45, 56, 67, 78, 89, 90, 72, 63, 55, 81]
plt.hist(marks, bins=5, color='green', edgecolor='black')
plt.title("Distribution of Marks")
plt.xlabel("Marks Range")
plt.ylabel("Number of Students")
plt.show()Use Cases:
- Age distribution
- Income distribution
- Test score distribution
- Height distribution
Difference from Bar Chart: Histogram bars touch (continuous data); bar chart bars are separate (categorical data).
8. Scatter Plot
Purpose: Show relationship between two numerical variables.
hours = [1, 2, 3, 4, 5, 6]
marks = [40, 50, 55, 65, 75, 85]
plt.scatter(hours, marks, color='red')
plt.title("Study Hours vs Marks")
plt.xlabel("Study Hours")
plt.ylabel("Marks")
plt.show()Use Cases:
- Height vs weight
- Advertising spend vs sales
- Study time vs marks
- Temperature vs ice cream sales
9. Pie Chart
Purpose: Show parts of a whole.
labels = ["Python", "Java", "C++", "R"]
students = [40, 25, 20, 15]
plt.pie(students, labels=labels, autopct="%1.1f%%")
plt.title("Programming Language Preference")
plt.show()Use Cases:
- Market share
- Budget allocation
- Survey responses
- Population by category
Limitation: Hard to compare many slices; best for 2–5 categories.
10. Common pyplot Functions
| Function | Purpose |
|---|---|
plot() | Line plot |
bar() | Bar chart |
hist() | Histogram |
scatter() | Scatter plot |
pie() | Pie chart |
title() | Chart title |
xlabel(), ylabel() | Axis labels |
legend() | Chart legend |
show() | Display figure |
savefig() | Save as image |
11. Visualization Rule
Every chart should have:
- A clear title
- Axis labels where applicable
- Readable category names
- A clear purpose
Avoid adding unnecessary decoration.
12. Summary
| Chart | Best For | Avoid When |
|---|---|---|
| Line | Trends over time | Categorical data |
| Bar | Comparing categories | Continuous data |
| Histogram | Distributions | Categorical data |
| Scatter | Relationships | Too many points |
| Pie | Parts of a whole | More than 5 categories |
Q23. Describe a complete mini-project in which a CSV file is loaded using Pandas, numerical operations are performed using NumPy, and results are visualized using Matplotlib.
Answer:
1. Project Title: Student Performance Analysis
Objective: Analyze student performance data from a CSV file, compute statistics, and visualize results.
2. Project Workflow
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Load CSV │ → │ 2. Clean & │ → │ 3. Compute │ → │ 4. Visualize │
│ (Pandas) │ │ Transform │ │ (NumPy) │ │ (Matplotlib)│
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘3. Sample Data (students.csv)
Name,Department,Internal,External,Attendance
Amit,CSE,25,60,85
Ravi,CSE,22,55,78
Neha,ECE,28,65,92
Sara,ECE,20,50,70
John,MECH,24,58,88
Priya,CSE,27,62,95
Arjun,MECH,18,45,65
Meena,ECE,26,60,804. Complete Program
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# ─────────────────────────────────────────────
# Step 1: Load Data (Pandas)
# ─────────────────────────────────────────────
df = pd.read_csv("students.csv")
print("=== First 5 rows ===")
print(df.head())
print("\n=== Data Info ===")
print(df.info())
print("\n=== Statistical Summary ===")
print(df.describe())
# ─────────────────────────────────────────────
# Step 2: Clean and Transform Data (Pandas)
# ─────────────────────────────────────────────
# Check for missing values
print("\n=== Missing Values ===")
print(df.isnull().sum())
# Remove duplicates
df = df.drop_duplicates()
# Create new columns
df["Total"] = df["Internal"] + df["External"]
df["Average"] = np.round(df["Total"] / 2, 2)
# Grade assignment
def assign_grade(avg):
if avg >= 85:
return "A"
elif avg >= 70:
return "B"
elif avg >= 55:
return "C"
else:
return "D"
df["Grade"] = df["Average"].apply(assign_grade)
print("\n=== Transformed Data ===")
print(df)
# ─────────────────────────────────────────────
# Step 3: Numerical Operations (NumPy)
# ─────────────────────────────────────────────
mean_avg = np.mean(df["Average"])
median_avg = np.median(df["Average"])
std_avg = np.std(df["Average"])
max_avg = np.max(df["Average"])
min_avg = np.min(df["Average"])
print("\n=== Statistical Measures (NumPy) ===")
print(f"Mean Average:{mean_avg:.2f}")
print(f"Median Average:{median_avg:.2f}")
print(f"Std Deviation:{std_avg:.2f}")
print(f"Max Average:{max_avg:.2f}")
print(f"Min Average:{min_avg:.2f}")
# Filter top performers
top = df[df["Average"] > 80]
print(f"\nTop performers ({len(top)}):")
print(top[["Name", "Average", "Grade"]])
# Filter low attendance
low_att = df[df["Attendance"] < 75]
print(f"\nLow attendance ({len(low_att)}):")
print(low_att[["Name", "Attendance"]])
# ─────────────────────────────────────────────
# Step 4: Grouping and Aggregation (Pandas)
# ─────────────────────────────────────────────
dept_avg = df.groupby("Department")["Average"].mean().reset_index()
print("\n=== Department-wise Average ===")
print(dept_avg)
# ─────────────────────────────────────────────
# Step 5: Visualization (Matplotlib)
# ─────────────────────────────────────────────
fig, axes = plt.subplots(2, 2, figsize=(14, 10))
# Chart 1: Bar chart — Average marks by student
axes[0, 0].bar(df["Name"], df["Average"], color="steelblue")
axes[0, 0].axhline(y=mean_avg, color="red", linestyle="--", label=f"Mean ={mean_avg:.1f}")
axes[0, 0].set_title("Average Marks by Student")
axes[0, 0].set_xlabel("Student")
axes[0, 0].set_ylabel("Average Marks")
axes[0, 0].legend()
axes[0, 0].tick_params(axis="x", rotation=45)
# Chart 2: Histogram — Distribution of averages
axes[0, 1].hist(df["Average"], bins=5, color="green", edgecolor="black")
axes[0, 1].set_title("Distribution of Average Marks")
axes[0, 1].set_xlabel("Average Marks")
axes[0, 1].set_ylabel("Number of Students")
# Chart 3: Scatter plot — Attendance vs Average
axes[1, 0].scatter(df["Attendance"], df["Average"], color="red", s=80)
axes[1, 0].set_title("Attendance vs Average Marks")
axes[1, 0].set_xlabel("Attendance (%)")
axes[1, 0].set_ylabel("Average Marks")
axes[1, 0].grid(True, alpha=0.3)
# Chart 4: Bar chart — Department-wise average
axes[1, 1].bar(dept_avg["Department"], dept_avg["Average"], color="purple")
axes[1, 1].set_title("Department-wise Average Marks")
axes[1, 1].set_xlabel("Department")
axes[1, 1].set_ylabel("Average Marks")
plt.tight_layout()
plt.savefig("student_performance.png")
plt.show()
# ─────────────────────────────────────────────
# Step 6: Export Results (Pandas)
# ─────────────────────────────────────────────
df.to_csv("analyzed_students.csv", index=False)
dept_avg.to_csv("department_averages.csv", index=False)
print("\n=== Analysis Complete ===")
print("Files saved: student_performance.png, analyzed_students.csv, department_averages.csv")5. Sample Output
=== First 5 rows ===
Name Department Internal External Attendance
0 Amit CSE 25 60 85
1 Ravi CSE 22 55 78
2 Neha ECE 28 65 92
3 Sara ECE 20 50 70
4 John MECH 24 58 88
=== Statistical Measures (NumPy) ===
Mean Average: 71.75
Median Average: 72.50
Std Deviation: 11.36
Max Average: 89.00
Min Average: 52.50
=== Department-wise Average ===
Department Average
0 CSE 75.00
1 ECE 72.75
2 MECH 66.256. Role of Each Library
| Library | Role in the Project |
|---|---|
| Pandas | Loads CSV, cleans data, creates columns, groups, exports |
| NumPy | Computes mean, median, std, max, min; rounds values |
| Matplotlib | Creates bar charts, histogram, scatter plot; saves image |
7. Key Insights from the Analysis
- CSE has the highest average marks (75.0).
- MECH has the lowest average (66.25).
- Students with attendance > 85% tend to score above 80.
- The distribution of averages is roughly symmetric.
- Top performers: Neha (89.0), Amit (85.0), Priya (85.0).
8. Extensions
- Add correlation analysis between attendance and marks.
- Include more visualizations (pie chart for grade distribution).
- Build an interactive dashboard using Plotly.
- Apply machine learning to predict student performance.
9. Conclusion
This mini-project demonstrates the complete data analysis pipeline:
- Load data with Pandas
- Clean and transform with Pandas
- Compute statistics with NumPy
- Visualize with Matplotlib
- Export results with Pandas
Q24. Web scraping is useful but must be used responsibly. Explain the workflow of web scraping, tools used, possible applications, limitations, and ethical/legal precautions.
Answer:
1. What is Web Scraping?
Web scraping is the process of extracting data from websites using programs. It is useful when information is available on web pages but not provided as a downloadable dataset or API.
2. Web Scraping Workflow
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Send HTTP │ → │ 2. Download │ → │ 3. Parse and │ → │ 4. Clean and │ → │ 5. Store or │
│ Request │ │ HTML │ │ Extract Data │ │ Structure │ │ Export Data │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘| Step | Meaning | Common Tool |
|---|---|---|
| 1. Send HTTP Request | Send a request to the target website | requests |
| 2. Download HTML | Retrieve the HTML content of the page | response.text |
| 3. Parse and Extract Data | Parse the HTML and extract the required data | BeautifulSoup |
| 4. Clean and Structure Data | Clean the extracted data and structure it for use | Pandas, manual cleaning |
| 5. Store or Export Data | Save the data to a file, database, or export it for analysis | CSV, JSON, Pandas DataFrame |
3. Tools Used
| Tool | Purpose | Example |
|---|---|---|
requests | Send HTTP requests | requests.get(url) |
BeautifulSoup | Parse HTML/XML | BeautifulSoup(html, "html.parser") |
lxml | Faster HTML/XML parser | BeautifulSoup(html, "lxml") |
Selenium | Browser automation for dynamic pages | driver.get(url) |
Scrapy | Full-featured scraping framework | scrapy crawl spider |
Pandas | Store and analyze scraped data | pd.DataFrame(data) |
4. Simple Web Scraping Example
import requests
from bs4 import BeautifulSoup
url = "https://example.com"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")
# Extract title
print(soup.title.text)
# Extract all links
for link in soup.find_all("a"):
print(link.get("href"))
# Extract all paragraphs
for p in soup.find_all("p"):
print(p.text)5. Possible Applications
| Application | Description |
|---|---|
| Price comparison | Track product prices across e-commerce sites |
| News aggregation | Collect headlines from multiple news sources |
| Sentiment analysis | Gather social media posts for analysis |
| Research | Collect data for academic studies |
| Lead generation | Gather business contact information |
| Real estate | Track property listings and prices |
| Job market analysis | Collect job postings and salary data |
| Sports statistics | Gather player and team performance data |
6. Limitations
| Limitation | Explanation |
|---|---|
| Dynamic content | JavaScript-rendered content not available in HTML |
| Anti-scraping measures | CAPTCHAs, IP blocking, rate limiting |
| Changing structure | Website layout changes break scrapers |
| Legal restrictions | Terms of service may prohibit scraping |
| Data quality | Scraped data may be incomplete or noisy |
| Maintenance | Scrapers require regular updates |
| Performance | Large-scale scraping is resource-intensive |
7. Ethical and Legal Precautions
| Precaution | Explanation |
|---|---|
| Terms of Service | Check if the website allows automated access |
| robots.txt | Follow the rules specified in the site’s robots.txt file |
| Copyright | Do not reproduce copyrighted content without permission |
| Privacy | Do not collect personal or restricted data without consent |
| Server Load | Avoid sending too many requests that could overload the server |
| Official APIs | Prefer official APIs whenever available |
| Rate Limiting | Add delays between requests (time.sleep(2)) |
| User-Agent | Identify your bot transparently |
| Data Usage | Use scraped data only for intended purposes |
| Attribution | Credit the source where appropriate |
8. Responsible Scraping Example
import requests
from bs4 import BeautifulSoup
import time
headers = {"User-Agent": "Mozilla/5.0 (educational scraper)"}
for page in range(1, 10):
url = f"https://example.com/page{page}"
try:
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
print(f"Page{page}:{soup.title.text}")
time.sleep(2) # Respectful delay
except requests.RequestException as e:
print(f"Error on page{page}:{e}")9. Alternatives to Web Scraping
| Alternative | Description |
|---|---|
| Official APIs | Structured data provided by the website |
| Public datasets | Pre-collected data from Kaggle, UCI, etc. |
| RSS feeds | Standardized content feeds |
| Data dumps | Bulk data downloads offered by platforms |
| Partnerships | Direct data-sharing agreements |
10. Summary
Web scraping is a powerful technique for collecting data from websites, but it must be used responsibly. Key points:
- Workflow: Request → Download → Parse → Clean → Store
- Tools:
requests,BeautifulSoup,Selenium,Scrapy - Applications: Price comparison, news, research, sentiment analysis
- Limitations: Dynamic content, anti-scraping, legal restrictions
- Ethics: Respect robots.txt, terms of service, privacy, server load
- Best Practice: Prefer official APIs whenever available
PART D — ANALYTICAL / CASE-BASED QUESTIONS
Q25 (Case Study). A college has student performance data stored in a CSV file with columns such as Name, Department, Internal Marks, External Marks, Attendance, and Result. The management wants a Python program to read the data, calculate total marks, find average marks by department, identify students with low attendance, and visualize department-wise performance.
Questions:
- Which Pandas functions will be useful for reading and inspecting the data?
- How can new columns such as Total and Average be created?
- How can students with attendance below 75% be filtered?
- Which Matplotlib charts would be suitable for showing department-wise averages and result distribution?
- Where can NumPy be used in this problem?
Answer:
1. Pandas Functions for Reading and Inspecting Data
| Function | Purpose |
|---|---|
pd.read_csv() | Read the CSV file into a DataFrame |
df.head() | Display first 5 rows |
df.tail() | Display last 5 rows |
df.shape | Show (rows, columns) |
df.columns | List column names |
df.info() | Show data types and non-null counts |
df.describe() | Statistical summary of numerical columns |
df.isnull().sum() | Count missing values per column |
df.dtypes | Data types of columns |
Example:
import pandas as pd
df = pd.read_csv("student_performance.csv")
print(df.head())
print(df.shape)
print(df.info())
print(df.describe())
print(df.isnull().sum())2. Creating New Columns: Total and Average
# Create Total column
df["Total"] = df["Internal Marks"] + df["External Marks"]
# Create Average column
df["Average"] = df["Total"] / 2
# Alternative using NumPy for rounding
import numpy as np
df["Average"] = np.round(df["Total"] / 2, 2)Explanation:
- Pandas allows column creation by assignment:
df["NewCol"] = values. - Operations are vectorized (applied to all rows at once).
- NumPy can be used for rounding:
np.round().
3. Filtering Students with Attendance Below 75%
low_attendance = df[df["Attendance"] < 75]
print(low_attendance[["Name", "Department", "Attendance"]])
# Count
print(f"Students with low attendance:{len(low_attendance)}")
# Save to file
low_attendance.to_csv("low_attendance.csv", index=False)Explanation:
- Boolean filtering:
df[df["Attendance"] < 75]returns rows where the condition is True. - Multiple conditions:
df[(df["Attendance"] < 75) & (df["Result"] == "Fail")].
4. Suitable Matplotlib Charts
| Objective | Chart Type | Justification |
|---|---|---|
| Department-wise averages | Bar chart | Compares categories clearly |
| Result distribution | Pie chart or bar chart | Shows parts of a whole (Pass/Fail) |
| Attendance vs Average | Scatter plot | Shows relationship |
| Distribution of marks | Histogram | Shows frequency distribution |
Example — Department-wise Average (Bar Chart):
import matplotlib.pyplot as plt
dept_avg = df.groupby("Department")["Average"].mean().reset_index()
plt.bar(dept_avg["Department"], dept_avg["Average"], color="steelblue")
plt.title("Department-wise Average Marks")
plt.xlabel("Department")
plt.ylabel("Average Marks")
plt.show()Example — Result Distribution (Pie Chart):
result_counts = df["Result"].value_counts()
plt.pie(result_counts, labels=result_counts.index, autopct="%1.1f%%")
plt.title("Result Distribution")
plt.show()5. Where NumPy Can Be Used
| Task | NumPy Function |
|---|---|
| Rounding averages | np.round(df["Average"], 2) |
| Computing mean | np.mean(df["Average"]) |
| Computing median | np.median(df["Average"]) |
| Computing standard deviation | np.std(df["Average"]) |
| Finding max/min | np.max(), np.min() |
| Correlation | np.corrcoef(df["Attendance"], df["Average"]) |
| Filtering with conditions | df["Average"].values > 80 |
Example:
import numpy as np
mean_avg = np.mean(df["Average"])
std_avg = np.std(df["Average"])
max_avg = np.max(df["Average"])
correlation = np.corrcoef(df["Attendance"], df["Average"])[0, 1]
print(f"Mean:{mean_avg:.2f}")
print(f"Std:{std_avg:.2f}")
print(f"Max:{max_avg:.2f}")
print(f"Correlation (Attendance vs Average):{correlation:.2f}")6. Complete Program
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# Step 1: Read data
df = pd.read_csv("student_performance.csv")
print(df.head())
print(df.info())
# Step 2: Calculate Total and Average
df["Total"] = df["Internal Marks"] + df["External Marks"]
df["Average"] = np.round(df["Total"] / 2, 2)
# Step 3: Find average by department
dept_avg = df.groupby("Department")["Average"].mean().reset_index()
print(dept_avg)
# Step 4: Identify low attendance students
low_att = df[df["Attendance"] < 75]
print(f"Low attendance:{len(low_att)} students")
# Step 5: Visualize
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
# Department-wise average
axes[0].bar(dept_avg["Department"], dept_avg["Average"], color="steelblue")
axes[0].set_title("Department-wise Average Marks")
axes[0].set_xlabel("Department")
axes[0].set_ylabel("Average Marks")
# Result distribution
result_counts = df["Result"].value_counts()
axes[1].pie(result_counts, labels=result_counts.index, autopct="%1.1f%%")
axes[1].set_title("Result Distribution")
plt.tight_layout()
plt.savefig("college_analysis.png")
plt.show()
# Step 6: Export
dept_avg.to_csv("department_averages.csv", index=False)
low_att.to_csv("low_attendance.csv", index=False)7. Key Insights
- Department-wise averages identify strong and weak departments.
- Low attendance students may need intervention.
- Result distribution shows overall pass/fail rates.
- Correlation between attendance and marks indicates if attendance affects performance.
Q26 (Compare and Analyse). Construct a table comparing Pandas, NumPy, and Matplotlib using at least eight dimensions such as purpose, data structure, common functions, input data type, output type, speed, use case, and example code. After the table, explain how the three libraries work together in a real data analysis project.
Answer:
1. Comparison Table
| Dimension | Pandas | NumPy | Matplotlib |
|---|---|---|---|
| Purpose | Tabular data analysis and manipulation | Numerical computation and array operations | Data visualization |
| Data Structure | Series (1D), DataFrame (2D) | ndarray (N-dimensional) | Figure, Axes |
| Common Functions | read_csv(), head(), merge(), groupby(), fillna() | array(), zeros(), sort(), where(), concatenate() | plot(), bar(), hist(), scatter(), pie() |
| Input Data Type | CSV, JSON, Excel, SQL, dict, list | Lists, tuples, arrays, scalars | Lists, arrays, Series, DataFrame columns |
| Output Type | DataFrame, Series | ndarray, scalar | Charts, plots, figures |
| Speed | Moderate (built on NumPy) | Very fast (C-based) | Moderate (rendering overhead) |
| Memory Efficiency | Moderate | High (contiguous storage) | N/A (visualization) |
| Use Case | Data cleaning, analysis, aggregation | Mathematical operations, linear algebra | Visual communication, EDA |
| Learning Curve | Moderate | Low to moderate | Low |
| Integration | Built on NumPy; integrates with Matplotlib | Foundation for Pandas, SciPy, ML | Works with Pandas and NumPy |
| Example Code | df.groupby("Dept")["Marks"].mean() | np.mean(arr) | plt.bar(x, y) |
| Handles Missing Data | Yes (dropna(), fillna()) | No (uses NaN) | N/A |
| Labelled Data | Yes (row/column labels) | No (position-based) | N/A |
| Typical Users | Data analysts, data scientists | Scientists, engineers | Analysts, communicators |
2. How the Three Libraries Work Together
In a real data analysis project, the three libraries form a complete pipeline:
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ PANDAS │ → │ NUMPY │ → │ MATPLOTLIB │
│ Load & Clean │ │ Compute & │ │ Visualize & │
│ Data │ │ Analyze │ │ Communicate │
└──────────────┘ └──────────────┘ └──────────────┘3. Real-World Data Analysis Project
Scenario: Analyze sales data from a retail company.
Step 1: Load Data (Pandas)
import pandas as pd
df = pd.read_csv("sales.csv")
print(df.head())
print(df.info())Step 2: Clean Data (Pandas)
df = df.dropna()
df = df.drop_duplicates()
df["Total"] = df["Quantity"] * df["Price"]Step 3: Compute Statistics (NumPy)
import numpy as np
mean_sales = np.mean(df["Total"])
std_sales = np.std(df["Total"])
max_sales = np.max(df["Total"])Step 4: Aggregate (Pandas)
monthly = df.groupby("Month")["Total"].sum().reset_index()
product_avg = df.groupby("Product")["Total"].mean().reset_index()Step 5: Visualize (Matplotlib)
import matplotlib.pyplot as plt
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
# Monthly sales trend
axes[0].plot(monthly["Month"], monthly["Total"], marker="o")
axes[0].set_title("Monthly Sales Trend")
axes[0].set_xlabel("Month")
axes[0].set_ylabel("Total Sales")
# Product-wise average
axes[1].bar(product_avg["Product"], product_avg["Total"], color="steelblue")
axes[1].set_title("Product-wise Average Sales")
axes[1].set_xlabel("Product")
axes[1].set_ylabel("Average Sales")
plt.tight_layout()
plt.savefig("sales_analysis.png")
plt.show()Step 6: Export Results (Pandas)
monthly.to_csv("monthly_sales.csv", index=False)
product_avg.to_csv("product_averages.csv", index=False)4. Role of Each Library in the Project
| Stage | Library | Role |
|---|---|---|
| Data Loading | Pandas | Read CSV/JSON/Excel into DataFrames |
| Data Cleaning | Pandas | Handle missing values, duplicates |
| Feature Creation | Pandas | Create new columns (Total, Average) |
| Numerical Operations | NumPy | Compute mean, std, max, min |
| Aggregation | Pandas | Group by month, product, department |
| Visualization | Matplotlib | Create line, bar, pie, scatter charts |
| Export | Pandas | Write results to CSV/JSON |
5. Benefits of Using All Three Together
| Benefit | Explanation |
|---|---|
| Complete Pipeline | From raw data to actionable insights |
| Efficiency | NumPy provides fast computation; Pandas provides structured handling |
| Clarity | Matplotlib communicates results visually |
| Flexibility | Each library handles its strengths |
| Integration | Pandas is built on NumPy; Matplotlib accepts Pandas/NumPy data |
| Reproducibility | Code can be rerun on new data |
| Scalability | Handles small to large datasets |
6. Summary
| Library | Strengths | Weaknesses |
|---|---|---|
| Pandas | Tabular data, labelled indexing, I/O | Slower than NumPy for pure numerics |
| NumPy | Fast numerics, vectorization | No labels, no missing value handling |
| Matplotlib | Flexible visualization | Verbose syntax, static by default |
Key Takeaway: These three libraries are complementary. Pandas handles structured data, NumPy handles numerical computation, and Matplotlib handles visualization. Together, they form the foundation of Python data analysis.
Summary of All Answers
| Part | Q# | Topic | Key Answer |
|---|---|---|---|
| A | 1 | Tabular data library | (b) Pandas |
| A | 2 | Series definition | (b) One-dimensional labelled array |
| A | 3 | Read CSV | (a) pd.read_csv() |
| A | 4 | NumPy array object | (c) ndarray |
| A | 5 | Sort array | (b) np.sort() |
| A | 6 | Display chart | (c) plt.show() |
| A | 7 | Trend over time | (b) Line plot |
| A | 8 | HTML parsing | (a) BeautifulSoup |
| B | 9 | Pandas advantages | Data loading, cleaning, selection, transformation, aggregation |
| B | 10 | Series vs DataFrame | 1D vs 2D; single column vs full table |
| B | 11 | loc[] vs iloc[] | Label-based vs position-based |
| B | 12 | Missing values | dropna() removes; fillna() replaces |
| B | 13 | CSV and JSON I/O | read_csv(), to_csv(), read_json(), to_json() |
| B | 14 | ndarray attributes | shape, ndim, size, dtype |
| B | 15 | Indexing and slicing | 1D and 2D examples |
| B | 16 | Joining and splitting | concatenate(), vstack(), array_split(), reshape() |
| B | 17 | Search, sort, filter | where(), sort(), boolean masks |
| B | 18 | Web scraping | Workflow, tools, ethics |
| B | 19 | Visualization need | Trends, comparisons, distributions, relationships |
| C | 20 | Pandas workflow | Load → Inspect → Clean → Transform → Combine → Export |
| C | 21 | NumPy in detail | Creation, attributes, indexing, operations, random |
| C | 22 | Matplotlib charts | Line, bar, histogram, scatter, pie |
| C | 23 | Mini-project | Complete pipeline with all three libraries |
| C | 24 | Web scraping | Workflow, tools, applications, limitations, ethics |
| D | 25 | College case study | Pandas + NumPy + Matplotlib solution |
| D | 26 | Comparison table | 14-dimension comparison + integration |