BTCE | 5th Sem
Adv-Python SubjectExtra Questions

Adv-Py: Unit 4 (Book-back Questions)

Generated and Prepared By Thiruselvan (ThiruXD)

PART A — MULTIPLE CHOICE QUESTIONS (1 Mark Each)


Q1. Which library is mainly used for tabular data analysis in Python?

  1. NumPy
  2. Pandas
  3. Matplotlib
  4. Requests

Answer: (b) Pandas

Explanation: Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures (Series and DataFrame) for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs. NumPy is for numerical computation; Matplotlib is for visualization; Requests is for HTTP requests.


Q2. A Pandas Series is best described as:

  1. A two-dimensional table
  2. A one-dimensional labelled array
  3. A plotting function
  4. A web scraping tool

Answer: (b) A one-dimensional labelled array

Explanation: A Series is a one-dimensional labelled array in Pandas. It can store integers, floats, strings, or other Python objects, and is similar to a single column in a spreadsheet. A DataFrame is the two-dimensional structure.


Q3. Which function reads a CSV file into a DataFrame?

  1. pd.read_csv()
  2. pd.open_csv()
  3. np.read_csv()
  4. plt.csv()

Answer: (a) pd.read_csv()

Explanation: pd.read_csv() is the Pandas function for reading CSV files into a DataFrame. pd.open_csv() and plt.csv() do not exist; np.read_csv() is not a NumPy function (NumPy uses np.loadtxt() or np.genfromtxt()).


Q4. What is the main array object in NumPy called?

  1. DataFrame
  2. Series
  3. ndarray
  4. pyplot

Answer: (c) ndarray

Explanation: The main array object in NumPy is called ndarray (N-dimensional array). It stores elements of the same data type in a compact and efficient way. DataFrame and Series are Pandas objects; pyplot is a Matplotlib module.


Q5. Which NumPy function is used to sort an array?

  1. np.order()
  2. np.sort()
  3. np.filter()
  4. np.arange()

Answer: (b) np.sort()

Explanation: np.sort() returns a sorted copy of the array. np.order() does not exist; np.filter() does not exist; np.arange() creates values in a range.


Q6. Which Matplotlib function is used to display a chart?

  1. plt.display()
  2. plt.view()
  3. plt.show()
  4. plt.open()

Answer: (c) plt.show()

Explanation: plt.show() displays the figure on screen. plt.display(), plt.view(), and plt.open() are not Matplotlib functions.


Q7. Which chart is most suitable for showing trend over time?

  1. Pie chart
  2. Line plot
  3. Histogram
  4. Box plot

Answer: (b) Line plot

Explanation: Line plots connect data points chronologically to show trends over time. Pie charts show parts of a whole; histograms show distributions; box plots show five-number summaries.


Q8. Which library is commonly used for parsing HTML in basic web scraping?

  1. BeautifulSoup
  2. NumPy
  3. Matplotlib
  4. json

Answer: (a) BeautifulSoup

Explanation: BeautifulSoup is a Python library for parsing HTML and XML documents. It is commonly used with requests for web scraping. NumPy is for numerical computation; Matplotlib is for visualization; json is for JSON data handling.


PART B — SHORT ANSWER QUESTIONS (5 Marks Each)


Q9. Define Pandas. Explain any five advantages of using Pandas for data analysis.

Answer:

Definition:Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures (Series and DataFrame) and functions for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs. The name Pandas is derived from Panel Data, which refers to multidimensional structured datasets.

Five Advantages of Pandas:

#AdvantageExplanationExample
1Data LoadingReads data from many file formats and sourcesread_csv(), read_json(), read_excel()
2Data CleaningHandles missing values, duplicates, and incorrect formatsdropna(), fillna(), drop_duplicates()
3Data SelectionSelects rows, columns, and subsets efficientlydf["Name"], loc[], iloc[]
4Data TransformationCreates new columns and modifies existing datadf["Total"] = df["A"] + df["B"]
5Data AggregationGroups and summarizes data easilygroupby(), mean(), sum()
6Data ExportWrites results back to filesto_csv(), to_json()

Additional advantages: Handles mixed data types, supports labelled indexing, integrates well with NumPy and Matplotlib, and is faster than using basic Python lists and dictionaries for tabular data.


Q10. Differentiate between Pandas Series and DataFrame with suitable examples.

Answer:

AspectPandas SeriesPandas DataFrame
DefinitionOne-dimensional labelled arrayTwo-dimensional labelled data structure
Dimensions1D2D (rows × columns)
Similar toA single column in a spreadsheetA complete table with rows and columns
IndexSingle indexRow index + column index
Data TypesHomogeneous (single dtype)Heterogeneous (mixed dtypes)
Creationpd.Series([1, 2, 3])pd.DataFrame({"A": [1, 2], "B": [3, 4]})
Accessseries[0], series["label"]df["col"], df.loc[], df.iloc[]
Use CaseSingle variable analysisMulti-variable tabular analysis

Example — Series:

import pandas as pd

marks = pd.Series([78, 85, 92, 66], index=["Amit", "Ravi", "Neha", "Sara"])
print(marks)
print(marks["Neha"])   # 92

Output:

Amit    78
Ravi    85
Neha    92
Sara    66
dtype: int64
92

Example — DataFrame:

import pandas as pd

data = {
    "Name": ["Amit", "Ravi", "Neha", "Sara"],
    "Age": [21, 22, 20, 23],
    "Marks": [78, 85, 92, 66]
}
df = pd.DataFrame(data)
print(df)

Output:

   Name  Age  Marks
0  Amit   21     78
1  Ravi   22     85
2  Neha   20     92
3  Sara   23     66

Key Difference: A Series is a single column; a DataFrame is a collection of Series sharing the same row index.


Q11. Explain how to select rows and columns from a DataFrame using loc[] and iloc[].

Answer:

Pandas provides two primary methods for selecting rows and columns:

  • loc[] — label-based indexing (uses row/column labels)
  • iloc[] — integer-position-based indexing (uses numeric positions)

Syntax:

  • df.loc[row_label, column_label]
  • df.iloc[row_position, column_position]

Example DataFrame:

import pandas as pd

data = {
    "Name": ["Amit", "Ravi", "Neha", "Sara"],
    "Age": [21, 22, 20, 23],
    "Marks": [78, 85, 92, 66]
}
df = pd.DataFrame(data)

Using loc[] (Label-Based):

# Select row with label 0
print(df.loc[0])

# Select rows 0 to 2 (inclusive)
print(df.loc[0:2])

# Select specific rows and columns
print(df.loc[0:2, ["Name", "Marks"]])

# Conditional selection
print(df.loc[df["Marks"] > 80])

Using iloc[] (Position-Based):

# Select first row
print(df.iloc[0])

# Select rows 0 to 2 (exclusive of 3)
print(df.iloc[0:3])

# Select rows 0-2, columns 0-1
print(df.iloc[0:3, 0:2])

# Select last row
print(df.iloc[-1])

Comparison Table:

MethodBased OnExampleResult
df.loc[0]Labeldf.loc[0]Row with index label 0
df.iloc[0]Positiondf.iloc[0]First row
df.loc[0:2]Label (inclusive)df.loc[0:2]Rows 0, 1, 2
df.iloc[0:3]Position (exclusive)df.iloc[0:3]Rows 0, 1, 2
df.loc[:, "Name"]Labeldf.loc[:, "Name"]All rows, “Name” column
df.iloc[:, 0]Positiondf.iloc[:, 0]All rows, first column

Key Difference: loc[] uses labels (inclusive of endpoint); iloc[] uses integer positions (exclusive of endpoint, like Python slicing).


Q12. Write short notes on handling missing values in Pandas using dropna() and fillna().

Answer:

Missing values are common in real datasets. Pandas represents missing values using NaN (Not a Number). Two primary functions handle missing values:

1. dropna() — Remove Missing Values

Purpose: Removes rows or columns containing missing values.

Syntax:

df.dropna(axis=0, how='any', thresh=None, subset=None, inplace=False)

Parameters:

ParameterMeaningDefault
axis0 = drop rows; 1 = drop columns0
how‘any’ = drop if any NaN; ‘all’ = drop if all NaN‘any’
threshMinimum non-NaN values required to keepNone
subsetColumns to considerNone
inplaceModify DataFrame in placeFalse

Example:

import pandas as pd

data = {"Name": ["Amit", "Ravi", "Neha"], "Marks": [78, None, 92]}
df = pd.DataFrame(data)

print(df.dropna())         # Remove rows with any NaN
print(df.dropna(axis=1))   # Remove columns with any NaN

2. fillna() — Replace Missing Values

Purpose: Replaces missing values with a specified value or method.

Syntax:

df.fillna(value=None, method=None, axis=None, inplace=False)

Parameters:

ParameterMeaning
valueValue to replace NaN with (scalar, dict, Series)
method‘ffill’ (forward fill) or ‘bfill’ (backward fill)
axis0 = rows; 1 = columns
inplaceModify DataFrame in place

Example:

print(df.fillna(0))              # Replace NaN with 0
print(df.fillna(df.mean()))      # Replace with column mean
print(df.fillna(method='ffill')) # Forward fill

3. Other Missing Value Functions

FunctionPurpose
isnull()Detects missing values (returns True/False)
notnull()Detects non-missing values
drop_duplicates()Removes duplicate rows

Complete Example:

import pandas as pd

data = {"Name": ["Amit", "Ravi", "Neha"], "Marks": [78, None, 92]}
df = pd.DataFrame(data)

print(df.isnull())        # Check missing values
print(df.dropna())        # Remove rows with missing values
print(df.fillna(0))       # Replace missing values with 0
print(df.fillna(df["Marks"].mean()))  # Replace with mean

Best Practice: Choose the method based on the data and analysis goal. Dropping is suitable when missing data is random and small in proportion; filling is better when missing data is substantial or when preserving rows is important.


Q13. Explain how CSV and JSON files can be read and written using Pandas.

Answer:

Pandas can read data from many file formats. Two of the most common formats are CSV (Comma-Separated Values) and JSON (JavaScript Object Notation).

1. Reading CSV Files

Syntax: pd.read_csv(filepath, sep=',', header='infer', names=None, usecols=None)

import pandas as pd

# Read a CSV file
students = pd.read_csv("students.csv")

# Inspect the data
print(students.head())     # First 5 rows
print(students.shape)      # (rows, columns)
print(students.info())     # Column types and non-null counts
print(students.describe()) # Statistical summary

Common Parameters:

ParameterPurpose
sepDelimiter (default ‘,’)
headerRow to use as column names
namesCustom column names
usecolsSubset of columns to read
dtypeData types for columns

2. Writing CSV Files

Syntax: df.to_csv(filepath, index=False)

# Save a DataFrame to a CSV file
students.to_csv("cleaned_students.csv", index=False)

Common Parameters:

ParameterPurpose
indexInclude row index (default True)
columnsSubset of columns to write
sepDelimiter
headerInclude column names

3. Reading JSON Files

Syntax: pd.read_json(filepath, orient=None)

import pandas as pd

# Read a JSON file
orders = pd.read_json("orders.json")
print(orders.head())

Common Parameters:

ParameterPurpose
orientJSON orientation: ‘records’, ‘index’, ‘columns’, ‘values’, ‘split’, ‘table’
linesRead line-delimited JSON

4. Writing JSON Files

Syntax: df.to_json(filepath, orient='records', indent=4)

# Save a DataFrame to a JSON file
orders.to_json("cleaned_orders.json", orient="records", indent=4)

Common Parameters:

ParameterPurpose
orientJSON orientation
indentPretty-print indentation
date_formatFormat for datetime values

5. File Reading and Writing Functions Summary

FunctionPurposeCommon Parameter
read_csv()Reads a CSV file into a DataFramesep, header, names, usecols
to_csv()Writes a DataFrame to a CSV fileindex=False
read_json()Reads a JSON file into a DataFrameorient
to_json()Writes a DataFrame to a JSON fileorient, indent

Practical Tip: After reading any file, always use head(), shape, info(), and describe() to understand the structure and quality of the imported data.


Q14. Define NumPy ndarray. Explain shape, ndim, size, and dtype attributes.

Answer:

Definition:NumPy (Numerical Python) is a fundamental Python library for scientific computing and numerical operations. It provides the ndarray object (N-dimensional array), which stores elements of the same data type in a compact and efficient way.

Compared with normal Python lists, NumPy arrays are:

  • Faster — optimized C implementation
  • More memory efficient — contiguous storage
  • Support vectorized operations — operations on entire arrays without explicit loops

Creating an ndarray:

import numpy as np

arr = np.array([10, 20, 30, 40])
print(arr)          # [10 20 30 40]
print(type(arr))    # <class 'numpy.ndarray'>

Attributes of ndarray:

1. shape

Returns the dimensions of the array as a tuple (rows, columns, …).

a = np.array([[1, 2, 3], [4, 5, 6]])
print(a.shape)   # (2, 3) — 2 rows, 3 columns

2. ndim

Returns the number of dimensions (axes) of the array.

print(a.ndim)    # 2 — two-dimensional array

3. size

Returns the total number of elements in the array.

print(a.size)    # 6 — 2 × 3 = 6 elements

4. dtype

Returns the data type of the elements in the array.

print(a.dtype)   # int64 (or int32 depending on platform)

Complete Example:

import numpy as np

a = np.array([[1, 2, 3], [4, 5, 6]])

print(a.shape)   # (2, 3)
print(a.ndim)    # 2
print(a.size)    # 6
print(a.dtype)   # int64

Summary Table:

AttributeMeaningExample Output
shapeDimensions of the array(2, 3)
ndimNumber of dimensions2
sizeTotal number of elements6
dtypeData type of elementsint64

Other Useful Attributes:

AttributeMeaning
itemsizeSize in bytes of each element
nbytesTotal bytes consumed by the array
TTransposed view of the array

Why These Attributes Matter: They help you understand the structure and memory layout of your data before performing operations. For example, shape is essential for reshaping, and dtype determines which operations are valid.


Q15. Explain indexing and slicing in one-dimensional and two-dimensional NumPy arrays.

Answer:

Indexing is used to access individual elements, while slicing is used to access a range or subset of elements. NumPy indexing starts from 0, just like normal Python lists.

1. One-Dimensional Arrays

Syntax: array[start:stop:step]

import numpy as np

a = np.array([10, 20, 30, 40, 50])

print(a[0])      # 10 — first element
print(a[-1])     # 50 — last element
print(a[1:4])    # [20, 30, 40] — elements from index 1 to 3
print(a[:3])     # [10, 20, 30] — first three elements
print(a[::2])    # [10, 30, 50] — every second element
print(a[::-1])   # [50, 40, 30, 20, 10] — reversed

Indexing and Slicing Examples:

ExpressionMeaningResult
a[0]First element10
a[-1]Last element50
a[1:4]Elements from index 1 to 3[20, 30, 40]
a[:3]First three elements[10, 20, 30]
a[::2]Every second element[10, 30, 50]

2. Two-Dimensional Arrays

Syntax: array[row, column] or array[row_start:row_stop, col_start:col_stop]

import numpy as np

matrix = np.array([[1, 2, 3],
                   [4, 5, 6],
                   [7, 8, 9]])

print(matrix[0, 0])       # 1 — row 0, column 0
print(matrix[1, 2])       # 6 — row 1, column 2
print(matrix[:, 1])       # [2, 5, 8] — all rows, column 1
print(matrix[0:2, 1:3])   # Sub-array [[2, 3], [5, 6]]
print(matrix[1, :])       # [4, 5, 6] — row 1, all columns
print(matrix[:, 0])       # [1, 4, 7] — all rows, column 0

2D Indexing and Slicing Examples:

ExpressionMeaningResult
matrix[0, 0]Element at row 0, column 01
matrix[1, 2]Element at row 1, column 26
matrix[:, 1]All rows, column 1[2, 5, 8]
matrix[0:2, 1:3]Rows 0-1, columns 1-2[[2, 3], [5, 6]]
matrix[1, :]Row 1, all columns[4, 5, 6]
matrix[:, 0]All rows, column 0[1, 4, 7]

3. Key Points

  • Indexing accesses a single element: a[0], matrix[1, 2].
  • Slicing accesses a range: a[1:4], matrix[0:2, 1:3].
  • Colon (:) means “all” in that dimension.
  • Negative indices count from the end: a[-1] is the last element.
  • Step can be used: a[::2] takes every second element.
  • Slicing returns a view, not a copy — modifying the slice modifies the original array.

Example of View vs Copy:

a = np.array([1, 2, 3, 4, 5])
b = a[1:4]     # b is a view
b[0] = 99
print(a)       # [1, 99, 3, 4, 5] — original modified

To create a copy, use .copy():

b = a[1:4].copy()

Q16. Discuss joining and splitting operations in NumPy with examples.

Answer:

NumPy provides functions to join (combine) and split (divide) arrays. These operations are essential for data manipulation and reshaping.

1. Joining Arrays

(a) concatenate() — Join Along an Existing Axis

import numpy as np

a = np.array([1, 2, 3])
b = np.array([4, 5, 6])

joined = np.concatenate((a, b))
print(joined)   # [1 2 3 4 5 6]

For 2D arrays:

x = np.array([[1, 2], [3, 4]])
y = np.array([[5, 6], [7, 8]])

print(np.concatenate((x, y), axis=0))  # Vertical (rows)
print(np.concatenate((x, y), axis=1))  # Horizontal (columns)

(b) vstack() — Vertical Stacking

Stacks arrays vertically (row-wise).

x = np.array([[1, 2], [3, 4]])
y = np.array([[5, 6], [7, 8]])

print(np.vstack((x, y)))
# [[1 2]
#  [3 4]
#  [5 6]
#  [7 8]]

(c) hstack() — Horizontal Stacking

Stacks arrays horizontally (column-wise).

print(np.hstack((x, y)))
# [[1 2 5 6]
#  [3 4 7 8]]

2. Splitting Arrays

(a) array_split() — Split into Multiple Parts

import numpy as np

a = np.array([1, 2, 3, 4, 5, 6])

parts = np.array_split(a, 3)
print(parts)
# [array([1, 2]), array([3, 4]), array([5, 6])]

(b) split() — Equal Division

parts = np.split(a, 3)
print(parts)   # Same as above but requires equal division

(c) hsplit() and vsplit() — Split 2D Arrays

matrix = np.array([[1, 2, 3, 4], [5, 6, 7, 8]])

print(np.hsplit(matrix, 2))   # Split horizontally
print(np.vsplit(matrix, 2))   # Split vertically

3. Reshaping

reshape() — Change Shape Without Changing Data

a = np.array([1, 2, 3, 4, 5, 6])
reshaped = a.reshape(2, 3)
print(reshaped)
# [[1 2 3]
#  [4 5 6]]

4. Summary Table

FunctionPurposeExample
concatenate()Joins arrays along an existing axisnp.concatenate((a, b))
vstack()Stacks arrays verticallynp.vstack((x, y))
hstack()Stacks arrays horizontallynp.hstack((x, y))
array_split()Splits an array into multiple partsnp.array_split(a, 3)
split()Splits into equal partsnp.split(a, 3)
hsplit()Splits horizontallynp.hsplit(matrix, 2)
vsplit()Splits verticallynp.vsplit(matrix, 2)
reshape()Changes shape without changing dataa.reshape(2, 3)

5. Practical Example

import numpy as np

# Joining
sem1 = np.array([85, 78, 92])
sem2 = np.array([88, 82, 95])
all_marks = np.concatenate((sem1, sem2))
print("Joined:", all_marks)

# Splitting
data = np.array([10, 20, 30, 40, 50, 60])
parts = np.array_split(data, 3)
print("Split:", parts)

# Reshaping
matrix = np.arange(1, 13).reshape(3, 4)
print("Reshaped:\n", matrix)

Output:

Joined: [85 78 92 88 82 95]
Split: [array([10, 20]), array([30, 40]), array([50, 60])]
Reshaped:
 [[ 1  2  3  4]
  [ 5  6  7  8]
  [ 9 10 11 12]]

Q17. Explain searching, sorting, and filtering arrays in NumPy.

Answer:

NumPy provides efficient functions for searching, sorting, and filtering arrays. These operations are essential for extracting useful information from data.

1. Searching Arrays

np.where() — Find Indices Where Condition is True

import numpy as np

a = np.array([10, 20, 30, 20, 40])
result = np.where(a == 20)
print(result)   # (array([1, 3]),)

With conditions:

print(np.where(a > 25))   # (array([2, 4]),)

np.searchsorted() — Find Insertion Position

sorted_arr = np.array([10, 20, 30, 40, 50])
print(np.searchsorted(sorted_arr, 35))   # 3 (insert before index 3)

np.argmax() and np.argmin() — Index of Max/Min

a = np.array([10, 50, 30, 20])
print(np.argmax(a))   # 1 (index of 50)
print(np.argmin(a))   # 0 (index of 10)

2. Sorting Arrays

np.sort() — Returns Sorted Copy

import numpy as np

a = np.array([40, 10, 30, 20])
print(np.sort(a))   # [10 20 30 40]

array.sort() — Sorts In-Place

a = np.array([40, 10, 30, 20])
a.sort()
print(a)   # [10 20 30 40]

Sorting 2D Arrays

matrix = np.array([[3, 1, 2], [6, 5, 4]])
print(np.sort(matrix, axis=0))   # Sort each column
print(np.sort(matrix, axis=1))   # Sort each row

np.argsort() — Indices That Would Sort the Array

a = np.array([40, 10, 30, 20])
print(np.argsort(a))   # [1 3 2 0]

3. Filtering Arrays

Filtering uses boolean masks — conditions that produce True/False for each element.

import numpy as np

a = np.array([10, 25, 30, 45, 50])

# Filter values greater than 30
filtered = a[a > 30]
print(filtered)   # [45 50]

# Filter with multiple conditions
filtered = a[(a > 20) & (a < 50)]
print(filtered)   # [25 30 45]

# Filter even numbers
evens = a[a % 2 == 0]
print(evens)   # [10 30 50]

Boolean Mask:

mask = a > 30
print(mask)   # [False False False  True  True]
print(a[mask])  # [45 50]

4. Summary Table

OperationExampleOutput Meaning
Searchnp.where(a == 20)Returns indexes where condition is true
Sortnp.sort(a)Returns sorted copy of the array
Filtera[a > 30]Returns values greater than 30
Boolean maska % 2 == 0Creates True/False condition for each value
Argmaxnp.argmax(a)Index of maximum value
Argminnp.argmin(a)Index of minimum value
Argsortnp.argsort(a)Indices that would sort the array

5. Complete Example

import numpy as np

marks = np.array([78, 85, 92, 66, 74, 88, 95, 60])

# Searching
print("Students scoring 85:", np.where(marks == 85))

# Sorting
print("Sorted marks:", np.sort(marks))
print("Ranking indices:", np.argsort(marks))

# Filtering
print("Marks above 80:", marks[marks > 80])
print("Marks between 70 and 90:", marks[(marks >= 70) & (marks <= 90)])

# Top scorer
print("Top scorer marks:", marks.max())
print("Top scorer index:", marks.argmax())

Output:

Students scoring 85: (array([1]),)
Sorted marks: [60 66 74 78 85 88 92 95]
Ranking indices: [7 3 4 0 1 5 2 6]
Marks above 80: [85 92 88 95]
Marks between 70 and 90: [78 85 74 88]
Top scorer marks: 95
Top scorer index: 6

Important: Filtering in NumPy is usually done using boolean conditions. The condition creates a True/False mask, and the mask is used to select matching elements.


Q18. What is web scraping? Explain the basic steps and ethical precautions involved.

Answer:

Definition:Web scraping is the process of extracting data from websites using programs. It is useful when information is available on web pages but not provided as a downloadable dataset or API.

1. Basic Web Scraping Workflow

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Send HTTP │ →  │ 2. Download  │ →  │ 3. Parse and │ →  │ 4. Clean and │ →  │ 5. Store or  │
│   Request    │    │    HTML      │    │ Extract Data │    │  Structure   │    │  Export Data │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘

2. Steps in a Basic Web Scraping Workflow

StepMeaningCommon Tool
1. Send HTTP RequestSend a request to the target websiterequests
2. Download HTMLRetrieve the HTML content of the pageresponse.text
3. Parse and Extract DataParse the HTML and extract the required dataBeautifulSoup
4. Clean and Structure DataClean the extracted data and structure it for usePandas, manual cleaning
5. Store or Export DataSave the data to a file, database, or export it for analysisCSV, JSON, Pandas DataFrame

3. Simple Web Scraping Example

import requests
from bs4 import BeautifulSoup

url = "https://example.com"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")

print(soup.title.text)

Explanation:

  • requests.get(url) sends an HTTP GET request.
  • response.text contains the HTML source.
  • BeautifulSoup(response.text, "html.parser") parses the HTML.
  • soup.title.text extracts the page title.

4. Common Tools

ToolPurpose
requestsSend HTTP requests
BeautifulSoupParse HTML/XML
lxmlFaster HTML/XML parser
SeleniumBrowser automation for dynamic pages
ScrapyFull-featured scraping framework

5. Ethical Precautions

Web scraping should be performed responsibly. Programmers must respect:

PrecautionExplanation
Terms of ServiceCheck if the website allows automated access
robots.txtFollow the rules specified in the site’s robots.txt file
CopyrightDo not reproduce copyrighted content without permission
PrivacyDo not collect personal or restricted data without consent
Server LoadAvoid sending too many requests that could overload the server
Official APIsPrefer official APIs whenever available
Rate LimitingAdd delays between requests
User-AgentIdentify your bot transparently

Ethical Reminder: Before scraping a website, check whether the website allows automated access. Prefer official APIs whenever available.

6. Applications of Web Scraping

  • Price comparison: Track product prices across e-commerce sites.
  • News aggregation: Collect headlines from multiple news sources.
  • Sentiment analysis: Gather social media posts for analysis.
  • Research: Collect data for academic studies.
  • Lead generation: Gather business contact information.
  • Real estate: Track property listings and prices.

Q19. Explain the need for data visualization and list common Matplotlib chart types.

Answer:

Need for Data Visualization: Data visualization converts numbers into patterns. A good chart can reveal:

  • Trends — how values change over time
  • Comparisons — differences between categories
  • Distributions — how data is spread
  • Relationships — connections between variables
  • Outliers — unusual data points

These may not be clear from a table alone. Visualization also helps communicate insights to non-technical audiences.

Common Matplotlib Chart Types:

Chart TypeUsed ForExample Question
Line plotShowing trends over timeHow did sales change month by month?
Bar chartComparing categoriesWhich department scored highest?
HistogramShowing distribution of numerical valuesWhat is the distribution of exam marks?
Scatter plotShowing relationship between two numerical variablesIs study time related to marks?
Pie chartShowing parts of a wholeWhat percentage of students selected each elective?

Common Matplotlib pyplot Functions:

FunctionPurpose
plot()Creates a line plot
bar()Creates a bar chart
hist()Creates a histogram
scatter()Creates a scatter plot
pie()Creates a pie chart
title()Adds chart title
xlabel(), ylabel()Adds axis labels
legend()Displays chart legend
show()Displays the figure
savefig()Saves chart as an image file

Visualization Rule: Every chart should have a clear title, axis labels where applicable, readable category names, and a purpose. Avoid adding unnecessary decoration.

Matplotlib Visualization Workflow

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Prepare   │ →  │ 2. Create    │ →  │ 3. Plot Data │ →  │ 4. Customize │ →  │ 5. Show or   │
│    Data      │    │   Figure     │    │              │    │    Plot      │    │    Save      │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘
StepActionCode Example
1. Prepare DataLoad and organize data for visualizationx = [1, 2, 3], y = [10, 15, 13]
2. Create FigureCreate a figure and axesplt.figure(figsize=(8, 5))
3. Plot DataPlot data using the appropriate plot typeplt.plot(x, y)
4. Customize PlotAdd titles, labels, legends, colorsplt.title("Title"), plt.xlabel("X")
5. Show or SaveDisplay or save the plotplt.show(), plt.savefig("plot.png")

Complete Example:

import matplotlib.pyplot as plt

# Step 1: Prepare Data
x = [1, 2, 3, 4, 5]
y = [10, 15, 13, 18, 20]

# Step 2: Create Figure (optional)
plt.figure(figsize=(8, 5))

# Step 3: Plot Data
plt.plot(x, y)

# Step 4: Customize Plot
plt.title("Simple Line Plot")
plt.xlabel("X Values")
plt.ylabel("Y Values")

# Step 5: Show or Save
plt.show()

Chart Examples:

Bar Chart:

subjects = ["Python", "Maths", "DBMS", "OS"]
marks = [85, 78, 92, 74]
plt.bar(subjects, marks)
plt.title("Marks by Subject")
plt.show()

Histogram:

marks = [45, 56, 67, 78, 89, 90, 72, 63, 55, 81]
plt.hist(marks, bins=5)
plt.title("Distribution of Marks")
plt.show()

Scatter Plot:

hours = [1, 2, 3, 4, 5, 6]
marks = [40, 50, 55, 65, 75, 85]
plt.scatter(hours, marks)
plt.title("Study Hours vs Marks")
plt.show()

Pie Chart:

labels = ["Python", "Java", "C++", "R"]
students = [40, 25, 20, 15]
plt.pie(students, labels=labels, autopct="%1.1f%%")
plt.title("Programming Language Preference")
plt.show()

PART C — LONG ANSWER / ESSAY QUESTIONS (10 Marks Each)


Q20. Explain the complete Pandas workflow for loading, inspecting, cleaning, transforming, combining, and exporting tabular data. Support your answer with suitable code examples.

Answer:

1. Introduction to Pandas

Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures (Series and DataFrame) and functions for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs.

2. Pandas Workflow Overview

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Import    │ →  │ 2. Create    │ →  │ 3. Clean &   │ →  │ 4. Combine   │ →  │ 5. Analyze   │
│    Data      │    │  Series/DF   │    │   Transform  │    │   Datasets   │    │  & Export    │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘

3. Step 1: Import Data

Pandas can read data from many file formats.

import pandas as pd

# Read CSV
df = pd.read_csv("students.csv")

# Read JSON
df = pd.read_json("students.json")

# Read Excel
df = pd.read_excel("students.xlsx")

Common Parameters:

ParameterPurpose
sepDelimiter
headerRow to use as column names
namesCustom column names
usecolsSubset of columns
dtypeData types

4. Step 2: Create and Inspect DataFrame

Creating a DataFrame:

import pandas as pd

data = {
    "Name": ["Amit", "Ravi", "Neha", "Sara"],
    "Age": [21, 22, 20, 23],
    "Marks": [78, 85, 92, 66]
}
df = pd.DataFrame(data)

Inspection Methods:

MethodPurpose
df.head()First 5 rows
df.tail()Last 5 rows
df.shape(rows, columns)
df.columnsColumn names
df.info()Column types and non-null counts
df.describe()Statistical summary
df.dtypesData types of columns
print(df.head())
print(df.shape)
print(df.info())
print(df.describe())

5. Step 3: Clean and Transform Data

Handling Missing Values:

# Detect
print(df.isnull())

# Remove
df = df.dropna()

# Fill
df = df.fillna(0)
df = df.fillna(df.mean())

Removing Duplicates:

df = df.drop_duplicates()

Selecting Rows and Columns:

# Select columns
df["Name"]
df[["Name", "Marks"]]

# Select rows by label
df.loc[0]
df.loc[df["Marks"] > 80]

# Select rows by position
df.iloc[0:3, 1:3]

Creating New Columns:

df["Total"] = df["Python"] + df["Maths"]
df["Average"] = df["Total"] / 2
df["Grade"] = df["Average"].apply(lambda x: "A" if x > 85 else "B")

Filtering:

high_scorers = df[df["Marks"] > 80]
low_attendance = df[df["Attendance"] < 75]

Sorting:

df = df.sort_values("Marks", ascending=False)

Grouping and Aggregation:

dept_avg = df.groupby("Department")["Marks"].mean()
dept_summary = df.groupby("Department").agg({
    "Marks": ["mean", "max", "min"],
    "Attendance": "mean"
})

6. Step 4: Combine DataFrames

Concatenation (stacking):

df1 = pd.DataFrame({"ID": [1, 2], "Name": ["Amit", "Ravi"]})
df2 = pd.DataFrame({"ID": [3, 4], "Name": ["Neha", "Sara"]})
combined = pd.concat([df1, df2], ignore_index=True)

Merging (common column):

marks = pd.DataFrame({"ID": [1, 2, 3], "Marks": [78, 85, 92]})
students = pd.DataFrame({"ID": [1, 2, 3], "Name": ["Amit", "Ravi", "Neha"]})
result = pd.merge(students, marks, on="ID")

Joining (index-based):

a = pd.DataFrame({"Marks": [78, 85]}, index=[1, 2])
b = pd.DataFrame({"Name": ["Amit", "Ravi"]}, index=[1, 2])
joined = a.join(b)

Combining Methods Summary:

OperationMeaningWhen to Use
concat()Stacks DataFramesSame columns or same index
merge()Combines using common columnTables share a key such as ID
join()Combines using indexIndex labels are meaningful

7. Step 5: Analyze and Export

Export to CSV:

df.to_csv("cleaned_students.csv", index=False)

Export to JSON:

df.to_json("cleaned_students.json", orient="records", indent=4)

Export to Excel:

df.to_excel("cleaned_students.xlsx", index=False)

8. Complete Integrated Example

import pandas as pd

# Step 1: Load
df = pd.read_csv("students.csv")

# Step 2: Inspect
print(df.head())
print(df.info())

# Step 3: Clean
df = df.dropna()
df = df.drop_duplicates()

# Step 4: Transform
df["Total"] = df["Internal"] + df["External"]
df["Average"] = df["Total"] / 2

# Step 5: Filter
top = df[df["Average"] > 80]

# Step 6: Group
dept_avg = df.groupby("Department")["Average"].mean()

# Step 7: Combine
extra = pd.read_csv("extra_students.csv")
combined = pd.concat([df, extra], ignore_index=True)

# Step 8: Export
combined.to_csv("final_results.csv", index=False)

print("Analysis complete.")

9. Benefits of the Pandas Workflow

  • Systematic: Each stage has clear inputs and outputs.
  • Reproducible: Code can be rerun on new data.
  • Maintainable: Easy to modify individual steps.
  • Integrated: Works with NumPy, Matplotlib, and scikit-learn.
  • Efficient: Handles large datasets faster than pure Python.

Q21. Discuss NumPy arrays in detail. Explain array creation, attributes, indexing, slicing, vectorized operations, joining, splitting, searching, sorting, filtering, and random number generation.

Answer:

1. Introduction to NumPy

NumPy (Numerical Python) is a fundamental Python library for scientific computing and numerical operations. It provides the ndarray object (N-dimensional array), which stores elements of the same data type in a compact and efficient way.

Advantages over Python lists:

  • Faster — optimized C implementation
  • More memory efficient — contiguous storage
  • Support vectorized operations — no explicit loops
  • Rich function library — linear algebra, statistics, random

2. Array Creation

FunctionPurposeExample
np.array()Creates array from list/tuplenp.array([1, 2, 3])
np.zeros()Array filled with zerosnp.zeros(5)
np.ones()Array filled with onesnp.ones((2, 3))
np.arange()Values in a rangenp.arange(1, 10, 2)
np.linspace()Evenly spaced valuesnp.linspace(0, 1, 5)
np.eye()Identity matrixnp.eye(3)
import numpy as np

a = np.array([10, 20, 30, 40])
b = np.zeros(5)
c = np.ones((2, 3))
d = np.arange(1, 10, 2)
e = np.linspace(0, 1, 5)
f = np.eye(3)

3. ndarray Attributes

AttributeMeaning
shapeDimensions (rows, columns)
ndimNumber of dimensions
sizeTotal number of elements
dtypeData type of elements
itemsizeSize in bytes of each element
nbytesTotal bytes consumed
a = np.array([[1, 2, 3], [4, 5, 6]])
print(a.shape)   # (2, 3)
print(a.ndim)    # 2
print(a.size)    # 6
print(a.dtype)   # int64

4. Indexing and Slicing

1D Arrays:

a = np.array([10, 20, 30, 40, 50])
print(a[0])      # 10
print(a[-1])     # 50
print(a[1:4])    # [20 30 40]
print(a[::2])    # [10 30 50]

2D Arrays:

matrix = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
print(matrix[0, 0])       # 1
print(matrix[1, 2])       # 6
print(matrix[:, 1])       # [2 5 8]
print(matrix[0:2, 1:3])   # [[2 3], [5 6]]

5. Vectorized Operations

Operations applied to entire arrays without explicit loops.

a = np.array([10, 20, 30])
b = np.array([1, 2, 3])

print(a + b)   # [11 22 33]
print(a - b)   # [9 18 27]
print(a * b)   # [10 40 90]
print(a / b)   # [10. 10. 10.]
print(a ** 2)  # [100 400 900]

6. Joining Arrays

FunctionPurpose
concatenate()Joins along an existing axis
vstack()Stacks vertically
hstack()Stacks horizontally
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
print(np.concatenate((a, b)))   # [1 2 3 4 5 6]

x = np.array([[1, 2], [3, 4]])
y = np.array([[5, 6], [7, 8]])
print(np.vstack((x, y)))
print(np.hstack((x, y)))

7. Splitting Arrays

FunctionPurpose
array_split()Split into multiple parts
split()Split into equal parts
hsplit()Split horizontally
vsplit()Split vertically
a = np.array([1, 2, 3, 4, 5, 6])
print(np.array_split(a, 3))
# [array([1, 2]), array([3, 4]), array([5, 6])]

8. Searching Arrays

FunctionPurpose
np.where()Indices where condition is true
np.argmax()Index of maximum value
np.argmin()Index of minimum value
np.searchsorted()Insertion position
a = np.array([10, 20, 30, 20, 40])
print(np.where(a == 20))   # (array([1, 3]),)
print(np.argmax(a))        # 4
print(np.argmin(a))        # 0

9. Sorting Arrays

FunctionPurpose
np.sort()Returns sorted copy
array.sort()Sorts in-place
np.argsort()Indices that would sort
a = np.array([40, 10, 30, 20])
print(np.sort(a))       # [10 20 30 40]
print(np.argsort(a))    # [1 3 2 0]

10. Filtering Arrays

Boolean masks select matching elements.

a = np.array([10, 25, 30, 45, 50])
print(a[a > 30])              # [45 50]
print(a[(a > 20) & (a < 50)]) # [25 30 45]
print(a[a % 2 == 0])          # [10 30 50]

11. Random Number Generation

FunctionPurpose
np.random.rand()Random floats 0 to 1
np.random.randint()Random integers
np.random.choice()Random selection
np.random.seed()Reproducibility
np.random.seed(10)
print(np.random.randint(1, 100, 5))   # Same every time
print(np.random.rand(3))              # 3 random floats
print(np.random.choice([1, 2, 3, 4])) # Random selection

12. Summary

NumPy provides:

  • Fast array operations
  • Flexible indexing and slicing
  • Powerful joining and splitting
  • Efficient searching, sorting, and filtering
  • Reproducible random number generation
  • Foundation for Pandas, SciPy, and machine learning

Q22. Explain data visualization using Matplotlib. Compare line plot, bar chart, histogram, scatter plot, and pie chart with suitable examples and use cases.

Answer:

1. Introduction to Matplotlib

Matplotlib is a Python library used for creating visualizations. It allows programmers to create line graphs, bar charts, histograms, scatter plots, pie charts, and many other types of plots.

The most commonly used Matplotlib module is pyplot, usually imported as plt.

import matplotlib.pyplot as plt

2. Why Visualize Data?

Visualization helps convert numbers into patterns. A good chart can reveal:

  • Trends — how values change over time
  • Comparisons — differences between categories
  • Distributions — how data is spread
  • Relationships — connections between variables
  • Outliers — unusual data points

3. Matplotlib Visualization Workflow

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Prepare   │ →  │ 2. Create    │ →  │ 3. Plot Data │ →  │ 4. Customize │ →  │ 5. Show or   │
│    Data      │    │   Figure     │    │              │    │    Plot      │    │    Save      │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘

4. Comparison of Chart Types

Chart TypeUsed ForExample Question
Line plotShowing trends over timeHow did sales change month by month?
Bar chartComparing categoriesWhich department scored highest?
HistogramShowing distribution of numerical valuesWhat is the distribution of exam marks?
Scatter plotShowing relationship between two numerical variablesIs study time related to marks?
Pie chartShowing parts of a wholeWhat percentage of students selected each elective?

5. Line Plot

Purpose: Show trends over time.

import matplotlib.pyplot as plt

x = [1, 2, 3, 4, 5]
y = [10, 15, 13, 18, 20]

plt.plot(x, y, marker='o', color='blue')
plt.title("Simple Line Plot")
plt.xlabel("X Values")
plt.ylabel("Y Values")
plt.grid(True, alpha=0.3)
plt.show()

Use Cases:

  • Stock prices over time
  • Monthly sales trends
  • Temperature changes
  • Website traffic

6. Bar Chart

Purpose: Compare categories.

subjects = ["Python", "Maths", "DBMS", "OS"]
marks = [85, 78, 92, 74]

plt.bar(subjects, marks, color='steelblue')
plt.title("Marks by Subject")
plt.xlabel("Subject")
plt.ylabel("Marks")
plt.show()

Use Cases:

  • Sales by product
  • Population by country
  • Survey responses
  • Department comparisons

7. Histogram

Purpose: Show distribution of numerical values.

marks = [45, 56, 67, 78, 89, 90, 72, 63, 55, 81]

plt.hist(marks, bins=5, color='green', edgecolor='black')
plt.title("Distribution of Marks")
plt.xlabel("Marks Range")
plt.ylabel("Number of Students")
plt.show()

Use Cases:

  • Age distribution
  • Income distribution
  • Test score distribution
  • Height distribution

Difference from Bar Chart: Histogram bars touch (continuous data); bar chart bars are separate (categorical data).

8. Scatter Plot

Purpose: Show relationship between two numerical variables.

hours = [1, 2, 3, 4, 5, 6]
marks = [40, 50, 55, 65, 75, 85]

plt.scatter(hours, marks, color='red')
plt.title("Study Hours vs Marks")
plt.xlabel("Study Hours")
plt.ylabel("Marks")
plt.show()

Use Cases:

  • Height vs weight
  • Advertising spend vs sales
  • Study time vs marks
  • Temperature vs ice cream sales

9. Pie Chart

Purpose: Show parts of a whole.

labels = ["Python", "Java", "C++", "R"]
students = [40, 25, 20, 15]

plt.pie(students, labels=labels, autopct="%1.1f%%")
plt.title("Programming Language Preference")
plt.show()

Use Cases:

  • Market share
  • Budget allocation
  • Survey responses
  • Population by category

Limitation: Hard to compare many slices; best for 2–5 categories.

10. Common pyplot Functions

FunctionPurpose
plot()Line plot
bar()Bar chart
hist()Histogram
scatter()Scatter plot
pie()Pie chart
title()Chart title
xlabel(), ylabel()Axis labels
legend()Chart legend
show()Display figure
savefig()Save as image

11. Visualization Rule

Every chart should have:

  • A clear title
  • Axis labels where applicable
  • Readable category names
  • A clear purpose

Avoid adding unnecessary decoration.

12. Summary

ChartBest ForAvoid When
LineTrends over timeCategorical data
BarComparing categoriesContinuous data
HistogramDistributionsCategorical data
ScatterRelationshipsToo many points
PieParts of a wholeMore than 5 categories

Q23. Describe a complete mini-project in which a CSV file is loaded using Pandas, numerical operations are performed using NumPy, and results are visualized using Matplotlib.

Answer:

1. Project Title: Student Performance Analysis

Objective: Analyze student performance data from a CSV file, compute statistics, and visualize results.

2. Project Workflow

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Load CSV  │ →  │ 2. Clean &   │ →  │ 3. Compute   │ →  │ 4. Visualize │
│  (Pandas)    │    │  Transform   │    │  (NumPy)     │    │  (Matplotlib)│
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘

3. Sample Data (students.csv)

Name,Department,Internal,External,Attendance
Amit,CSE,25,60,85
Ravi,CSE,22,55,78
Neha,ECE,28,65,92
Sara,ECE,20,50,70
John,MECH,24,58,88
Priya,CSE,27,62,95
Arjun,MECH,18,45,65
Meena,ECE,26,60,80

4. Complete Program

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# ─────────────────────────────────────────────
# Step 1: Load Data (Pandas)
# ─────────────────────────────────────────────
df = pd.read_csv("students.csv")

print("=== First 5 rows ===")
print(df.head())

print("\n=== Data Info ===")
print(df.info())

print("\n=== Statistical Summary ===")
print(df.describe())

# ─────────────────────────────────────────────
# Step 2: Clean and Transform Data (Pandas)
# ─────────────────────────────────────────────
# Check for missing values
print("\n=== Missing Values ===")
print(df.isnull().sum())

# Remove duplicates
df = df.drop_duplicates()

# Create new columns
df["Total"] = df["Internal"] + df["External"]
df["Average"] = np.round(df["Total"] / 2, 2)

# Grade assignment
def assign_grade(avg):
    if avg >= 85:
        return "A"
    elif avg >= 70:
        return "B"
    elif avg >= 55:
        return "C"
    else:
        return "D"

df["Grade"] = df["Average"].apply(assign_grade)

print("\n=== Transformed Data ===")
print(df)

# ─────────────────────────────────────────────
# Step 3: Numerical Operations (NumPy)
# ─────────────────────────────────────────────
mean_avg = np.mean(df["Average"])
median_avg = np.median(df["Average"])
std_avg = np.std(df["Average"])
max_avg = np.max(df["Average"])
min_avg = np.min(df["Average"])

print("\n=== Statistical Measures (NumPy) ===")
print(f"Mean Average:{mean_avg:.2f}")
print(f"Median Average:{median_avg:.2f}")
print(f"Std Deviation:{std_avg:.2f}")
print(f"Max Average:{max_avg:.2f}")
print(f"Min Average:{min_avg:.2f}")

# Filter top performers
top = df[df["Average"] > 80]
print(f"\nTop performers ({len(top)}):")
print(top[["Name", "Average", "Grade"]])

# Filter low attendance
low_att = df[df["Attendance"] < 75]
print(f"\nLow attendance ({len(low_att)}):")
print(low_att[["Name", "Attendance"]])

# ─────────────────────────────────────────────
# Step 4: Grouping and Aggregation (Pandas)
# ─────────────────────────────────────────────
dept_avg = df.groupby("Department")["Average"].mean().reset_index()
print("\n=== Department-wise Average ===")
print(dept_avg)

# ─────────────────────────────────────────────
# Step 5: Visualization (Matplotlib)
# ─────────────────────────────────────────────
fig, axes = plt.subplots(2, 2, figsize=(14, 10))

# Chart 1: Bar chart — Average marks by student
axes[0, 0].bar(df["Name"], df["Average"], color="steelblue")
axes[0, 0].axhline(y=mean_avg, color="red", linestyle="--", label=f"Mean ={mean_avg:.1f}")
axes[0, 0].set_title("Average Marks by Student")
axes[0, 0].set_xlabel("Student")
axes[0, 0].set_ylabel("Average Marks")
axes[0, 0].legend()
axes[0, 0].tick_params(axis="x", rotation=45)

# Chart 2: Histogram — Distribution of averages
axes[0, 1].hist(df["Average"], bins=5, color="green", edgecolor="black")
axes[0, 1].set_title("Distribution of Average Marks")
axes[0, 1].set_xlabel("Average Marks")
axes[0, 1].set_ylabel("Number of Students")

# Chart 3: Scatter plot — Attendance vs Average
axes[1, 0].scatter(df["Attendance"], df["Average"], color="red", s=80)
axes[1, 0].set_title("Attendance vs Average Marks")
axes[1, 0].set_xlabel("Attendance (%)")
axes[1, 0].set_ylabel("Average Marks")
axes[1, 0].grid(True, alpha=0.3)

# Chart 4: Bar chart — Department-wise average
axes[1, 1].bar(dept_avg["Department"], dept_avg["Average"], color="purple")
axes[1, 1].set_title("Department-wise Average Marks")
axes[1, 1].set_xlabel("Department")
axes[1, 1].set_ylabel("Average Marks")

plt.tight_layout()
plt.savefig("student_performance.png")
plt.show()

# ─────────────────────────────────────────────
# Step 6: Export Results (Pandas)
# ─────────────────────────────────────────────
df.to_csv("analyzed_students.csv", index=False)
dept_avg.to_csv("department_averages.csv", index=False)

print("\n=== Analysis Complete ===")
print("Files saved: student_performance.png, analyzed_students.csv, department_averages.csv")

5. Sample Output

=== First 5 rows ===
    Name Department  Internal  External  Attendance
0   Amit        CSE        25        60          85
1   Ravi        CSE        22        55          78
2   Neha        ECE        28        65          92
3   Sara        ECE        20        50          70
4   John       MECH        24        58          88

=== Statistical Measures (NumPy) ===
Mean Average:   71.75
Median Average: 72.50
Std Deviation:  11.36
Max Average:    89.00
Min Average:    52.50

=== Department-wise Average ===
  Department  Average
0        CSE    75.00
1        ECE    72.75
2       MECH    66.25

6. Role of Each Library

LibraryRole in the Project
PandasLoads CSV, cleans data, creates columns, groups, exports
NumPyComputes mean, median, std, max, min; rounds values
MatplotlibCreates bar charts, histogram, scatter plot; saves image

7. Key Insights from the Analysis

  1. CSE has the highest average marks (75.0).
  2. MECH has the lowest average (66.25).
  3. Students with attendance > 85% tend to score above 80.
  4. The distribution of averages is roughly symmetric.
  5. Top performers: Neha (89.0), Amit (85.0), Priya (85.0).

8. Extensions

  • Add correlation analysis between attendance and marks.
  • Include more visualizations (pie chart for grade distribution).
  • Build an interactive dashboard using Plotly.
  • Apply machine learning to predict student performance.

9. Conclusion

This mini-project demonstrates the complete data analysis pipeline:

  • Load data with Pandas
  • Clean and transform with Pandas
  • Compute statistics with NumPy
  • Visualize with Matplotlib
  • Export results with Pandas

Q24. Web scraping is useful but must be used responsibly. Explain the workflow of web scraping, tools used, possible applications, limitations, and ethical/legal precautions.

Answer:

1. What is Web Scraping?

Web scraping is the process of extracting data from websites using programs. It is useful when information is available on web pages but not provided as a downloadable dataset or API.

2. Web Scraping Workflow

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│ 1. Send HTTP │ →  │ 2. Download  │ →  │ 3. Parse and │ →  │ 4. Clean and │ →  │ 5. Store or  │
│   Request    │    │    HTML      │    │ Extract Data │    │  Structure   │    │  Export Data │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘
StepMeaningCommon Tool
1. Send HTTP RequestSend a request to the target websiterequests
2. Download HTMLRetrieve the HTML content of the pageresponse.text
3. Parse and Extract DataParse the HTML and extract the required dataBeautifulSoup
4. Clean and Structure DataClean the extracted data and structure it for usePandas, manual cleaning
5. Store or Export DataSave the data to a file, database, or export it for analysisCSV, JSON, Pandas DataFrame

3. Tools Used

ToolPurposeExample
requestsSend HTTP requestsrequests.get(url)
BeautifulSoupParse HTML/XMLBeautifulSoup(html, "html.parser")
lxmlFaster HTML/XML parserBeautifulSoup(html, "lxml")
SeleniumBrowser automation for dynamic pagesdriver.get(url)
ScrapyFull-featured scraping frameworkscrapy crawl spider
PandasStore and analyze scraped datapd.DataFrame(data)

4. Simple Web Scraping Example

import requests
from bs4 import BeautifulSoup

url = "https://example.com"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")

# Extract title
print(soup.title.text)

# Extract all links
for link in soup.find_all("a"):
    print(link.get("href"))

# Extract all paragraphs
for p in soup.find_all("p"):
    print(p.text)

5. Possible Applications

ApplicationDescription
Price comparisonTrack product prices across e-commerce sites
News aggregationCollect headlines from multiple news sources
Sentiment analysisGather social media posts for analysis
ResearchCollect data for academic studies
Lead generationGather business contact information
Real estateTrack property listings and prices
Job market analysisCollect job postings and salary data
Sports statisticsGather player and team performance data

6. Limitations

LimitationExplanation
Dynamic contentJavaScript-rendered content not available in HTML
Anti-scraping measuresCAPTCHAs, IP blocking, rate limiting
Changing structureWebsite layout changes break scrapers
Legal restrictionsTerms of service may prohibit scraping
Data qualityScraped data may be incomplete or noisy
MaintenanceScrapers require regular updates
PerformanceLarge-scale scraping is resource-intensive
PrecautionExplanation
Terms of ServiceCheck if the website allows automated access
robots.txtFollow the rules specified in the site’s robots.txt file
CopyrightDo not reproduce copyrighted content without permission
PrivacyDo not collect personal or restricted data without consent
Server LoadAvoid sending too many requests that could overload the server
Official APIsPrefer official APIs whenever available
Rate LimitingAdd delays between requests (time.sleep(2))
User-AgentIdentify your bot transparently
Data UsageUse scraped data only for intended purposes
AttributionCredit the source where appropriate

8. Responsible Scraping Example

import requests
from bs4 import BeautifulSoup
import time

headers = {"User-Agent": "Mozilla/5.0 (educational scraper)"}

for page in range(1, 10):
    url = f"https://example.com/page{page}"
    try:
        response = requests.get(url, headers=headers, timeout=10)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, "html.parser")
        print(f"Page{page}:{soup.title.text}")
        time.sleep(2)   # Respectful delay
    except requests.RequestException as e:
        print(f"Error on page{page}:{e}")

9. Alternatives to Web Scraping

AlternativeDescription
Official APIsStructured data provided by the website
Public datasetsPre-collected data from Kaggle, UCI, etc.
RSS feedsStandardized content feeds
Data dumpsBulk data downloads offered by platforms
PartnershipsDirect data-sharing agreements

10. Summary

Web scraping is a powerful technique for collecting data from websites, but it must be used responsibly. Key points:

  • Workflow: Request → Download → Parse → Clean → Store
  • Tools: requests, BeautifulSoup, Selenium, Scrapy
  • Applications: Price comparison, news, research, sentiment analysis
  • Limitations: Dynamic content, anti-scraping, legal restrictions
  • Ethics: Respect robots.txt, terms of service, privacy, server load
  • Best Practice: Prefer official APIs whenever available

PART D — ANALYTICAL / CASE-BASED QUESTIONS


Q25 (Case Study). A college has student performance data stored in a CSV file with columns such as Name, Department, Internal Marks, External Marks, Attendance, and Result. The management wants a Python program to read the data, calculate total marks, find average marks by department, identify students with low attendance, and visualize department-wise performance.

Questions:

  1. Which Pandas functions will be useful for reading and inspecting the data?
  2. How can new columns such as Total and Average be created?
  3. How can students with attendance below 75% be filtered?
  4. Which Matplotlib charts would be suitable for showing department-wise averages and result distribution?
  5. Where can NumPy be used in this problem?

Answer:

1. Pandas Functions for Reading and Inspecting Data

FunctionPurpose
pd.read_csv()Read the CSV file into a DataFrame
df.head()Display first 5 rows
df.tail()Display last 5 rows
df.shapeShow (rows, columns)
df.columnsList column names
df.info()Show data types and non-null counts
df.describe()Statistical summary of numerical columns
df.isnull().sum()Count missing values per column
df.dtypesData types of columns

Example:

import pandas as pd

df = pd.read_csv("student_performance.csv")
print(df.head())
print(df.shape)
print(df.info())
print(df.describe())
print(df.isnull().sum())

2. Creating New Columns: Total and Average

# Create Total column
df["Total"] = df["Internal Marks"] + df["External Marks"]

# Create Average column
df["Average"] = df["Total"] / 2

# Alternative using NumPy for rounding
import numpy as np
df["Average"] = np.round(df["Total"] / 2, 2)

Explanation:

  • Pandas allows column creation by assignment: df["NewCol"] = values.
  • Operations are vectorized (applied to all rows at once).
  • NumPy can be used for rounding: np.round().

3. Filtering Students with Attendance Below 75%

low_attendance = df[df["Attendance"] < 75]
print(low_attendance[["Name", "Department", "Attendance"]])

# Count
print(f"Students with low attendance:{len(low_attendance)}")

# Save to file
low_attendance.to_csv("low_attendance.csv", index=False)

Explanation:

  • Boolean filtering: df[df["Attendance"] < 75] returns rows where the condition is True.
  • Multiple conditions: df[(df["Attendance"] < 75) & (df["Result"] == "Fail")].

4. Suitable Matplotlib Charts

ObjectiveChart TypeJustification
Department-wise averagesBar chartCompares categories clearly
Result distributionPie chart or bar chartShows parts of a whole (Pass/Fail)
Attendance vs AverageScatter plotShows relationship
Distribution of marksHistogramShows frequency distribution

Example — Department-wise Average (Bar Chart):

import matplotlib.pyplot as plt

dept_avg = df.groupby("Department")["Average"].mean().reset_index()

plt.bar(dept_avg["Department"], dept_avg["Average"], color="steelblue")
plt.title("Department-wise Average Marks")
plt.xlabel("Department")
plt.ylabel("Average Marks")
plt.show()

Example — Result Distribution (Pie Chart):

result_counts = df["Result"].value_counts()

plt.pie(result_counts, labels=result_counts.index, autopct="%1.1f%%")
plt.title("Result Distribution")
plt.show()

5. Where NumPy Can Be Used

TaskNumPy Function
Rounding averagesnp.round(df["Average"], 2)
Computing meannp.mean(df["Average"])
Computing mediannp.median(df["Average"])
Computing standard deviationnp.std(df["Average"])
Finding max/minnp.max(), np.min()
Correlationnp.corrcoef(df["Attendance"], df["Average"])
Filtering with conditionsdf["Average"].values > 80

Example:

import numpy as np

mean_avg = np.mean(df["Average"])
std_avg = np.std(df["Average"])
max_avg = np.max(df["Average"])
correlation = np.corrcoef(df["Attendance"], df["Average"])[0, 1]

print(f"Mean:{mean_avg:.2f}")
print(f"Std:{std_avg:.2f}")
print(f"Max:{max_avg:.2f}")
print(f"Correlation (Attendance vs Average):{correlation:.2f}")

6. Complete Program

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# Step 1: Read data
df = pd.read_csv("student_performance.csv")
print(df.head())
print(df.info())

# Step 2: Calculate Total and Average
df["Total"] = df["Internal Marks"] + df["External Marks"]
df["Average"] = np.round(df["Total"] / 2, 2)

# Step 3: Find average by department
dept_avg = df.groupby("Department")["Average"].mean().reset_index()
print(dept_avg)

# Step 4: Identify low attendance students
low_att = df[df["Attendance"] < 75]
print(f"Low attendance:{len(low_att)} students")

# Step 5: Visualize
fig, axes = plt.subplots(1, 2, figsize=(14, 5))

# Department-wise average
axes[0].bar(dept_avg["Department"], dept_avg["Average"], color="steelblue")
axes[0].set_title("Department-wise Average Marks")
axes[0].set_xlabel("Department")
axes[0].set_ylabel("Average Marks")

# Result distribution
result_counts = df["Result"].value_counts()
axes[1].pie(result_counts, labels=result_counts.index, autopct="%1.1f%%")
axes[1].set_title("Result Distribution")

plt.tight_layout()
plt.savefig("college_analysis.png")
plt.show()

# Step 6: Export
dept_avg.to_csv("department_averages.csv", index=False)
low_att.to_csv("low_attendance.csv", index=False)

7. Key Insights

  1. Department-wise averages identify strong and weak departments.
  2. Low attendance students may need intervention.
  3. Result distribution shows overall pass/fail rates.
  4. Correlation between attendance and marks indicates if attendance affects performance.

Q26 (Compare and Analyse). Construct a table comparing Pandas, NumPy, and Matplotlib using at least eight dimensions such as purpose, data structure, common functions, input data type, output type, speed, use case, and example code. After the table, explain how the three libraries work together in a real data analysis project.

Answer:

1. Comparison Table

DimensionPandasNumPyMatplotlib
PurposeTabular data analysis and manipulationNumerical computation and array operationsData visualization
Data StructureSeries (1D), DataFrame (2D)ndarray (N-dimensional)Figure, Axes
Common Functionsread_csv(), head(), merge(), groupby(), fillna()array(), zeros(), sort(), where(), concatenate()plot(), bar(), hist(), scatter(), pie()
Input Data TypeCSV, JSON, Excel, SQL, dict, listLists, tuples, arrays, scalarsLists, arrays, Series, DataFrame columns
Output TypeDataFrame, Seriesndarray, scalarCharts, plots, figures
SpeedModerate (built on NumPy)Very fast (C-based)Moderate (rendering overhead)
Memory EfficiencyModerateHigh (contiguous storage)N/A (visualization)
Use CaseData cleaning, analysis, aggregationMathematical operations, linear algebraVisual communication, EDA
Learning CurveModerateLow to moderateLow
IntegrationBuilt on NumPy; integrates with MatplotlibFoundation for Pandas, SciPy, MLWorks with Pandas and NumPy
Example Codedf.groupby("Dept")["Marks"].mean()np.mean(arr)plt.bar(x, y)
Handles Missing DataYes (dropna(), fillna())No (uses NaN)N/A
Labelled DataYes (row/column labels)No (position-based)N/A
Typical UsersData analysts, data scientistsScientists, engineersAnalysts, communicators

2. How the Three Libraries Work Together

In a real data analysis project, the three libraries form a complete pipeline:

┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│   PANDAS     │ →  │    NUMPY     │ →  │  MATPLOTLIB  │
│ Load & Clean │    │  Compute &   │    │  Visualize & │
│    Data      │    │   Analyze    │    │  Communicate │
└──────────────┘    └──────────────┘    └──────────────┘

3. Real-World Data Analysis Project

Scenario: Analyze sales data from a retail company.

Step 1: Load Data (Pandas)

import pandas as pd

df = pd.read_csv("sales.csv")
print(df.head())
print(df.info())

Step 2: Clean Data (Pandas)

df = df.dropna()
df = df.drop_duplicates()
df["Total"] = df["Quantity"] * df["Price"]

Step 3: Compute Statistics (NumPy)

import numpy as np

mean_sales = np.mean(df["Total"])
std_sales = np.std(df["Total"])
max_sales = np.max(df["Total"])

Step 4: Aggregate (Pandas)

monthly = df.groupby("Month")["Total"].sum().reset_index()
product_avg = df.groupby("Product")["Total"].mean().reset_index()

Step 5: Visualize (Matplotlib)

import matplotlib.pyplot as plt

fig, axes = plt.subplots(1, 2, figsize=(14, 5))

# Monthly sales trend
axes[0].plot(monthly["Month"], monthly["Total"], marker="o")
axes[0].set_title("Monthly Sales Trend")
axes[0].set_xlabel("Month")
axes[0].set_ylabel("Total Sales")

# Product-wise average
axes[1].bar(product_avg["Product"], product_avg["Total"], color="steelblue")
axes[1].set_title("Product-wise Average Sales")
axes[1].set_xlabel("Product")
axes[1].set_ylabel("Average Sales")

plt.tight_layout()
plt.savefig("sales_analysis.png")
plt.show()

Step 6: Export Results (Pandas)

monthly.to_csv("monthly_sales.csv", index=False)
product_avg.to_csv("product_averages.csv", index=False)

4. Role of Each Library in the Project

StageLibraryRole
Data LoadingPandasRead CSV/JSON/Excel into DataFrames
Data CleaningPandasHandle missing values, duplicates
Feature CreationPandasCreate new columns (Total, Average)
Numerical OperationsNumPyCompute mean, std, max, min
AggregationPandasGroup by month, product, department
VisualizationMatplotlibCreate line, bar, pie, scatter charts
ExportPandasWrite results to CSV/JSON

5. Benefits of Using All Three Together

BenefitExplanation
Complete PipelineFrom raw data to actionable insights
EfficiencyNumPy provides fast computation; Pandas provides structured handling
ClarityMatplotlib communicates results visually
FlexibilityEach library handles its strengths
IntegrationPandas is built on NumPy; Matplotlib accepts Pandas/NumPy data
ReproducibilityCode can be rerun on new data
ScalabilityHandles small to large datasets

6. Summary

LibraryStrengthsWeaknesses
PandasTabular data, labelled indexing, I/OSlower than NumPy for pure numerics
NumPyFast numerics, vectorizationNo labels, no missing value handling
MatplotlibFlexible visualizationVerbose syntax, static by default

Key Takeaway: These three libraries are complementary. Pandas handles structured data, NumPy handles numerical computation, and Matplotlib handles visualization. Together, they form the foundation of Python data analysis.


Summary of All Answers

PartQ#TopicKey Answer
A1Tabular data library(b) Pandas
A2Series definition(b) One-dimensional labelled array
A3Read CSV(a) pd.read_csv()
A4NumPy array object(c) ndarray
A5Sort array(b) np.sort()
A6Display chart(c) plt.show()
A7Trend over time(b) Line plot
A8HTML parsing(a) BeautifulSoup
B9Pandas advantagesData loading, cleaning, selection, transformation, aggregation
B10Series vs DataFrame1D vs 2D; single column vs full table
B11loc[] vs iloc[]Label-based vs position-based
B12Missing valuesdropna() removes; fillna() replaces
B13CSV and JSON I/Oread_csv(), to_csv(), read_json(), to_json()
B14ndarray attributesshape, ndim, size, dtype
B15Indexing and slicing1D and 2D examples
B16Joining and splittingconcatenate(), vstack(), array_split(), reshape()
B17Search, sort, filterwhere(), sort(), boolean masks
B18Web scrapingWorkflow, tools, ethics
B19Visualization needTrends, comparisons, distributions, relationships
C20Pandas workflowLoad → Inspect → Clean → Transform → Combine → Export
C21NumPy in detailCreation, attributes, indexing, operations, random
C22Matplotlib chartsLine, bar, histogram, scatter, pie
C23Mini-projectComplete pipeline with all three libraries
C24Web scrapingWorkflow, tools, applications, limitations, ethics
D25College case studyPandas + NumPy + Matplotlib solution
D26Comparison table14-dimension comparison + integration

On this page

PART A — MULTIPLE CHOICE QUESTIONS (1 Mark Each)PART B — SHORT ANSWER QUESTIONS (5 Marks Each)Q9. Define Pandas. Explain any five advantages of using Pandas for data analysis.Q10. Differentiate between Pandas Series and DataFrame with suitable examples.Q11. Explain how to select rows and columns from a DataFrame using loc[] and iloc[].Q12. Write short notes on handling missing values in Pandas using dropna() and fillna().1. dropna() — Remove Missing Values2. fillna() — Replace Missing Values3. Other Missing Value FunctionsQ13. Explain how CSV and JSON files can be read and written using Pandas.1. Reading CSV Files2. Writing CSV Files3. Reading JSON Files4. Writing JSON Files5. File Reading and Writing Functions SummaryQ14. Define NumPy ndarray. Explain shape, ndim, size, and dtype attributes.1. shape2. ndim3. size4. dtypeQ15. Explain indexing and slicing in one-dimensional and two-dimensional NumPy arrays.1. One-Dimensional Arrays2. Two-Dimensional Arrays3. Key PointsQ16. Discuss joining and splitting operations in NumPy with examples.1. Joining Arrays(a) concatenate() — Join Along an Existing Axis(b) vstack() — Vertical Stacking(c) hstack() — Horizontal Stacking2. Splitting Arrays(a) array_split() — Split into Multiple Parts(b) split() — Equal Division(c) hsplit() and vsplit() — Split 2D Arrays3. Reshapingreshape() — Change Shape Without Changing Data4. Summary Table5. Practical ExampleQ17. Explain searching, sorting, and filtering arrays in NumPy.1. Searching Arraysnp.where() — Find Indices Where Condition is Truenp.searchsorted() — Find Insertion Positionnp.argmax() and np.argmin() — Index of Max/Min2. Sorting Arraysnp.sort() — Returns Sorted Copyarray.sort() — Sorts In-PlaceSorting 2D Arraysnp.argsort() — Indices That Would Sort the Array3. Filtering Arrays4. Summary Table5. Complete ExampleQ18. What is web scraping? Explain the basic steps and ethical precautions involved.1. Basic Web Scraping Workflow2. Steps in a Basic Web Scraping Workflow3. Simple Web Scraping Example4. Common Tools5. Ethical Precautions6. Applications of Web ScrapingQ19. Explain the need for data visualization and list common Matplotlib chart types.Matplotlib Visualization WorkflowPART C — LONG ANSWER / ESSAY QUESTIONS (10 Marks Each)Q20. Explain the complete Pandas workflow for loading, inspecting, cleaning, transforming, combining, and exporting tabular data. Support your answer with suitable code examples.1. Introduction to Pandas2. Pandas Workflow Overview3. Step 1: Import Data4. Step 2: Create and Inspect DataFrame5. Step 3: Clean and Transform Data6. Step 4: Combine DataFrames7. Step 5: Analyze and Export8. Complete Integrated Example9. Benefits of the Pandas WorkflowQ21. Discuss NumPy arrays in detail. Explain array creation, attributes, indexing, slicing, vectorized operations, joining, splitting, searching, sorting, filtering, and random number generation.1. Introduction to NumPy2. Array Creation3. ndarray Attributes4. Indexing and Slicing5. Vectorized Operations6. Joining Arrays7. Splitting Arrays8. Searching Arrays9. Sorting Arrays10. Filtering Arrays11. Random Number Generation12. SummaryQ22. Explain data visualization using Matplotlib. Compare line plot, bar chart, histogram, scatter plot, and pie chart with suitable examples and use cases.1. Introduction to Matplotlib2. Why Visualize Data?3. Matplotlib Visualization Workflow4. Comparison of Chart Types5. Line Plot6. Bar Chart7. Histogram8. Scatter Plot9. Pie Chart10. Common pyplot Functions11. Visualization Rule12. SummaryQ23. Describe a complete mini-project in which a CSV file is loaded using Pandas, numerical operations are performed using NumPy, and results are visualized using Matplotlib.1. Project Title: Student Performance Analysis2. Project Workflow3. Sample Data (students.csv)4. Complete Program5. Sample Output6. Role of Each Library7. Key Insights from the Analysis8. Extensions9. ConclusionQ24. Web scraping is useful but must be used responsibly. Explain the workflow of web scraping, tools used, possible applications, limitations, and ethical/legal precautions.1. What is Web Scraping?2. Web Scraping Workflow3. Tools Used4. Simple Web Scraping Example5. Possible Applications6. Limitations7. Ethical and Legal Precautions8. Responsible Scraping Example9. Alternatives to Web Scraping10. SummaryPART D — ANALYTICAL / CASE-BASED QUESTIONSQ25 (Case Study). A college has student performance data stored in a CSV file with columns such as Name, Department, Internal Marks, External Marks, Attendance, and Result. The management wants a Python program to read the data, calculate total marks, find average marks by department, identify students with low attendance, and visualize department-wise performance.1. Pandas Functions for Reading and Inspecting Data2. Creating New Columns: Total and Average3. Filtering Students with Attendance Below 75%4. Suitable Matplotlib Charts5. Where NumPy Can Be Used6. Complete Program7. Key InsightsQ26 (Compare and Analyse). Construct a table comparing Pandas, NumPy, and Matplotlib using at least eight dimensions such as purpose, data structure, common functions, input data type, output type, speed, use case, and example code. After the table, explain how the three libraries work together in a real data analysis project.1. Comparison Table2. How the Three Libraries Work Together3. Real-World Data Analysis Project4. Role of Each Library in the Project5. Benefits of Using All Three Together6. SummarySummary of All Answers