Adv-Python Unit 4: Question with Answers
Unit IV: Pandas, NumPy, and Matplotlib -> Generated and Prepared By Thiruselvan (ThiruXD)
SECTION A: MULTIPLE CHOICE QUESTIONS (50 MCQs)
Introduction to Pandas
Q1. Which library is mainly used for tabular data analysis in Python?
- NumPy
- Pandas
- Matplotlib
- Requests
Answer: B) Pandas
Explanation: Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures (Series and DataFrame) for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs.
Q2. The name “Pandas” is derived from:
- Personal Data Analysis
- Panel Data
- Pandas Animal
- Parallel Data
Answer: B) Panel Data
Explanation: The name Pandas is derived from Panel Data, which refers to multidimensional structured datasets.
Q3. Which of the following is NOT a feature of Pandas?
- Data loading
- Data cleaning
- Image processing
- Data aggregation
Answer: C) Image processing
Explanation: Pandas features include data loading, cleaning, selection, transformation, aggregation, and export. Image processing is not a Pandas feature.
Q4. What is the standard import convention for Pandas?
import pandasimport pandas as pdfrom pandas import pdimport pd as pandas
Answer: B) import pandas as pd
Explanation: The standard import convention is import pandas as pd, which creates the alias pd for convenience.
Q5. Which function is used to read a CSV file into a DataFrame?
pd.read_csv()pd.open_csv()np.read_csv()plt.csv()
Answer: A) pd.read_csv()
Explanation: pd.read_csv() is the Pandas function for reading CSV files into a DataFrame. The other options do not exist or belong to different libraries.
Pandas Series and DataFrames
Q6. A Pandas Series is best described as:
- A two-dimensional table
- A one-dimensional labelled array
- A plotting function
- A web scraping tool
Answer: B) A one-dimensional labelled array
Explanation: A Series is a one-dimensional labelled array in Pandas. It can store integers, floats, strings, or other Python objects, similar to a single column in a spreadsheet.
Q7. A Pandas DataFrame is:
- A one-dimensional array
- A two-dimensional labelled data structure
- A plotting function
- A JSON parser
Answer: B) A two-dimensional labelled data structure
Explanation: A DataFrame is a two-dimensional labelled data structure with rows and columns. It is the most commonly used data structure in Pandas.
Q8. Which method displays the first five rows of a DataFrame?
df.first()df.head()df.top()df.start()
Answer: B) df.head()
Explanation: df.head() displays the first five rows of a DataFrame by default. You can specify a different number as an argument.
Q9. Which attribute shows the number of rows and columns in a DataFrame?
df.sizedf.dimdf.shapedf.length
Answer: C) df.shape
Explanation: df.shape returns a tuple (rows, columns) representing the dimensions of the DataFrame.
Q10. Which method gives a statistical summary of numerical columns?
df.info()df.summary()df.describe()df.stats()
Answer: C) df.describe()
Explanation: df.describe() displays a statistical summary (count, mean, std, min, quartiles, max) of numerical columns.
Data Manipulation and Cleaning
Q11. Which method selects rows by label in Pandas?
df.iloc[]df.loc[]df.sel[]df.pick[]
Answer: B) df.loc[]
Explanation: df.loc[] selects rows and columns by label. df.iloc[] selects by integer position.
Q12. Which method selects rows by integer position?
df.loc[]df.iloc[]df.pos[]df.idx[]
Answer: B) df.iloc[]
Explanation: df.iloc[] selects rows and columns by integer position (0-based indexing), similar to Python list slicing.
Q13. Which function detects missing values in a DataFrame?
df.isnull()df.missing()df.nan()df.empty()
Answer: A) df.isnull()
Explanation: df.isnull() returns a DataFrame of Boolean values indicating where missing values (NaN) are present.
Q14. Which function removes rows with missing values?
df.remove()df.dropna()df.clean()df.delete()
Answer: B) df.dropna()
Explanation: df.dropna() removes rows (or columns) containing missing values. Use axis=1 to drop columns.
Q15. Which function replaces missing values with a specified value?
df.fillna(value)df.replace(value)df.insert(value)df.update(value)
Answer: A) df.fillna(value)
Explanation: df.fillna(value) replaces missing values (NaN) with the specified value. You can also use methods like ‘ffill’ or ‘bfill’.
Q16. Which function removes duplicate rows?
df.drop_duplicates()df.remove_duplicates()df.unique()df.distinct()
Answer: A) df.drop_duplicates()
Explanation: df.drop_duplicates() removes duplicate rows from the DataFrame, keeping the first occurrence by default.
Q17. Which operation stacks DataFrames vertically or horizontally?
pd.merge()pd.concat()pd.join()pd.combine()
Answer: B) pd.concat()
Explanation: pd.concat() stacks DataFrames row-wise or column-wise. Use axis=0 for vertical and axis=1 for horizontal stacking.
Q18. Which operation combines DataFrames using a common column?
pd.concat()pd.merge()pd.stack()pd.append()
Answer: B) pd.merge()
Explanation: pd.merge() combines DataFrames using a common column (key), similar to SQL JOIN operations.
Q19. Which operation combines DataFrames using their index?
pd.concat()pd.merge()pd.join()pd.combine()
Answer: C) pd.join()
Explanation: pd.join() combines DataFrames using their index labels. It is a convenience method based on merge().
Q20. Which function is used to create a new column in a DataFrame?
df.add_column()df["NewCol"] = valuesdf.insert_column()df.create()
Answer: B) df["NewCol"] = values
Explanation: You can create a new column by assigning values to a new column name using bracket notation, e.g., df["Total"] = df["A"] + df["B"].
Reading and Writing Files
Q21. Which function reads a JSON file into a DataFrame?
pd.read_json()pd.load_json()pd.open_json()pd.parse_json()
Answer: A) pd.read_json()
Explanation: pd.read_json() reads a JSON file into a DataFrame. It supports various orientations via the orient parameter.
Q22. Which function writes a DataFrame to a CSV file?
df.write_csv()df.to_csv()df.save_csv()df.export_csv()
Answer: B) df.to_csv()
Explanation: df.to_csv() writes a DataFrame to a CSV file. Use index=False to omit the row index.
Q23. Which parameter in to_csv() prevents writing the row index?
header=Falseindex=Falserows=Falselabels=False
Answer: B) index=False
Explanation: index=False prevents Pandas from writing the row index as the first column in the CSV file.
Q24. Which function writes a DataFrame to a JSON file?
df.to_json()df.write_json()df.save_json()df.export_json()
Answer: A) df.to_json()
Explanation: df.to_json() writes a DataFrame to a JSON file. The orient parameter controls the JSON structure.
NumPy Basics
Q25. What does NumPy stand for?
- Numerical Python
- Number Python
- New Python
- Numeric Py
Answer: A) Numerical Python
Explanation: NumPy stands for Numerical Python. It is a fundamental Python library for scientific computing and numerical operations.
Q26. What is the main array object in NumPy called?
- DataFrame
- Series
- ndarray
- pyplot
Answer: C) ndarray
Explanation: The main array object in NumPy is called ndarray (N-dimensional array). It stores elements of the same data type in a compact and efficient way.
Q27. Which NumPy function creates an array from a list?
np.list()np.array()np.create()np.make()
Answer: B) np.array()
Explanation: np.array() creates an array from a list, tuple, or other sequence. Example: np.array([1, 2, 3]).
Q28. Which NumPy function creates an array filled with zeros?
np.empty()np.zeros()np.null()np.nan()
Answer: B) np.zeros()
Explanation: np.zeros(shape) creates an array filled with zeros. Example: np.zeros(5) creates [0. 0. 0. 0. 0.].
Q29. Which NumPy function creates an identity matrix?
np.identity()np.eye()- Both A and B
np.matrix()
Answer: C) Both A and B
Explanation: Both np.eye(n) and np.identity(n) create an n×n identity matrix. np.eye() is more flexible (allows offset diagonals).
Q30. Which attribute returns the dimensions of a NumPy array?
a.dima.shapea.sizea.length
Answer: B) a.shape
Explanation: a.shape returns a tuple representing the dimensions of the array. For a 2D array, it returns (rows, columns).
Q31. Which attribute returns the number of dimensions of a NumPy array?
a.shapea.ndima.dimsa.rank
Answer: B) a.ndim
Explanation: a.ndim returns the number of dimensions (axes) of the array. A 1D array has ndim=1; a 2D array has ndim=2.
Q32. Which attribute returns the total number of elements in a NumPy array?
a.counta.sizea.totala.length
Answer: B) a.size
Explanation: a.size returns the total number of elements in the array. For a (2, 3) array, size = 6.
Q33. Which attribute returns the data type of elements in a NumPy array?
a.typea.dtypea.datatypea.kind
Answer: B) a.dtype
Explanation: a.dtype returns the data type of elements in the array, such as int64, float64, or bool.
NumPy Indexing and Operations
Q34. In NumPy, what does a[1:4] return for a = np.array([10, 20, 30, 40, 50])?
[10, 20, 30][20, 30, 40][20, 30, 40, 50][10, 20, 30, 40]
Answer: B) [20, 30, 40]
Explanation: Slicing a[1:4] returns elements from index 1 to 3 (exclusive of 4): [20, 30, 40].
Q35. For a 2D array matrix[1, 2], what does this access?
- Row 1, all columns
- All rows, column 2
- Element at row 1, column 2
- Row 2, column 1
Answer: C) Element at row 1, column 2
Explanation: matrix[1, 2] accesses the element at row index 1 and column index 2 (0-based indexing).
Q36. Which NumPy function joins arrays along an existing axis?
np.join()np.concatenate()np.merge()np.combine()
Answer: B) np.concatenate()
Explanation: np.concatenate((a, b)) joins arrays along an existing axis. Use axis=0 for vertical and axis=1 for horizontal.
Q37. Which NumPy function stacks arrays vertically?
np.hstack()np.vstack()np.dstack()np.stack()
Answer: B) np.vstack()
Explanation: np.vstack((x, y)) stacks arrays vertically (row-wise). np.hstack() stacks horizontally.
Q38. Which NumPy function splits an array into multiple parts?
np.divide()np.array_split()np.partition()np.separate()
Answer: B) np.array_split()
Explanation: np.array_split(a, 3) splits an array into 3 parts. It allows unequal division, unlike np.split().
Q39. Which NumPy function returns indices where a condition is true?
np.find()np.where()np.search()np.locate()
Answer: B) np.where()
Explanation: np.where(condition) returns the indices where the condition is true. Example: np.where(a == 20).
Q40. Which NumPy function sorts an array?
np.order()np.sort()np.filter()np.arange()
Answer: B) np.sort()
Explanation: np.sort(a) returns a sorted copy of the array. Use a.sort() to sort in-place.
Q41. Which NumPy operation filters an array?
a.filter(30)a[a > 30]a.select(30)a.where(30)
Answer: B) a[a > 30]
Explanation: Filtering uses boolean masks: a[a > 30] returns elements greater than 30. The condition creates a True/False mask.
Q42. Which NumPy function generates random integers within a range?
np.random.rand()np.random.randint()np.random.choice()np.random.float()
Answer: B) np.random.randint()
Explanation: np.random.randint(low, high, size) generates random integers within a range. Example: np.random.randint(1, 100, 5).
Q43. What is the purpose of np.random.seed()?
- Generate random numbers faster
- Reproduce the same random results
- Stop random generation
- Sort random numbers
Answer: B) Reproduce the same random results
Explanation: np.random.seed(n) sets the seed for the random number generator, ensuring the same sequence of random numbers is produced each time for reproducibility.
Web Scraping
Q44. What is web scraping?
- Creating websites
- Extracting data from websites using programs
- Browsing websites
- Hosting websites
Answer: B) Extracting data from websites using programs
Explanation: Web scraping is the process of extracting data from websites using programs. It is useful when information is available on web pages but not as a downloadable dataset or API.
Q45. Which library is commonly used for parsing HTML in web scraping?
- BeautifulSoup
- NumPy
- Matplotlib
- json
Answer: A) BeautifulSoup
Explanation: BeautifulSoup is a Python library for parsing HTML and XML documents. It is commonly used with requests for web scraping.
Q46. Which library is used to send HTTP requests in Python?
requestsBeautifulSoupNumPyPandas
Answer: A) requests
Explanation: The requests library is used to send HTTP requests (GET, POST, etc.) to web servers. It returns the response content for parsing.
Matplotlib
Q47. Which Matplotlib function is used to display a chart?
plt.display()plt.view()plt.show()plt.open()
Answer: C) plt.show()
Explanation: plt.show() displays the figure on screen. It should be called after creating the plot and customizing it.
Q48. Which Matplotlib function creates a line plot?
plt.line()plt.plot()plt.chart()plt.graph()
Answer: B) plt.plot()
Explanation: plt.plot(x, y) creates a line plot connecting data points. It is the most basic Matplotlib plotting function.
Q49. Which chart is most suitable for showing trend over time?
- Pie chart
- Line plot
- Histogram
- Box plot
Answer: B) Line plot
Explanation: Line plots connect data points chronologically to show trends over time. Pie charts show parts of a whole; histograms show distributions.
Q50. Which Matplotlib function creates a histogram?
plt.bar()plt.hist()plt.pie()plt.scatter()
Answer: B) plt.hist()
Explanation: plt.hist(data, bins=n) creates a histogram showing the frequency distribution of numerical data. plt.bar() creates bar charts.
SECTION B: THEORY QUESTIONS (20)
Q1. Define Pandas. Explain any five advantages of using Pandas for data analysis.
Answer:Pandas is a powerful open-source Python library used for data analysis and manipulation. It provides simple data structures (Series and DataFrame) and functions for working with structured data such as tables, spreadsheets, CSV files, JSON files, and database outputs. The name Pandas is derived from Panel Data, which refers to multidimensional structured datasets.
Five Advantages:
| # | Advantage | Explanation | Example |
|---|---|---|---|
| 1 | Data Loading | Reads data from many file formats and sources | read_csv(), read_json(), read_excel() |
| 2 | Data Cleaning | Handles missing values, duplicates, and incorrect formats | dropna(), fillna(), drop_duplicates() |
| 3 | Data Selection | Selects rows, columns, and subsets efficiently | df["Name"], loc[], iloc[] |
| 4 | Data Transformation | Creates new columns and modifies existing data | df["Total"] = df["A"] + df["B"] |
| 5 | Data Aggregation | Groups and summarizes data easily | groupby(), mean(), sum() |
| 6 | Data Export | Writes results back to files | to_csv(), to_json() |
Additional advantages: Handles mixed data types, supports labelled indexing, integrates well with NumPy and Matplotlib, and is faster than using basic Python lists and dictionaries for tabular data.
Q2. Differentiate between Pandas Series and DataFrame with suitable examples.
Answer:
| Aspect | Pandas Series | Pandas DataFrame |
|---|---|---|
| Definition | One-dimensional labelled array | Two-dimensional labelled data structure |
| Dimensions | 1D | 2D (rows × columns) |
| Similar to | A single column in a spreadsheet | A complete table with rows and columns |
| Index | Single index | Row index + column index |
| Data Types | Homogeneous (single dtype) | Heterogeneous (mixed dtypes) |
| Creation | pd.Series([1, 2, 3]) | pd.DataFrame({"A": [1, 2], "B": [3, 4]}) |
| Access | series[0], series["label"] | df["col"], df.loc[], df.iloc[] |
| Use Case | Single variable analysis | Multi-variable tabular analysis |
Example — Series:
import pandas as pd
marks = pd.Series([78, 85, 92, 66], index=["Amit", "Ravi", "Neha", "Sara"])
print(marks)
print(marks["Neha"]) # 92Example — DataFrame:
import pandas as pd
data = {
"Name": ["Amit", "Ravi", "Neha", "Sara"],
"Age": [21, 22, 20, 23],
"Marks": [78, 85, 92, 66]
}
df = pd.DataFrame(data)
print(df)Key Difference: A Series is a single column; a DataFrame is a collection of Series sharing the same row index.
Q3. Explain how to select rows and columns from a DataFrame using loc[] and iloc[].
Answer:
Pandas provides two primary methods for selecting rows and columns:
loc[]— label-based indexing (uses row/column labels)iloc[]— integer-position-based indexing (uses numeric positions)
Syntax:
df.loc[row_label, column_label]df.iloc[row_position, column_position]
Using loc[] (Label-Based):
print(df.loc[0]) # Row with label 0
print(df.loc[0:2]) # Rows 0 to 2 (inclusive)
print(df.loc[0:2, ["Name", "Marks"]]) # Specific rows and columns
print(df.loc[df["Marks"] > 80]) # Conditional selectionUsing iloc[] (Position-Based):
print(df.iloc[0]) # First row
print(df.iloc[0:3]) # Rows 0 to 2 (exclusive of 3)
print(df.iloc[0:3, 0:2]) # Rows 0-2, columns 0-1
print(df.iloc[-1]) # Last rowComparison Table:
| Method | Based On | Example | Result |
|---|---|---|---|
df.loc[0] | Label | df.loc[0] | Row with index label 0 |
df.iloc[0] | Position | df.iloc[0] | First row |
df.loc[0:2] | Label (inclusive) | df.loc[0:2] | Rows 0, 1, 2 |
df.iloc[0:3] | Position (exclusive) | df.iloc[0:3] | Rows 0, 1, 2 |
df.loc[:, "Name"] | Label | df.loc[:, "Name"] | All rows, “Name” column |
df.iloc[:, 0] | Position | df.iloc[:, 0] | All rows, first column |
Key Difference: loc[] uses labels (inclusive of endpoint); iloc[] uses integer positions (exclusive of endpoint, like Python slicing).
Q4. Write short notes on handling missing values in Pandas using dropna() and fillna().
Answer:
Missing values are common in real datasets. Pandas represents missing values using NaN (Not a Number). Two primary functions handle missing values:
1. dropna() — Remove Missing Values
Purpose: Removes rows or columns containing missing values.
Syntax: df.dropna(axis=0, how='any', thresh=None, subset=None, inplace=False)
Parameters:
| Parameter | Meaning | Default |
|---|---|---|
axis | 0 = drop rows; 1 = drop columns | 0 |
how | ‘any’ = drop if any NaN; ‘all’ = drop if all NaN | ‘any’ |
thresh | Minimum non-NaN values required to keep | None |
subset | Columns to consider | None |
inplace | Modify DataFrame in place | False |
Example:
import pandas as pd
data = {"Name": ["Amit", "Ravi", "Neha"], "Marks": [78, None, 92]}
df = pd.DataFrame(data)
print(df.dropna()) # Remove rows with any NaN
print(df.dropna(axis=1)) # Remove columns with any NaN2. fillna() — Replace Missing Values
Purpose: Replaces missing values with a specified value or method.
Syntax: df.fillna(value=None, method=None, axis=None, inplace=False)
Parameters:
| Parameter | Meaning |
|---|---|
value | Value to replace NaN with (scalar, dict, Series) |
method | ‘ffill’ (forward fill) or ‘bfill’ (backward fill) |
axis | 0 = rows; 1 = columns |
inplace | Modify DataFrame in place |
Example:
print(df.fillna(0)) # Replace NaN with 0
print(df.fillna(df.mean())) # Replace with column mean
print(df.fillna(method='ffill')) # Forward fill3. Other Missing Value Functions
| Function | Purpose |
|---|---|
isnull() | Detects missing values (returns True/False) |
notnull() | Detects non-missing values |
drop_duplicates() | Removes duplicate rows |
Best Practice: Choose the method based on the data and analysis goal. Dropping is suitable when missing data is random and small in proportion; filling is better when missing data is substantial or when preserving rows is important.
Q5. Explain how CSV and JSON files can be read and written using Pandas.
Answer:
Pandas can read data from many file formats. Two of the most common formats are CSV (Comma-Separated Values) and JSON (JavaScript Object Notation).
1. Reading CSV Files
Syntax: pd.read_csv(filepath, sep=',', header='infer', names=None, usecols=None)
import pandas as pd
# Read a CSV file
students = pd.read_csv("students.csv")
# Inspect the data
print(students.head()) # First 5 rows
print(students.shape) # (rows, columns)
print(students.info()) # Column types and non-null counts
print(students.describe()) # Statistical summaryCommon Parameters:
| Parameter | Purpose |
|---|---|
sep | Delimiter (default ‘,’) |
header | Row to use as column names |
names | Custom column names |
usecols | Subset of columns to read |
dtype | Data types for columns |
2. Writing CSV Files
Syntax: df.to_csv(filepath, index=False)
# Save a DataFrame to a CSV file
students.to_csv("cleaned_students.csv", index=False)3. Reading JSON Files
Syntax: pd.read_json(filepath, orient=None)
import pandas as pd
# Read a JSON file
orders = pd.read_json("orders.json")
print(orders.head())4. Writing JSON Files
Syntax: df.to_json(filepath, orient='records', indent=4)
# Save a DataFrame to a JSON file
orders.to_json("cleaned_orders.json", orient="records", indent=4)5. File Reading and Writing Functions Summary
| Function | Purpose | Common Parameter |
|---|---|---|
read_csv() | Reads a CSV file into a DataFrame | sep, header, names, usecols |
to_csv() | Writes a DataFrame to a CSV file | index=False |
read_json() | Reads a JSON file into a DataFrame | orient |
to_json() | Writes a DataFrame to a JSON file | orient, indent |
Practical Tip: After reading any file, always use head(), shape, info(), and describe() to understand the structure and quality of the imported data.
Q6. Define NumPy ndarray. Explain shape, ndim, size, and dtype attributes.
Answer:
Definition:NumPy (Numerical Python) is a fundamental Python library for scientific computing and numerical operations. It provides the ndarray object (N-dimensional array), which stores elements of the same data type in a compact and efficient way.
Compared with normal Python lists, NumPy arrays are:
- Faster — optimized C implementation
- More memory efficient — contiguous storage
- Support vectorized operations — operations on entire arrays without explicit loops
Creating an ndarray:
import numpy as np
arr = np.array([10, 20, 30, 40])
print(arr) # [10 20 30 40]
print(type(arr)) # <class 'numpy.ndarray'>Attributes of ndarray:
1. shape
Returns the dimensions of the array as a tuple (rows, columns, …).
a = np.array([[1, 2, 3], [4, 5, 6]])
print(a.shape) # (2, 3) — 2 rows, 3 columns2. ndim
Returns the number of dimensions (axes) of the array.
print(a.ndim) # 2 — two-dimensional array3. size
Returns the total number of elements in the array.
print(a.size) # 6 — 2 × 3 = 6 elements4. dtype
Returns the data type of the elements in the array.
print(a.dtype) # int64 (or int32 depending on platform)Summary Table:
| Attribute | Meaning | Example Output |
|---|---|---|
shape | Dimensions of the array | (2, 3) |
ndim | Number of dimensions | 2 |
size | Total number of elements | 6 |
dtype | Data type of elements | int64 |
Other Useful Attributes:
| Attribute | Meaning |
|---|---|
itemsize | Size in bytes of each element |
nbytes | Total bytes consumed by the array |
T | Transposed view of the array |
Why These Attributes Matter: They help you understand the structure and memory layout of your data before performing operations. For example, shape is essential for reshaping, and dtype determines which operations are valid.
Q7. Explain indexing and slicing in one-dimensional and two-dimensional NumPy arrays.
Answer:
Indexing is used to access individual elements, while slicing is used to access a range or subset of elements. NumPy indexing starts from 0, just like normal Python lists.
1. One-Dimensional Arrays
Syntax: array[start:stop:step]
import numpy as np
a = np.array([10, 20, 30, 40, 50])
print(a[0]) # 10 — first element
print(a[-1]) # 50 — last element
print(a[1:4]) # [20, 30, 40] — elements from index 1 to 3
print(a[:3]) # [10, 20, 30] — first three elements
print(a[::2]) # [10, 30, 50] — every second element
print(a[::-1]) # [50, 40, 30, 20, 10] — reversedIndexing and Slicing Examples:
| Expression | Meaning | Result |
|---|---|---|
a[0] | First element | 10 |
a[-1] | Last element | 50 |
a[1:4] | Elements from index 1 to 3 | [20, 30, 40] |
a[:3] | First three elements | [10, 20, 30] |
a[::2] | Every second element | [10, 30, 50] |
2. Two-Dimensional Arrays
Syntax: array[row, column] or array[row_start:row_stop, col_start:col_stop]
import numpy as np
matrix = np.array([[1, 2, 3],
[4, 5, 6],
[7, 8, 9]])
print(matrix[0, 0]) # 1 — row 0, column 0
print(matrix[1, 2]) # 6 — row 1, column 2
print(matrix[:, 1]) # [2, 5, 8] — all rows, column 1
print(matrix[0:2, 1:3]) # Sub-array [[2, 3], [5, 6]]
print(matrix[1, :]) # [4, 5, 6] — row 1, all columns
print(matrix[:, 0]) # [1, 4, 7] — all rows, column 02D Indexing and Slicing Examples:
| Expression | Meaning | Result |
|---|---|---|
matrix[0, 0] | Element at row 0, column 0 | 1 |
matrix[1, 2] | Element at row 1, column 2 | 6 |
matrix[:, 1] | All rows, column 1 | [2, 5, 8] |
matrix[0:2, 1:3] | Rows 0-1, columns 1-2 | [[2, 3], [5, 6]] |
matrix[1, :] | Row 1, all columns | [4, 5, 6] |
matrix[:, 0] | All rows, column 0 | [1, 4, 7] |
3. Key Points
- Indexing accesses a single element:
a[0],matrix[1, 2]. - Slicing accesses a range:
a[1:4],matrix[0:2, 1:3]. - Colon (
:) means “all” in that dimension. - Negative indices count from the end:
a[-1]is the last element. - Step can be used:
a[::2]takes every second element. - Slicing returns a view, not a copy — modifying the slice modifies the original array.
Example of View vs Copy:
a = np.array([1, 2, 3, 4, 5])
b = a[1:4] # b is a view
b[0] = 99
print(a) # [1, 99, 3, 4, 5] — original modifiedTo create a copy, use .copy():
b = a[1:4].copy()Q8. Discuss joining and splitting operations in NumPy with examples.
Answer:
NumPy provides functions to join (combine) and split (divide) arrays. These operations are essential for data manipulation and reshaping.
1. Joining Arrays
(a) concatenate() — Join Along an Existing Axis
import numpy as np
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
joined = np.concatenate((a, b))
print(joined) # [1 2 3 4 5 6]For 2D arrays:
x = np.array([[1, 2], [3, 4]])
y = np.array([[5, 6], [7, 8]])
print(np.concatenate((x, y), axis=0)) # Vertical (rows)
print(np.concatenate((x, y), axis=1)) # Horizontal (columns)(b) vstack() — Vertical Stacking
Stacks arrays vertically (row-wise).
print(np.vstack((x, y)))
# [[1 2]
# [3 4]
# [5 6]
# [7 8]](c) hstack() — Horizontal Stacking
Stacks arrays horizontally (column-wise).
print(np.hstack((x, y)))
# [[1 2 5 6]
# [3 4 7 8]]2. Splitting Arrays
(a) array_split() — Split into Multiple Parts
a = np.array([1, 2, 3, 4, 5, 6])
parts = np.array_split(a, 3)
print(parts)
# [array([1, 2]), array([3, 4]), array([5, 6])](b) split() — Equal Division
parts = np.split(a, 3) # Requires equal division(c) hsplit() and vsplit() — Split 2D Arrays
matrix = np.array([[1, 2, 3, 4], [5, 6, 7, 8]])
print(np.hsplit(matrix, 2)) # Split horizontally
print(np.vsplit(matrix, 2)) # Split vertically3. Reshaping
reshape() — Change Shape Without Changing Data
a = np.array([1, 2, 3, 4, 5, 6])
reshaped = a.reshape(2, 3)
print(reshaped)
# [[1 2 3]
# [4 5 6]]4. Summary Table
| Function | Purpose | Example |
|---|---|---|
concatenate() | Joins arrays along an existing axis | np.concatenate((a, b)) |
vstack() | Stacks arrays vertically | np.vstack((x, y)) |
hstack() | Stacks arrays horizontally | np.hstack((x, y)) |
array_split() | Splits an array into multiple parts | np.array_split(a, 3) |
split() | Splits into equal parts | np.split(a, 3) |
hsplit() | Splits horizontally | np.hsplit(matrix, 2) |
vsplit() | Splits vertically | np.vsplit(matrix, 2) |
reshape() | Changes shape without changing data | a.reshape(2, 3) |
Q9. Explain searching, sorting, and filtering arrays in NumPy.
Answer:
NumPy provides efficient functions for searching, sorting, and filtering arrays.
1. Searching Arrays
np.where() — Find Indices Where Condition is True
import numpy as np
a = np.array([10, 20, 30, 20, 40])
result = np.where(a == 20)
print(result) # (array([1, 3]),)np.argmax() and np.argmin() — Index of Max/Min
a = np.array([10, 50, 30, 20])
print(np.argmax(a)) # 1 (index of 50)
print(np.argmin(a)) # 0 (index of 10)2. Sorting Arrays
np.sort() — Returns Sorted Copy
a = np.array([40, 10, 30, 20])
print(np.sort(a)) # [10 20 30 40]array.sort() — Sorts In-Place
a.sort()
print(a) # [10 20 30 40]np.argsort() — Indices That Would Sort the Array
print(np.argsort(a)) # [1 3 2 0]3. Filtering Arrays
Filtering uses boolean masks — conditions that produce True/False for each element.
a = np.array([10, 25, 30, 45, 50])
# Filter values greater than 30
filtered = a[a > 30]
print(filtered) # [45 50]
# Filter with multiple conditions
filtered = a[(a > 20) & (a < 50)]
print(filtered) # [25 30 45]
# Filter even numbers
evens = a[a % 2 == 0]
print(evens) # [10 30 50]4. Summary Table
| Operation | Example | Output Meaning |
|---|---|---|
| Search | np.where(a == 20) | Returns indexes where condition is true |
| Sort | np.sort(a) | Returns sorted copy of the array |
| Filter | a[a > 30] | Returns values greater than 30 |
| Boolean mask | a % 2 == 0 | Creates True/False condition for each value |
| Argmax | np.argmax(a) | Index of maximum value |
| Argmin | np.argmin(a) | Index of minimum value |
| Argsort | np.argsort(a) | Indices that would sort the array |
Important: Filtering in NumPy is usually done using boolean conditions. The condition creates a True/False mask, and the mask is used to select matching elements.
Q10. What is web scraping? Explain the basic steps and ethical precautions involved.
Answer:
Definition:Web scraping is the process of extracting data from websites using programs. It is useful when information is available on web pages but not provided as a downloadable dataset or API.
1. Basic Web Scraping Workflow
Send HTTP Request → Download HTML → Parse & Extract Data → Clean & Structure → Store/Export2. Steps in a Basic Web Scraping Workflow
| Step | Meaning | Common Tool |
|---|---|---|
| 1. Send HTTP Request | Send a request to the target website | requests |
| 2. Download HTML | Retrieve the HTML content of the page | response.text |
| 3. Parse and Extract Data | Parse the HTML and extract the required data | BeautifulSoup |
| 4. Clean and Structure Data | Clean the extracted data and structure it | Pandas, manual cleaning |
| 5. Store or Export Data | Save the data to a file, database, or export for analysis | CSV, JSON, Pandas DataFrame |
3. Simple Web Scraping Example
import requests
from bs4 import BeautifulSoup
url = "https://example.com"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")
print(soup.title.text)Explanation:
requests.get(url)sends an HTTP GET request.response.textcontains the HTML source.BeautifulSoup(response.text, "html.parser")parses the HTML.soup.title.textextracts the page title.
4. Ethical Precautions
Web scraping should be performed responsibly. Programmers must respect:
| Precaution | Explanation |
|---|---|
| Terms of Service | Check if the website allows automated access |
| robots.txt | Follow the rules specified in the site’s robots.txt file |
| Copyright | Do not reproduce copyrighted content without permission |
| Privacy | Do not collect personal or restricted data without consent |
| Server Load | Avoid sending too many requests that could overload the server |
| Official APIs | Prefer official APIs whenever available |
Ethical Reminder: Before scraping a website, check whether the website allows automated access. Prefer official APIs whenever available.
5. Applications of Web Scraping
- Price comparison: Track product prices across e-commerce sites.
- News aggregation: Collect headlines from multiple news sources.
- Sentiment analysis: Gather social media posts for analysis.
- Research: Collect data for academic studies.
- Lead generation: Gather business contact information.
Q11. Explain the need for data visualization and list common Matplotlib chart types.
Answer:
Need for Data Visualization: Data visualization converts numbers into patterns. A good chart can reveal:
- Trends — how values change over time
- Comparisons — differences between categories
- Distributions — how data is spread
- Relationships — connections between variables
- Outliers — unusual data points
These may not be clear from a table alone. Visualization also helps communicate insights to non-technical audiences.
Common Matplotlib Chart Types:
| Chart Type | Used For | Example Question |
|---|---|---|
| Line plot | Showing trends over time | How did sales change month by month? |
| Bar chart | Comparing categories | Which department scored highest? |
| Histogram | Showing distribution of numerical values | What is the distribution of exam marks? |
| Scatter plot | Showing relationship between two numerical variables | Is study time related to marks? |
| Pie chart | Showing parts of a whole | What percentage of students selected each elective? |
Common Matplotlib pyplot Functions:
| Function | Purpose |
|---|---|
plot() | Creates a line plot |
bar() | Creates a bar chart |
hist() | Creates a histogram |
scatter() | Creates a scatter plot |
pie() | Creates a pie chart |
title() | Adds chart title |
xlabel(), ylabel() | Adds axis labels |
legend() | Displays chart legend |
show() | Displays the figure |
savefig() | Saves chart as an image file |
Visualization Rule: Every chart should have a clear title, axis labels where applicable, readable category names, and a purpose. Avoid adding unnecessary decoration.
Q12. Explain the Matplotlib visualization workflow.
Answer:
The Matplotlib visualization workflow consists of five steps:
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Prepare │ → │ 2. Create │ → │ 3. Plot Data │ → │ 4. Customize │ → │ 5. Show or │
│ Data │ │ Figure │ │ │ │ Plot │ │ Save │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘| Step | Action | Code Example |
|---|---|---|
| 1. Prepare Data | Load and organize data for visualization | x = [1, 2, 3], y = [10, 15, 13] |
| 2. Create Figure | Create a figure and axes | plt.figure(figsize=(8, 5)) |
| 3. Plot Data | Plot data using the appropriate plot type | plt.plot(x, y) |
| 4. Customize Plot | Add titles, labels, legends, colors | plt.title("Title"), plt.xlabel("X") |
| 5. Show or Save | Display or save the plot | plt.show(), plt.savefig("plot.png") |
Complete Example:
import matplotlib.pyplot as plt
# Step 1: Prepare Data
x = [1, 2, 3, 4, 5]
y = [10, 15, 13, 18, 20]
# Step 2: Create Figure (optional)
plt.figure(figsize=(8, 5))
# Step 3: Plot Data
plt.plot(x, y)
# Step 4: Customize Plot
plt.title("Simple Line Plot")
plt.xlabel("X Values")
plt.ylabel("Y Values")
# Step 5: Show or Save
plt.show()Q13. Explain how Pandas, NumPy, and Matplotlib can be used together in a data analysis project.
Answer:
Pandas, NumPy, and Matplotlib form the foundation of Python data analysis. They work together in a typical workflow:
| Stage | Library | Role |
|---|---|---|
| Data Loading | Pandas | Read CSV/JSON files into DataFrames |
| Data Cleaning | Pandas | Handle missing values, duplicates |
| Numerical Operations | NumPy | Fast array-based computations |
| Analysis | Pandas + NumPy | Grouping, aggregation, statistics |
| Visualization | Matplotlib | Create charts and graphs |
| Export | Pandas | Write results to CSV/JSON |
Integrated Example:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# Step 1: Create sample data (Pandas)
students = pd.DataFrame({
"Name": ["Amit", "Ravi", "Neha", "Sara", "John"],
"Python": [78, 85, 92, 66, 74],
"Maths": [80, 70, 88, 60, 79]
})
# Step 2: Calculate total and average (Pandas + NumPy)
students["Total"] = students["Python"] + students["Maths"]
students["Average"] = np.round(students["Total"] / 2, 2)
print(students)
print(students.describe())
# Step 3: Plot average marks (Matplotlib)
plt.bar(students["Name"], students["Average"])
plt.title("Average Marks of Students")
plt.xlabel("Student")
plt.ylabel("Average Marks")
plt.show()Role of Each Library:
| Library | Role in the Example |
|---|---|
| Pandas | Creates and manages the student DataFrame |
| NumPy | Rounds numerical average values |
| Matplotlib | Visualizes average marks as a bar chart |
Benefits of Integration:
- Pandas handles structured data with labels.
- NumPy provides fast numerical operations.
- Matplotlib communicates results visually.
- Together they form a complete data analysis pipeline.
Q14. Explain the Pandas workflow for tabular data analysis.
Answer:
The Pandas workflow consists of five stages:
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Import │ → │ 2. Create │ → │ 3. Clean & │ → │ 4. Combine │ → │ 5. Analyze │
│ Data │ │ Series/DF │ │ Transform │ │ Datasets │ │ & Export │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘| Stage | Action | Pandas Functions |
|---|---|---|
| 1. Import Data | Load data from various sources | read_csv(), read_json(), read_excel() |
| 2. Create Series/DataFrame | Organize data into Series or DataFrame structures | pd.Series(), pd.DataFrame() |
| 3. Clean & Transform | Handle missing values, filter, sort, transform data | dropna(), fillna(), drop_duplicates() |
| 4. Combine Datasets | Merge, join, or concatenate multiple datasets | concat(), merge(), join() |
| 5. Analyze & Export | Analyze data to extract insights and export results | groupby(), mean(), to_csv() |
Complete Example:
import pandas as pd
# Step 1: Import Data
df = pd.read_csv("students.csv")
# Step 2: Create/Inspect DataFrame
print(df.head())
print(df.info())
# Step 3: Clean & Transform
df = df.dropna()
df["Total"] = df["Internal"] + df["External"]
df["Average"] = df["Total"] / 2
# Step 4: Combine Datasets
extra = pd.read_csv("extra_students.csv")
combined = pd.concat([df, extra], ignore_index=True)
# Step 5: Analyze & Export
dept_avg = combined.groupby("Department")["Average"].mean()
print(dept_avg)
dept_avg.to_csv("department_averages.csv")Benefits of This Workflow:
- Systematic approach to data analysis.
- Each stage has clear inputs and outputs.
- Reproducible and maintainable.
- Integrates with NumPy and Matplotlib.
Q15. Explain the NumPy workflow for array-based computation.
Answer:
The NumPy workflow consists of five stages:
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1. Create │ → │ 2. Index & │ → │ 3. Vectorized│ → │ 4. Join / │ → │ 5. Search / │
│ ndarray │ │ Slice │ │ Operations │ │ Split │ │ Sort/Filter │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘| Stage | Action | NumPy Functions |
|---|---|---|
| 1. Create ndarray | Create arrays from lists, tuples, or other data structures | array(), zeros(), ones(), arange() |
| 2. Index & Slice | Access and extract specific elements, rows, or columns | a[0], a[1:4], matrix[0, 1] |
| 3. Vectorized Operations | Perform element-wise operations and math computations | a + b, a * 2, np.sqrt(a) |
| 4. Join / Split | Split arrays into multiple parts or join arrays together | concatenate(), vstack(), array_split() |
| 5. Search / Sort / Filter | Search for elements, sort arrays, apply filters | where(), sort(), a[a > 30] |
Complete Example:
import numpy as np
# Step 1: Create ndarray
a = np.array([10, 25, 30, 45, 50])
b = np.array([1, 2, 3, 4, 5])
# Step 2: Index & Slice
print(a[0]) # 10
print(a[1:4]) # [25 30 45]
# Step 3: Vectorized Operations
print(a + b) # [11 27 33 49 55]
print(a * 2) # [20 50 60 90 100]
# Step 4: Join / Split
joined = np.concatenate((a, b))
parts = np.array_split(a, 5)
# Step 5: Search / Sort / Filter
print(np.where(a > 30)) # (array([3, 4]),)
print(np.sort(a)) # [10 25 30 45 50]
print(a[a > 30]) # [45 50]Benefits of NumPy Workflow:
- Fast numerical operations (C-based).
- Memory efficient.
- Vectorization eliminates explicit loops.
- Foundation for Pandas, SciPy, and machine learning libraries.
Q16. Explain the difference between np.sort() and array.sort().
Answer:
| Aspect | np.sort() | array.sort() |
|---|---|---|
| Type | Function | Method |
| Syntax | np.sort(a) | a.sort() |
| Return | Returns a sorted copy | Returns None; sorts in-place |
| Original Array | Unchanged | Modified |
| Use Case | When original order must be preserved | When in-place sorting is acceptable |
| Example | sorted_a = np.sort(a) | a.sort() |
Example:
import numpy as np
# np.sort() — returns copy
a = np.array([40, 10, 30, 20])
b = np.sort(a)
print(a) # [40 10 30 20] — unchanged
print(b) # [10 20 30 40] — sorted copy
# array.sort() — in-place
a = np.array([40, 10, 30, 20])
a.sort()
print(a) # [10 20 30 40] — modifiedKey Difference: np.sort() preserves the original array; array.sort() modifies it in-place and returns None.
Q17. Explain the difference between pd.concat(), pd.merge(), and pd.join().
Answer:
| Aspect | pd.concat() | pd.merge() | pd.join() |
|---|---|---|---|
| Purpose | Stacks DataFrames | Combines using common column | Combines using index |
| Axis | Row-wise (axis=0) or column-wise (axis=1) | Column-wise (like SQL JOIN) | Column-wise (index-based) |
| Key | No key required | Requires a common column | Uses index labels |
| Use Case | Same columns, different rows | Tables sharing a key (ID) | Index labels are meaningful |
| Example | pd.concat([df1, df2]) | pd.merge(a, b, on="ID") | a.join(b) |
Examples:
import pandas as pd
# concat — stacking
df1 = pd.DataFrame({"ID": [1, 2], "Name": ["Amit", "Ravi"]})
df2 = pd.DataFrame({"ID": [3, 4], "Name": ["Neha", "Sara"]})
combined = pd.concat([df1, df2], ignore_index=True)
# merge — common column
marks = pd.DataFrame({"ID": [1, 2, 3], "Marks": [78, 85, 92]})
students = pd.DataFrame({"ID": [1, 2, 3], "Name": ["Amit", "Ravi", "Neha"]})
result = pd.merge(students, marks, on="ID")
# join — index-based
a = pd.DataFrame({"Marks": [78, 85]}, index=[1, 2])
b = pd.DataFrame({"Name": ["Amit", "Ravi"]}, index=[1, 2])
joined = a.join(b)When to Use:
concat()— When datasets have the same columns or same index.merge()— When tables share a key such as ID.join()— When index labels are meaningful.
Q18. Explain the describe() method in Pandas.
Answer:
The describe() method provides a statistical summary of numerical columns in a DataFrame.
Syntax: df.describe()
Output Columns:
| Statistic | Meaning |
|---|---|
count | Number of non-null values |
mean | Average value |
std | Standard deviation |
min | Minimum value |
25% | First quartile (Q1) |
50% | Median (Q2) |
75% | Third quartile (Q3) |
max | Maximum value |
Example:
import pandas as pd
data = {
"Name": ["Amit", "Ravi", "Neha", "Sara"],
"Age": [21, 22, 20, 23],
"Marks": [78, 85, 92, 66]
}
df = pd.DataFrame(data)
print(df.describe())Output:
Age Marks
count 4.000000 4.000000
mean 21.500000 80.250000
std 1.290994 11.056000
min 20.000000 66.000000
25% 20.750000 75.000000
50% 21.500000 81.500000
75% 22.250000 87.250000
max 23.000000 92.000000For Categorical Columns:
print(df.describe(include="all"))This includes unique, top, and freq for categorical columns.
Use Case: Quickly understand the distribution, central tendency, and spread of numerical data before deeper analysis.
Q19. Explain the info() method in Pandas.
Answer:
The info() method provides a concise summary of a DataFrame, including:
- Number of rows and columns
- Column names
- Data types of each column
- Number of non-null values
- Memory usage
Syntax: df.info()
Example:
import pandas as pd
data = {
"Name": ["Amit", "Ravi", "Neha"],
"Age": [21, 22, 20],
"Marks": [78, None, 92]
}
df = pd.DataFrame(data)
df.info()Output:
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 3 entries, 0 to 2
Data columns (total 3 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 Name 3 non-null object
1 Age 3 non-null int64
2 Marks 2 non-null float64
dtypes: float64(1), int64(1), object(1)
memory usage: 200.0+ bytesKey Information Provided:
| Field | Meaning |
|---|---|
RangeIndex | Row index range |
Data columns | Number of columns |
Non-Null Count | Number of non-missing values |
Dtype | Data type of each column |
memory usage | Memory consumed by the DataFrame |
Use Case: Quickly identify missing values, data types, and memory usage before cleaning or analysis.
Q20. Explain the head() and tail() methods in Pandas.
Answer:
head() and tail() are used to preview the beginning and end of a DataFrame.
| Method | Purpose | Default |
|---|---|---|
df.head() | Displays first 5 rows | 5 |
df.tail() | Displays last 5 rows | 5 |
Syntax:
df.head(n) # First n rows
df.tail(n) # Last n rowsExample:
import pandas as pd
data = {
"Name": ["Amit", "Ravi", "Neha", "Sara", "John", "Priya"],
"Marks": [78, 85, 92, 66, 74, 88]
}
df = pd.DataFrame(data)
print(df.head()) # First 5 rows
print(df.head(3)) # First 3 rows
print(df.tail()) # Last 5 rows
print(df.tail(2)) # Last 2 rowsOutput:
# df.head()
Name Marks
0 Amit 78
1 Ravi 85
2 Neha 92
3 Sara 66
4 John 74
# df.tail(2)
Name Marks
4 John 74
5 Priya 88Use Case: Quickly inspect data structure, column names, and sample values after loading a file.
SECTION C: ANALYTICAL QUESTIONS (10)
Q1. Analyze the following code and predict the output. Explain each line.
import pandas as pd
data = {
"Name": ["Amit", "Ravi", "Neha"],
"Marks": [78, 85, 92]
}
df = pd.DataFrame(data)
df["Total"] = df["Marks"] + 10
print(df)
print(df.shape)Answer:
Output:
Name Marks Total
0 Amit 78 88
1 Ravi 85 95
2 Neha 92 102
(3, 3)Line-by-Line Explanation:
| Line | Explanation |
|---|---|
import pandas as pd | Imports Pandas with alias pd |
data = {...} | Creates a dictionary with two lists |
df = pd.DataFrame(data) | Creates a DataFrame with 3 rows and 2 columns |
df["Total"] = df["Marks"] + 10 | Creates a new column “Total” by adding 10 to “Marks” |
print(df) | Displays the DataFrame with 3 columns |
print(df.shape) | Displays (3, 3) — 3 rows, 3 columns |
Key Insight: Pandas allows easy creation of new columns through vectorized operations.
Q2. Analyze the following NumPy code and predict the output.
import numpy as np
a = np.array([10, 20, 30, 40, 50])
print(a[1:4])
print(a[a > 30])
print(np.where(a == 30))
print(np.sort(a)[::-1])Answer:
Output:
[20 30 40]
[40 50]
(array([2]),)
[50 40 30 20 10]Explanation:
| Line | Output | Explanation |
|---|---|---|
a[1:4] | [20 30 40] | Elements from index 1 to 3 |
a[a > 30] | [40 50] | Boolean filter for values > 30 |
np.where(a == 30) | (array([2]),) | Index 2 contains 30 |
np.sort(a)[::-1] | [50 40 30 20 10] | Sorts ascending, then reverses |
Key Insight: NumPy supports slicing, boolean filtering, searching, and sorting in a single line.
Q3. A student writes the following code to read a CSV file. Identify the issues and fix them.
import pandas as pd
df = pd.read_csv("students.csv")
print(df.head)
print(df.shape())
print(df.describe)Answer:
Issues Identified:
df.headmissing parentheses:headis a method; it must be called with(). Without parentheses, it returns the method object, not the data.df.shape()incorrect:shapeis an attribute, not a method. It should bedf.shapewithout parentheses.df.describemissing parentheses:describeis a method; it must be called with().
Corrected Code:
import pandas as pd
df = pd.read_csv("students.csv")
print(df.head()) # Method call with parentheses
print(df.shape) # Attribute without parentheses
print(df.describe()) # Method call with parenthesesExplanation:
- Methods (functions attached to objects) require
():head(),describe(),info(). - Attributes (properties) do not use
():shape,columns,dtypes.
Key Insight: Confusing methods and attributes is a common beginner mistake in Pandas.
Q4. Analyze the following code and identify what happens when the file is missing.
import pandas as pd
df = pd.read_csv("data.csv")
print(df.head())Answer:
Problem: If data.csv doesn’t exist, pd.read_csv() raises FileNotFoundError, and the program crashes.
Improved Code:
import pandas as pd
try:
df = pd.read_csv("data.csv")
print(df.head())
except FileNotFoundError:
print("Error: 'data.csv' not found.")
except pd.errors.EmptyDataError:
print("Error: File is empty.")
except Exception as error:
print("Unexpected error:", error)Improvements:
try-excepthandles missing file gracefully.EmptyDataErrorhandles empty CSV files.- Generic
exceptcatches unexpected errors. - User-friendly error messages.
Key Insight: Always handle file operations with exception handling to prevent crashes.
Q5. Analyze the following Matplotlib code and predict the output. Suggest improvements.
import matplotlib.pyplot as plt
x = [1, 2, 3, 4]
y = [10, 20, 30, 40]
plt.plot(x, y)
plt.show()Answer:
Output: A simple line plot connecting points (1,10), (2,20), (3,30), (4,40).
Issues:
- No title: The chart lacks a descriptive title.
- No axis labels: X and Y axes are unlabeled.
- No legend: Not needed for a single line, but useful for multiple lines.
- No grid: Grid lines improve readability.
- No color/style: Default styling is plain.
Improved Code:
import matplotlib.pyplot as plt
x = [1, 2, 3, 4]
y = [10, 20, 30, 40]
plt.figure(figsize=(8, 5))
plt.plot(x, y, marker='o', color='blue', linestyle='-', linewidth=2)
plt.title("Simple Line Plot")
plt.xlabel("X Values")
plt.ylabel("Y Values")
plt.grid(True, alpha=0.3)
plt.show()Improvements:
- Added title and axis labels.
- Added markers and line styling.
- Added grid for readability.
- Set figure size for better proportions.
Key Insight: Every chart should have a clear title, axis labels, and readable styling.
Q6. Analyze the following Pandas code for handling missing values. What is the output?
import pandas as pd
import numpy as np
data = {"Name": ["Amit", "Ravi", "Neha"], "Marks": [78, np.nan, 92]}
df = pd.DataFrame(data)
print(df.isnull())
print(df.dropna())
print(df.fillna(0))Answer:
Output:
print(df.isnull()):
Name Marks
0 False False
1 False True
2 False Falseprint(df.dropna()):
Name Marks
0 Amit 78.0
2 Neha 92.0print(df.fillna(0)):
Name Marks
0 Amit 78.0
1 Ravi 0.0
2 Neha 92.0Explanation:
| Function | Behavior |
|---|---|
isnull() | Returns True where values are NaN |
dropna() | Removes rows with any NaN (row 1 removed) |
fillna(0) | Replaces NaN with 0 (row 1’s Marks becomes 0.0) |
Key Insight: dropna() reduces data; fillna() preserves rows by substituting values.
Q7. A student wants to analyze student performance. Design a program using Pandas, NumPy, and Matplotlib.
Answer:
Scenario: Read student marks from a CSV, calculate total and average, identify top performers, and visualize results.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# Step 1: Load data
df = pd.read_csv("students.csv")
# Step 2: Inspect data
print(df.head())
print(df.info())
print(df.describe())
# Step 3: Calculate total and average
df["Total"] = df["Python"] + df["Maths"] + df["DBMS"]
df["Average"] = np.round(df["Total"] / 3, 2)
# Step 4: Identify top performers
top = df[df["Average"] > 85]
print("Top performers:")
print(top[["Name", "Average"]])
# Step 5: Visualize
plt.figure(figsize=(10, 5))
plt.bar(df["Name"], df["Average"], color="steelblue")
plt.axhline(y=df["Average"].mean(), color="red",
linestyle="--", label="Class Average")
plt.title("Student Average Marks")
plt.xlabel("Student")
plt.ylabel("Average Marks")
plt.legend()
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()Explanation:
- Pandas: Loads CSV, calculates columns, filters top performers.
- NumPy: Rounds average values.
- Matplotlib: Visualizes averages with a class average reference line.
Key Insight: The three libraries work together to form a complete data analysis pipeline.
Q8. Analyze the following NumPy code for joining and splitting. What is the output?
import numpy as np
a = np.array([1, 2, 3])
b = np.array([4, 5, 6])
joined = np.concatenate((a, b))
parts = np.array_split(joined, 3)
reshaped = joined.reshape(2, 3)
print(joined)
print(parts)
print(reshaped)Answer:
Output:
[1 2 3 4 5 6]
[array([1, 2]), array([3, 4]), array([5, 6])]
[[1 2 3]
[4 5 6]]Explanation:
| Line | Output | Explanation |
|---|---|---|
np.concatenate((a, b)) | [1 2 3 4 5 6] | Joins arrays |
np.array_split(joined, 3) | 3 sub-arrays | Splits into 3 parts |
joined.reshape(2, 3) | 2×3 matrix | Reshapes into 2 rows, 3 columns |
Key Insight: NumPy provides flexible joining, splitting, and reshaping operations for array manipulation.
Q9. Analyze the following web scraping code. Identify ethical issues and suggest improvements.
import requests
from bs4 import BeautifulSoup
for i in range(1, 1000):
url = f"https://example.com/page{i}"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")
print(soup.title.text)Answer:
Ethical Issues:
- No rate limiting: 1,000 rapid requests can overload the server.
- No robots.txt check: The program ignores the site’s scraping rules.
- No User-Agent header: Identifies the request as a bot without transparency.
- No error handling: Network errors or blocked requests crash the program.
- Possible Terms of Service violation: Scraping may be prohibited.
Improved Code:
import requests
from bs4 import BeautifulSoup
import time
headers = {"User-Agent": "Mozilla/5.0 (educational scraper)"}
for i in range(1, 20): # Fewer pages
url = f"https://example.com/page{i}"
try:
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
print(soup.title.text)
time.sleep(2) # 2-second delay between requests
except requests.RequestException as e:
print(f"Error fetching{url}:{e}")Improvements:
- Added User-Agent header for transparency.
- Added rate limiting (
time.sleep(2)). - Added error handling.
- Reduced number of requests.
- Respects server load.
Key Insight: Web scraping must be ethical — respect robots.txt, terms of service, privacy, and server load.
Q10. Design a complete program that combines Pandas, NumPy, Matplotlib, and file handling to analyze sales data.
Answer:
Scenario: Read sales data from a CSV, compute monthly totals using NumPy, and visualize trends with Matplotlib.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
def analyze_sales(input_file, output_file):
"""Read sales data, analyze, and visualize."""
try:
# Step 1: Load data (Pandas)
df = pd.read_csv(input_file)
print("Data loaded:")
print(df.head())
# Step 2: Inspect data
print("\nData info:")
print(df.info())
# Step 3: Calculate totals (NumPy)
df["Total"] = df["Quantity"] * df["Price"]
df["Total"] = np.round(df["Total"], 2)
# Step 4: Monthly aggregation (Pandas)
monthly = df.groupby("Month")["Total"].sum().reset_index()
# Step 5: Save results (Pandas)
monthly.to_csv(output_file, index=False)
print(f"\nMonthly totals saved to{output_file}")
# Step 6: Visualize (Matplotlib)
plt.figure(figsize=(10, 5))
plt.plot(monthly["Month"], monthly["Total"],
marker="o", color="green", linewidth=2)
plt.title("Monthly Sales Trend")
plt.xlabel("Month")
plt.ylabel("Total Sales")
plt.grid(True, alpha=0.3)
plt.xticks(rotation=45)
plt.tight_layout()
plt.savefig("sales_trend.png")
plt.show()
print("Chart saved as sales_trend.png")
except FileNotFoundError:
print(f"Error: '{input_file}' not found.")
except KeyError as e:
print(f"Missing column:{e}")
except Exception as error:
print("Unexpected error:", error)
# Usage
analyze_sales("sales.csv", "monthly_sales.csv")Sample Input (sales.csv):
Month,Product,Quantity,Price
Jan,A,10,100
Jan,B,5,200
Feb,A,15,100
Feb,B,8,200
Mar,A,20,100
Mar,B,12,200Sample Output (monthly_sales.csv):
Month,Total
Feb,3100.0
Jan,2000.0
Mar,4400.0Concepts Demonstrated:
| Concept | Implementation |
|---|---|
| File handling | pd.read_csv(), to_csv() |
| Pandas DataFrame | df.head(), df.info(), groupby() |
| NumPy operations | np.round() |
| Matplotlib visualization | plt.plot(), plt.savefig() |
| Exception handling | try-except blocks |
| Functions | analyze_sales() |
Key Insight: Combining these libraries creates a complete data analysis pipeline — from raw data to actionable insights and visualizations.
SUMMARY TABLE
| Section | Count | Topics Covered |
|---|---|---|
| MCQ | 50 | Pandas, Series/DataFrame, NumPy, indexing, joining, sorting, filtering, web scraping, Matplotlib |
| Theory | 20 | Pandas workflow, NumPy workflow, file I/O, missing values, visualization, integrated examples |
| Analytical | 10 | Code analysis, error identification, program design, integrated projects |