Adv-Python Unit 1: Complete Concept Guide
Unit I: Introduction to Computer Vision and Video Processing -> Generated and Prepared By Thiruselvan (ThiruXD)
1. Overview of Computer Vision
Computer Vision is a field of Artificial Intelligence that enables computers to interpret and understand visual information from the world (images and videos), similar to human vision.
Key Goals
- Acquire, process, analyze, and understand digital images/videos
- Extract meaningful information (objects, scenes, motion, text, etc.)
- Make decisions or take actions based on visual data
Difference from Image Processing
- Image Processing focuses on improving or transforming images (enhancement, filtering, compression).
- Computer Vision focuses on understanding the content (object detection, recognition, scene understanding).
Importance
- Autonomous vehicles, medical imaging, surveillance, robotics, AR/VR, industrial inspection, agriculture, retail (self-checkout), sports analytics, etc.
2. Introduction to OpenCV Library
OpenCV (Open Source Computer Vision Library) is the most widely used open-source library for computer vision, machine learning, and image processing.
- Originally developed by Intel (2000)
- Now maintained by OpenCV.org
- Written primarily in C++ with excellent Python, Java, and other bindings
- Cross-platform (Windows, Linux, macOS, Android, iOS)
- Highly optimized (uses multi-threading, SIMD, GPU acceleration via CUDA/OpenCL)
Installation (Python)
pip install opencv-python
# For extra modules (contrib):
pip install opencv-contrib-pythonBasic Import
import cv2
import numpy as np # Almost always used together with OpenCVChecking Version
print(cv2.__version__)3. Features and Capabilities of OpenCV
OpenCV provides modules for almost every computer-vision task:
| Module / Capability | Description | Common Use Cases |
|---|---|---|
| Core | Basic structures (Mat, arrays), operations | Image creation, arithmetic |
| Imgproc | Image processing (filtering, morphology, geometry) | Smoothing, edge detection, resizing |
| Highgui | GUI, image/video I/O, display | Reading/writing/displaying images |
| Videoio | Video capture and writing | Webcam, video files |
| Features2d | Feature detection & description (SIFT, ORB...) | Object matching, tracking |
| Objdetect | Object detection (Haar, HOG, Cascade) | Face, pedestrian detection |
| Calib3d | Camera calibration & 3D reconstruction | Stereo vision, AR |
| DNN | Deep Neural Network module | Running YOLO, ResNet, etc. |
| ML | Classical machine learning | SVM, KNN, decision trees |
| Photo | Computational photography | Denoising, HDR, inpainting |
| Stitching | Image stitching | Panorama creation |
Strengths
- Extremely fast (optimized C++ backend)
- Huge community and documentation
- Works with NumPy arrays natively in Python
- Supports real-time processing
4. Applications and Importance of Computer Vision
Real-world Applications
- Face detection & recognition (security, unlocking phones)
- Object detection & tracking (autonomous cars, drones)
- Medical image analysis (tumor detection, X-ray)
- Optical Character Recognition (OCR)
- Industrial quality control
- Augmented Reality / Virtual Reality
- Sports analytics and broadcast
- Agriculture (crop health monitoring)
- Retail (people counting, inventory)
Why it matters today Deep learning + powerful libraries like OpenCV have made previously impossible tasks practical and real-time.
5. Image Representation
A digital image is a 2-D or 3-D array of numbers.
- Grayscale image → 2D matrix (height × width)
- Color image → 3D array (height × width × channels)
In OpenCV (and most computer vision libraries):
- Images are stored as NumPy arrays of type
uint8(0–255) by default. - Coordinate system: Origin (0,0) is at the top-left corner.
- x increases to the right
- y increases downward
6. Pixels and Image Resolution
Pixel (Picture Element)
The smallest unit of a digital image. Each pixel stores intensity (and color) information.
Resolution
- Spatial resolution = number of pixels (e.g., 1920 × 1080)
- Higher resolution → more detail, larger file size, more computation
Bit Depth
- 8-bit → 256 intensity levels (most common)
- 16-bit → used in medical/scientific imaging
- 32-bit float → used in intermediate processing
Common Image Formats
- Lossy: JPEG
- Lossless: PNG, BMP, TIFF
- OpenCV supports almost all common formats.
7. Color Models
OpenCV uses BGR by default (not RGB)!
| Color Model | Channels | Description | OpenCV Flag |
|---|---|---|---|
| BGR | Blue, Green, Red | Default in OpenCV | – |
| RGB | Red, Green, Blue | Human-friendly, many other libraries | cv2.COLOR_BGR2RGB |
| Grayscale | Single intensity | 0 = black, 255 = white | cv2.COLOR_BGR2GRAY |
| HSV | Hue, Saturation, Value | Better for color-based segmentation | cv2.COLOR_BGR2HSV |
| HLS / LAB | Alternative perceptual models | Useful for advanced processing | Available |
Why HSV is popular
- Hue separates color information from intensity → robust to lighting changes.
Conversion Example
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)8. Reading Images
img = cv2.imread('path/to/image.jpg') # BGR, color
img = cv2.imread('path/to/image.jpg', 0) # Grayscale
img = cv2.imread('path/to/image.jpg', cv2.IMREAD_UNCHANGED) # With alpha if presentImportant Flags
cv2.IMREAD_COLOR(1) – defaultcv2.IMREAD_GRAYSCALE(0)cv2.IMREAD_UNCHANGED(-1)
Always check if image loaded successfully
if img is None:
print("Error: Image not found or cannot be read")9. Writing and Displaying Images
Saving
cv2.imwrite('output.jpg', img)
cv2.imwrite('output.png', img) # PNG is losslessDisplaying
cv2.imshow('Window Name', img)
cv2.waitKey(0) # Wait indefinitely for a key press
cv2.destroyAllWindows()Modern Alternative (especially in Jupyter / scripts) Many people now prefer Matplotlib or OpenCV’s highgui carefully, or use:
from matplotlib import pyplot as plt
plt.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))
plt.axis('off')
plt.show()10. Image Manipulation and Cropping
Basic Operations
# Resize
resized = cv2.resize(img, (width, height))
resized = cv2.resize(img, None, fx=0.5, fy=0.5) # scale factors
# Crop (NumPy slicing) – y first, then x
cropped = img[y1:y2, x1:x2]
# Flip
flipped = cv2.flip(img, 1) # 0=vertical, 1=horizontal, -1=both
# Rotate (using rotation matrix)
M = cv2.getRotationMatrix2D((cx, cy), angle, scale)
rotated = cv2.warpAffine(img, M, (w, h))Drawing on Images
cv2.line(img, pt1, pt2, color, thickness)
cv2.rectangle(img, pt1, pt2, color, thickness) # thickness=-1 → filled
cv2.circle(img, center, radius, color, thickness)
cv2.putText(img, text, org, font, scale, color, thickness)11. Understanding Image Properties
print(img.shape) # (height, width, channels)
print(img.shape[0]) # height
print(img.shape[1]) # width
print(img.shape[2]) # channels (if color)
print(img.dtype) # usually uint8
print(img.size) # total number of pixels × channels
print(type(img)) # <class 'numpy.ndarray'>Accessing Individual Pixels
# Color image
b, g, r = img[y, x]
img[y, x] = [0, 255, 0] # set pixel to green
# Grayscale
intensity = gray[y, x]ROI (Region of Interest)
roi = img[100:300, 200:400]
img[100:300, 200:400] = some_other_imageKey Takeaways
- OpenCV stores images as NumPy arrays in BGR order.
- Always check whether
cv2.imread()succeeded. - Most image operations are just NumPy array operations + specialized OpenCV functions.
- Understanding pixels, resolution, color spaces, and basic I/O is the foundation for everything that follows (video processing, object detection, etc.).
Advanced Python Index
Complete unit-wise directory mapping interactive concept guides, question banks, and source PDFs -> Generated and Prepared By Thiruselvan (ThiruXD).
Adv-Python Unit 1: Questions with Answers
Unit I: Introduction to Computer Vision and Video Processing -> Generated and Prepared By Thiruselvan (ThiruXD)