Adv-Python Unit 2: Complete Concept Guide
Unit II: Video Processing using OpenCV -> Generated and Prepared By Thiruselvan (ThiruXD)
2.1 Video Files and Formats in OpenCV
Concept Explanation:
A video is a collection of images (called frames) displayed rapidly one after another to create the illusion of motion.
Important components of a video:
- Frames: Individual images
- Frame Rate (FPS): Number of frames shown per second (common values: 24, 25, 30, 60)
- Resolution: Width × Height of each frame
- Codec: Algorithm used to compress and decompress the video
- Container Format: File type that holds the video data (.mp4, .avi, .mkv, etc.)
OpenCV does not handle audio properly. It mainly works with the video stream.
Common Formats & Codecs:
.avi→ Older but widely supported (codecs: XVID, MJPG).mp4→ Most popular modern format (codecs: H.264 / avc1, mp4v)- FourCC code is a 4-character code used by OpenCV to specify the codec while writing a video.
2.2 Extracting Frames from Videos
Concept Explanation:
Extracting frames means saving individual images from a video.
This is useful for:
- Creating datasets
- Analyzing specific moments
- Applying image processing on selected frames
We can extract:
- All frames
- Every nth frame
- Frame at a particular time (in seconds)
2.3 Reading Videos from File
Concept Explanation:
To process a video, we first need to open it using cv2.VideoCapture().
Working pattern (very important for exams):
- Create a VideoCapture object
- Check if the video opened successfully
- Run a loop to read frames one by one
- Process each frame
- Display or save the frame
- Release the resources
cap.read() returns two values:
ret→ Boolean (True if frame is read successfully)frame→ The actual image
2.4 Capturing Videos from Camera
Concept Explanation:
Instead of a video file, we can capture live video from a webcam or external camera.
cv2.VideoCapture(0)→ Default webcamcv2.VideoCapture(1)→ External camera
The rest of the process is almost the same as reading from a file.
Only difference: We usually use a very small delay in waitKey() (1 ms) for real-time feel.
2.5 Writing and Displaying Videos
Concept Explanation:
Displaying: Use cv2.imshow() inside the loop.
Writing (Saving) a video requires:
- Knowing the frame size (width, height)
- Knowing the FPS
- Selecting a codec (FourCC)
- Creating a
VideoWriterobject - Writing each processed frame using
out.write(frame)
(2.6, 2.7, 2.8) Drawing Lines, Rectangles, Circles and Adding Text (with Date & Time)
Concept Explanation:
All drawing functions work exactly the same way as on images.
Key Point: Drawing must be done inside the while loop on every frame, otherwise it will appear only on one frame.
Common drawing functions:
cv2.line()cv2.rectangle()cv2.circle()cv2.putText()
For live date and time, we use Python’s datetime module and update the text every frame.
2.9 Setting Video Properties
Concept Explanation:
OpenCV allows us to get and set various properties of a video or camera using:
cap.get(property_id)cap.set(property_id, value)
Important properties:
CAP_PROP_FRAME_WIDTHCAP_PROP_FRAME_HEIGHTCAP_PROP_FPSCAP_PROP_FRAME_COUNT(total frames)CAP_PROP_POS_FRAMES(current frame number)CAP_PROP_POS_MSEC(current time in milliseconds)
Note: Not all properties can be changed on all cameras/videos.
2.10 Motion Detection using OpenCV
Concept Explanation:
Motion detection means identifying moving objects in a video.
Basic Principle (Frame Differencing):
- Take two consecutive frames
- Find the difference between them
- Convert difference to grayscale
- Apply threshold to get a binary image
- Find contours of white regions
- Draw bounding boxes around large contours (moving objects)
Advanced methods (for knowledge):
- Background Subtraction (MOG2, KNN)
- Optical Flow
Exams usually expect the simple frame differencing method.
2.11 Object Tracking using OpenCV
Concept Explanation:
Object Tracking means locating the same object across multiple frames after it has been initially selected.
Difference from Detection:
- Detection → Finding objects in every frame independently
- Tracking → Following an object once it is detected/selected
OpenCV provides several trackers (in opencv-contrib):
- CSRT → Most accurate
- KCF → Good balance of speed and accuracy
- MOSSE → Very fast
- MIL, TLD, MEDIANFLOW, etc.
Basic steps:
- Select the object (ROI) in the first frame
- Initialize the tracker
- In every next frame, update the tracker
- Draw the updated bounding box
Practical Code Examples (Concept → Code)
1. Reading a Video File
import cv2
cap = cv2.VideoCapture('video.mp4')
if not cap.isOpened():
print("Error opening video")
else:
while True:
ret, frame = cap.read()
if not ret:
break
cv2.imshow('Video', frame)
if cv2.waitKey(25) & 0xFF == ord('q'):
break
cap.release()
cv2.destroyAllWindows()2. Capturing from Webcam
import cv2
cap = cv2.VideoCapture(0)
while True:
ret, frame = cap.read()
if not ret:
break
cv2.imshow('Webcam', frame)
if cv2.waitKey(1) & 0xFF == ord('q'):
break
cap.release()
cv2.destroyAllWindows()3. Writing (Saving) a Video
import cv2
cap = cv2.VideoCapture(0)
width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
fps = 20
fourcc = cv2.VideoWriter_fourcc(*'XVID')
out = cv2.VideoWriter('output.avi', fourcc, fps, (width, height))
while True:
ret, frame = cap.read()
if not ret:
break
out.write(frame)
cv2.imshow('Recording', frame)
if cv2.waitKey(1) & 0xFF == ord('q'):
break
cap.release()
out.release()
cv2.destroyAllWindows()4. Drawing + Date & Time on Video
import cv2
import datetime
cap = cv2.VideoCapture(0)
while True:
ret, frame = cap.read()
if not ret:
break
# Drawing
cv2.rectangle(frame, (100, 100), (300, 250), (0, 255, 0), 2)
cv2.circle(frame, (400, 200), 40, (0, 0, 255), -1)
cv2.line(frame, (50, 50), (250, 50), (255, 0, 0), 3)
# Date and Time
now = datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")
cv2.putText(frame, now, (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 0.8, (0, 255, 255), 2)
cv2.imshow('Annotated', frame)
if cv2.waitKey(1) & 0xFF == ord('q'):
break
cap.release()
cv2.destroyAllWindows()5. Basic Motion Detection
import cv2
cap = cv2.VideoCapture(0)
ret, frame1 = cap.read()
ret, frame2 = cap.read()
while True:
diff = cv2.absdiff(frame1, frame2)
gray = cv2.cvtColor(diff, cv2.COLOR_BGR2GRAY)
blur = cv2.GaussianBlur(gray, (5,5), 0)
_, thresh = cv2.threshold(blur, 20, 255, cv2.THRESH_BINARY)
dilated = cv2.dilate(thresh, None, iterations=3)
contours, _ = cv2.findContours(dilated, cv2.RETR_TREE, cv2.CHAIN_APPROX_SIMPLE)
for c in contours:
if cv2.contourArea(c) < 900:
continue
x, y, w, h = cv2.boundingRect(c)
cv2.rectangle(frame1, (x,y), (x+w, y+h), (0,255,0), 2)
cv2.putText(frame1, "Motion", (10,20), cv2.FONT_HERSHEY_SIMPLEX, 0.7, (0,0,255), 2)
cv2.imshow("Motion Detection", frame1)
frame1 = frame2
ret, frame2 = cap.read()
if cv2.waitKey(1) & 0xFF == ord('q'):
break
cap.release()
cv2.destroyAllWindows()Exam Tips for Unit II
- Always write the complete flow: open → check → loop → read → process → display → release
- Remember difference between
waitKey(1)(camera) andwaitKey(25)(video file) - Motion detection = Frame Differencing + Threshold + Contours
- Tracking needs initialization with ROI first
- FourCC is compulsory when writing videos
retvalue is very important (theory + coding questions)