Image ProcessingUnit 111 min read
Digital Image Processing Basics: Definitions, Steps & Applications
Unit 1 of Image Processing covers the foundational concepts of digital image processing, including its definition, key steps, spatial vs. frequency domains, and real-world applications in apps like eSewa and Ncell. This note explains how images are processed digitally, the role of pixels and bit planes, and how segment
Core Concepts: What is Digital Image Processing?
Digital Image Processing (DIP) is the use of computer algorithms to perform operations on digital images. These operations can enhance, restore, compress, or analyze images to extract meaningful information.
Why Process Images Digitally?
- Improve image quality (e.g., noise reduction in medical scans).
- Extract features (e.g., face recognition in smartphones).
- Compress images (e.g., JPEG in social media).
- Automate tasks (e.g., license plate detection in traffic cameras).
Key Definitions
| Term | Definition |
|---|---|
| Digital Image | A 2D array of pixels, each with an intensity value (e.g., 0–255 for grayscale). |
| Pixel | Smallest addressable element in an image (e.g., a single dot in a 1024×768 image). |
| Bit Plane | A binary layer representing a single bit of all pixels (e.g., the 4th bit plane of an 8-bit image). |
| Spatial Domain | Direct manipulation of pixel values (e.g., brightness adjustment). |
| Frequency Domain | Analysis using transforms (e.g., Fourier Transform) to study image textures and patterns. |
How Images Are Represented
1. Spatial Domain Representation
Example (3×3 grayscale image):
[ 50 60 70 ]
[ 80 90 100 ]
[110 120 130 ]
- Intensity Level: The brightness value (e.g., 0 = black, 255 = white in 8-bit images).
- Bit Plane Slicing: Separating an image into binary layers (e.g., an 8-bit image has 8 bit planes).
2. Frequency Domain Representation
- Low Frequencies: Represent smooth regions (e.g., sky in a landscape).
- High Frequencies: Represent edges and fine details (e.g., textures).
Example (1D DFT of a signal):
For a signal [1, 2, 3, 4], the DFT shows how much of each frequency (sine/cosine wave) exists in the signal.
(We’ll compute 2D DFT later in Unit 9.)
Steps in Digital Image Processing
Every DIP task follows a pipeline (sequence of operations). Here’s a typical workflow:
flowchart LR
A["Acquisition"] --> B["Preprocessing"]
B --> C["Enhancement"]
C --> D["Segmentation"]
D --> E["Analysis/Description"]
E --> F["Recognition"]
F --> G["Output"]Detailed Steps:
- Acquisition: Capturing the image (e.g., camera sensor, scanner).
- Preprocessing: Correcting distortions (e.g., noise removal, geometric corrections).
- Enhancement: Improving visual quality (e.g., contrast stretching, sharpening).
- Segmentation: Dividing the image into regions (e.g., separating a car from the background).
- Analysis/Description: Extracting features (e.g., edges, textures).
- Recognition: Classifying objects (e.g., "Is this a cat or a dog?").
- Output: Final result (e.g., a compressed image, a detected face).
In the Real World
1. eSewa & Kathmandu Traffic Management
- Application: eSewa uses image segmentation to verify signatures on digital payments.
- How? The app detects handwritten signatures by:
- Converting the signature image to grayscale.
- Applying thresholding to separate ink from background.
- Using edge detection (e.g., Sobel or Prewitt operators) to extract signature contours.
- Real Example: When you sign a payment request on eSewa, the system checks if the signature matches the stored template using pattern recognition.
2. Ncell’s Face Unlock Feature
- Application: Ncell’s face unlock uses frequency domain analysis and feature extraction.
- How?
- The camera captures an image in the spatial domain.
- The phone applies a 2D DFT to analyze facial features in the frequency domain.
- Eigenfaces (a technique from Principal Component Analysis) are used to compare the live face with stored templates.
- Real Example: When you unlock your phone with your face, the algorithm checks high-frequency details (like eye spacing) to ensure it’s you.
3. Daraz’s Product Image Compression
- Application: Daraz compresses product images to reduce storage and speed up loading.
- How?
- Uses Discrete Cosine Transform (DCT) (like JPEG compression).
- High-frequency components (fine details) are quantized (reduced in precision).
- Low-frequency components (smooth areas) are kept intact.
- Real Example: When you scroll through Daraz’s product gallery, the images load quickly because they’re stored in a compressed format.
Worked Example: Histogram Equalization (Spatial Domain)
Problem: Given a grayscale image with the following histogram:
| Gray Level | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|
| Frequency | 0 | 25 | 0 | 10 | 35 | 50 | 0 | 0 |
Goal: Stretch the histogram to use the full dynamic range (0–7) using histogram equalization.
Step-by-Step Solution:
Calculate Total Pixels: pixels.
Compute Cumulative Distribution Function (CDF): The CDF for gray level is: Where is the frequency of gray level .
Gray Level CDF 0 0 1 25 2 25 3 35 4 70 5 120 6 120 7 120 Normalize CDF to [0, 7]: Where (number of gray levels).
Gray Level CDF Normalized 0 0 0 1 25 → 1 2 25 1.46 → 1 3 35 → 2 4 70 → 4 5 120 7 6 120 7 7 120 7 Map Original to New Gray Levels:
- Original gray level 1 → New level 1
- Original gray level 3 → New level 2
- Original gray level 4 → New level 4
- Original gray level 5 → New level 7
Result: The enhanced image will have pixels spread across the full range (0–7), improving contrast.
Image Segmentation: Detecting Objects
Definition:
Types of Grey-Level Discontinuities (Edges)
Edges are detected using discontinuities in pixel intensity. Three main types:
| Type | Description | Example |
|---|---|---|
| Roof Edge | Sudden change in intensity (e.g., a white wall next to a dark shadow). | |
| Step Edge | Abrupt transition (e.g., a black line on a white background). | |
| Ramp Edge | Gradual change over multiple pixels (e.g., a shadow fading into light). |
Worked Example: Edge Detection with Prewitt Operator
Problem: Given the following 3×3 image, detect edges using the Prewitt operator (a gradient-based method).
| 0 | 30 | 60 |
|---|---|---|
| 5 | 32 | 62 |
| 10 | 38 | 64 |
Prewitt Operator Kernels:
- Horizontal Edge Detection (Gx):
[-1, 0, +1] [-1, 0, +1] [-1, 0, +1] - Vertical Edge Detection (Gy):
[-1, -1, -1] [ 0, 0, 0] [+1, +1, +1]
Step-by-Step Calculation:
Compute Gx (Horizontal Gradient): For the center pixel (32): (Note: For simplicity, we’ll compute for one pixel. In practice, you’d compute for all non-border pixels.)
Compute Gy (Vertical Gradient): For the center pixel (32):
Compute Gradient Magnitude and Direction: (This indicates a strong horizontal edge.)
Result: The center pixel has a strong horizontal edge. In practice, you’d threshold the magnitude to detect actual edges.
Spatial Domain vs. Frequency Domain: Key Differences
| Feature | Spatial Domain | Frequency Domain |
|---|---|---|
| Representation | Direct pixel manipulation. | Uses transforms (e.g., DFT, DCT). |
| Operations | Filtering, thresholding, histogram equalization. | Noise filtering, compression (JPEG). |
| Advantage | Simple, intuitive. | Better for global operations (e.g., removing periodic noise). |
| Disadvantage | Computationally expensive for large images. | Requires transforms (slower for real-time). |
| Example | Sharpening an image with a Laplacian filter. | Removing high-frequency noise using a low-pass filter. |
Exam Tip
What to Focus On:
- Definitions:
- Know the difference between spatial and frequency domains.
- Understand histogram equalization, edge detection, and image segmentation.
- Calculations:
- Practice histogram stretching and equalization (common in exams).
- Be able to compute gradient-based edge detection (Prewitt, Sobel).
- Real-World Applications:
- Link concepts to eSewa (segmentation), Ncell (face recognition), and Daraz (compression).
- Common Pitfalls:
- Forgetting to normalize CDF in histogram equalization.
- Misapplying kernels in edge detection (e.g., using Prewitt instead of Sobel).
- Ignoring border pixels in convolution operations.
Past Exam Patterns:
- Short Questions (5–10 marks): Define terms like "noise," "linear filter," or "pattern class."
- Numerical Problems (15–20 marks): Histogram equalization, edge detection, or DFT calculations.
- Theoretical Questions (10–15 marks): Explain steps in DIP, compare spatial vs. frequency domains.
Final Checklist Before Exam:
✅ Can you draw a flowchart of the DIP pipeline? ✅ Can you compute a histogram equalization from scratch? ✅ Can you detect edges using Prewitt/Sobel operators? ✅ Can you explain how eSewa or Ncell uses DIP in real life? ✅ Do you know the difference between spatial and frequency domains?
Based on the TU BSc CSIT syllabus for Image Processing (CSC332), unit 1.
Discussion
Loading…