MathJax 3

Showing posts with label color. Show all posts
Showing posts with label color. Show all posts

Saturday, 31 July 2021

Colour image histogram in Python

If you have missed the previous post on image channel decomposition, follow this link and then come back here. 

Colour image histogram

We can have two main types of colour image histogram: luminance histogram and component histogram. The principle used is the same as for the grey-level histogram in both the types, with one main difference between them:  the first type is obtained from a grey-scale version of the colour image, while the second type, from the image's three separate channels. 

Component histogram

To plot the component histogram of a colour image, after decomposing the image's channels into three separate images, each one containing only the colour for each channel, we must convert these images to grayscale. This step should account for the difference in human perception between red, green and blue, not just average the three channels' intensity, as explained here.

In the following picture, we can see the result of the above procedure executed on a 24 bits RGB image:

 

clockwise from top left:original image, red, blue and green channel

 

 

Some problems with the contrast are soon very evident, especially with the blue channel.

Once we have the arrays of the intensity values for each channel, we can apply the same steps to each channel's intensity array, as we would do for a grayscale image. That would produce the histogram for each one of the image's channels, as we can see in the next picture:

 

 
 
 
 
Interpreting the histogram
 
By analyzing the histogram above, it's evident that what it seems, after a quick look, a pretty and well-taken picture, in reality, it's instead a poor quality one and with a lot of room for improvements.
The red and the blue's dynamic range, ad example, it's too low. A little better for the greens, but still not enough intensity values. As a consequence, the contrast of the image lacks variance, and hence the colours look washed-up. 
However, we can use for colour images the same contrast improvement technique we learned here, like histogram equalization, which we will learn how to implement in the next post, together with other histogram operations.
 
 
The full source code for colour histogram:
 
from PIL import Image, ImageTk
import tkinter as tk
import sys

def rgb2gray(input_image):
   if
input_image.mode != 'RGB':
       return None
   else:
       output_image = Image.new('L', (input_image.width,
                                    input_image.height))

       for x in range(input_image.width):
           for y in range(input_image.height):
               pix = input_image.getpixel((x, y))
               pix = int(pix[0]*0.299 +
                         pix[1]*0.587 +
                         pix[2]*0.114)
               output_image.putpixel((x,y), pix)
       return output_image


def histogram(input_image):
   
if input_image.mode != 'L' and input_image.mode != 'P':
        return None
    else:
        IHIST = [0 for i in range(256)]
        HIST = [0 for i in range(256)]
        SUM = 0        
        for x in range(input_image.width):
            for y in range(input_image.height):
                pix = input_image.getpixel((x, y))
                IHIST[pix] = IHIST[pix]+1
                SUM += 1
        for i in range(256):
            HIST[i] = float(IHIST[i]/SUM)
            
        return HIST


def draw_histogram(canvas, IHIST, hist_w, hist_h, color, bins=False):
   hist_max = max(IHIST)## get the max value
   ## Normalize between 0 and hist_h
   for i in range(256):
      IHIST[i] = float(IHIST[i]/hist_max) * hist_h
   ## A bin is a bar of the histogram
   bin_w = round(float(hist_w/256))
   offset = hist_w + 20 ## where we draw the first
   prev_x = offset
   prev_y = hist_h
   for i in range(256):
      if bins:
         canvas.create_line(bin_w*i+offset, hist_h,
            bin_w*i+offset, hist_h-IHIST[i],fill=color)
      else:
         canvas.create_line(prev_x, prev_y,
            bin_w*i+offset, hist_h-IHIST[i],fill=color)
         prev_x = bin_w*i+offset
         prev_y =  hist_h-IHIST[i]

def get_channel(input_image, channel_key):
   if input_image.mode != 'RGB':
       return None
   else:
      channels_indices = {'R':0 ,'G':1 ,'B':2}
      output_image = Image.new('RGB', (input_image.width,
                                      input_image.height))
      ch_idx = channels_indices[channel_key]
      for x in range(input_image.width):
         for y in range(input_image.height):
            pix = input_image.getpixel((x, y))
            pix = pix[ch_idx]
            if ch_idx == 0:
               output_image.putpixel((x,y), (pix,0,0))
            elif ch_idx == 1:
               output_image.putpixel((x,y), (0,pix,0))
            elif ch_idx == 2:
               output_image.putpixel((x,y), (0,0,pix))
                  
       return output_image

if __name__ == "__main__":
   root = tk.Tk()
   img = Image.open("retriver.png")
   if img == None:
      print("Error opening image file"),
      sys.exit(1)

   width = img.width*2+20
   height = img.height+20
   root.title("RGB Image histogram")
   root.geometry(f'{width}x{height}')
   
   output_R = get_channel(img, 'R')  
   output_G = get_channel(img, 'G')
   output_B = get_channel(img, 'B')

   gray_R = rgb2gray(output_R)
   gray_G = rgb2gray(output_G)
   gray_B = rgb2gray(output_B)
   
   IHISTR = histogram(gray_R)
   IHISTG = histogram(gray_G)
   IHISTB = histogram(gray_B)
   
   hist_w = img.width
   hist_h = img.height
   
   if output_R != None and output_G!= None and output_B!=None:
      output_R = ImageTk.PhotoImage(output_R)
      output_G = ImageTk.PhotoImage(output_G)
      output_B = ImageTk.PhotoImage(output_B)
      input_im = ImageTk.PhotoImage(img)

      canvas = tk.Canvas(root, width=width, height=height, bg="#ffffff")
      canvas.create_image(width/4-1, height/2-1, image=input_im, state="normal")

      draw_histogram(canvas, IHISTR, hist_w, hist_h, "#f00")
      draw_histogram(canvas, IHISTG, hist_w, hist_h, "#0f0")
      draw_histogram(canvas, IHISTB, hist_w, hist_h, "#00f")
      
      canvas.place(x=0, y=0)
      canvas.pack()

      root.mainloop()
   else:
        print("Input image must be in 'RGB' colour space")
 
 
The code has been already explained in the previous posts. But, if something is not clear, ask for it in the comment section.

Friday, 30 July 2021

RGB channel decomposition - Python implementation

In the previous post, we have seen histogram equalization for grayscale images. Before learning about colour image histogram, we must know how to separate the channels of an image to be able to process them separately. We will learn how to do that in this post.
 
For this, we will use an RGB image with a bit depth of 8 bits per channel. I wrote a simple function which returns a copy of the input image containing only the data for a single channel. It's possible to select the required channel using a key ('R', 'G' or 'B') passed as an argument to the function: 

 

def get_channel(input_image, channel_key):
   if input_image.mode != 'RGB':
       return None
   else:
      channels_indices = {'R':0 ,'G':1 ,'B':2}
      output_image = Image.new('RGB', (input_image.width,
                                          
                    input_image.height))
      ch_idx = channels_indices[channel_key]
      for x in range(input_image.width):
         for y in range(input_image.height):
            pix = input_image.getpixel((x, y))
            pix = pix[ch_idx]
            if ch_idx == 0:
               output_image.putpixel((x,y), (pix,0,0))
            elif ch_idx == 1:
               output_image.putpixel((x,y), (0,pix,0))
            elif ch_idx == 2:
               output_image.putpixel((x,y), (0,0,pix))
                  
       return output_image

 

the function above takes an input image and a key as arguments. The key is used to choose the relevant channel's index for the pixel's tuple. Let's go through the most relevant lines:

channels_indices dict to map keys to indices,

channels_indices = {'R':0 ,'G':1 ,'B':2} 

ch_idx gets 0, 1 or 2 based on channel_key value

ch_idx = channels_indices[channel_key] 

first we get the pixel's tuple, then we get the channel's value from it:  

pix = input_image.getpixel((x, y))
pix = pix[ch_idx]

we also use ch_idx to select the output channel:

 if ch_idx == 0:
   output_image.putpixel((x,y), (pix,0,0))
 elif ch_idx == 1:
   output_image.putpixel((x,y), (0,pix,0))
 elif ch_idx == 2:
   output_image.putpixel((x,y), (0,0,pix))

In the following picture we can see the original image with the three channels separated:

 


 

and this is the full code:

 

from PIL import Image, ImageTk
import tkinter as tk


def get_channel(input_image, channel_key):
   if input_image.mode != 'RGB':
       return None
   else:
      channels_indices = {'R':0 ,'G':1 ,'B':2}
      output_image = Image.new('RGB', (input_image.width,
                                      input_image.height))
      ch_idx = channels_indices[channel_key]
      for x in range(input_image.width):
         for y in range(input_image.height):
            pix = input_image.getpixel((x, y))
            pix = pix[ch_idx]
            if ch_idx == 0:
               output_image.putpixel((x,y), (pix,0,0))
            elif ch_idx == 1:
               output_image.putpixel((x,y), (0,pix,0))
            elif ch_idx == 2:
               output_image.putpixel((x,y), (0,0,pix))
                  
       return output_image

if __name__ == "__main__":
   root = tk.Tk()
   img = Image.open("retriver.png")

   width = img.width*2+20
   height = img.height*2+20
   root.title("RGB channel decomposition")
   root.geometry(f'{width}x{height}')

   output_R = get_channel(img, 'R')  
   output_G = get_channel(img, 'G')
   output_B = get_channel(img, 'B')
     
   hist_w = img.width
   hist_h = img.height
   
   if output_R != None and output_G!= None and output_B!=None:
      output_R = ImageTk.PhotoImage(output_R)
      output_G = ImageTk.PhotoImage(output_G)
      output_B = ImageTk.PhotoImage(output_B)
      input_im = ImageTk.PhotoImage(img)

      canvas = tk.Canvas(root, width=width, height=height, bg="#ffffff")
      canvas.create_image(width/4-1, height/4-1, image=input_im, state="normal")
      canvas.create_image(3*width/4, height/4-1, image=output_R, state="normal")
      canvas.create_image(width/4-1, 3*height/4-1, image=output_G, state="normal")
      canvas.create_image(3*width/4, 3*height/4-1, image=output_B, state="normal")
      
      canvas.place(x=0, y=0)
      canvas.pack()

      root.mainloop()
   else:
        print("Input image must be in 'RGB' colour space")

 
 

Now we can learn how to plot an histogram for a colour image. Follow me in the next post to see how it is done and how it can help us.

Thursday, 15 July 2021

Image processing in Python

This is the first post of a series on image processing theory and practice. We will study some of the most important algorithms used in this field and we will learn how to implement them. I chose Python as the implementation language because this will make this series approachable for a large audience.

 

Python has become the language of choice for many scientists, students and engineers involved in machine vision and A.I. in general. Probably, the reasons for that are Python's easy syntax and Python's portability. But, without doubt, Python owns a good part of its success also to the large number of powerful and well-established libraries existing for the language, which allow users to perform nearly any kind of task with images and other types of data. Why bother to follow this tutorial, then? Firstly, because learning it's fun. Secondly, because if you are serious about your domain of expertise, you should have a good understanding of what a library (or whatever tool you use) is doing with your data, so to be able to choose the right functionality for your tasks, and in the right order.

 

What's image processing and why do we need it

 

A digital image is a signal, thus as such, it can be analysed and processed, to modify or to improve its properties. As with all signals, an image can have some noise in it that has to be removed, or maybe it needs to go through some transformation step before it can be used for some scope. Image processing is the study and the application of the mathematics needed to perform these improvements and transformations, utilizing computers.

 

A digital image can be represented as a bidimensional array of pixels, where each pixel is a sample of the image. Depending on the type of image, a pixel can have 1, 3 or 4 values associated with it. Such values are also called channels. Each channel holds a component of the image, like the level of one of the primary colours. The components of a pixel may vary a lot, depending on the encoding of the image (e.g. the file format) and the implementation algorithm. Also, the bit depth of an image can vary. The bit depth is the number of memory's bits used to hold the value associated with one pixel's channel.

 

When we consider all the components of a pixel and the range of values they can hold, we call these the image's colour space. Among all, the most common colour space used in digital images is probably the RGB, used also in PC's monitors and scanners. Printers use CMYK (cyan, magenta, yellow, and black). 

The pixel of an RGB image has usually 8 bits per channel, thus in such a case, we say that the image has a bit depth of 24 bits, for a total of 3 bytes of data per pixel. 

 

Another important type of image is the greyscale. This image is characterized by the absence of colours, except for the black, white and a range of shades of grey obtained by computing the amount of light carried by each pixel. If only black and white are present, we have a binary image, another important image type for machine vision applications. Grayscale images are lightweight images as they only need 1 byte per pixel (8 bits). In grayscale images, each pixel can be rendered with a value between 0 and 255, and that takes the same amount of memory's space as an ASCII character. These properties make this type of image more suitable for tasks where a lot of memory and computational power is needed. It's not for a case that in many algorithms, the first processing step on a colour image is to convert it from its colour space to grayscale.

 

Colour space conversion

 

As I mentioned above, sometimes we need to convert the colour space of our images to a model more suitable for our needs. We can convert any colour space to another, but the most important conversion it's, without doubt, the conversion to greyscale. The simplest method to achieve this with an RGB image is to average the value of the three channels:  

                                         

\[P = (R+G+B)/3\]

 

where P is the value of the pixel in the output image. Other algorithms are less naive, and take into consideration how we perceive the colours and their properties, as in the case of the gamma compression and the channel-dependent luminance algorithms. The former considers how luminance affects the way we perceive differences in colours, the latter how our eyes are more sensitive to green rather than blue or red. The implementation of gamma correction is expensive but gives a more accurate result, while the opposite is true for the channel-dependent luminance. Another algorithm do exist which provide a good balance between computational costs and luminance correctness, and that is the linear approximation, also used by some popular imaging libraries, like Opencv and Pillow: 

                                                

 \[P = 0.299R + 0.587G + 0.114B\]



Implementing RGB to Grayscale conversion

 

If you don't have Pillow (PIL fork) installed, you need it if you want to try the code in this tutorial. If you use pip as a package manager, just install it with the command pip install pillow. 

 

Pillow has quite a few ready-made filters implemented, but we will use it only because makes it very easy for us to access an image's data. We will also use Tkinter, from Python's standard library, as a display interface.

 

Let's import our libraries:

 

from PIL import Image, ImageTk

import tkinter as tk

 

follows our conversion function:

 

def rgb2gray(input_image):

   if input_image.mode != 'RGB':

       return None

   else:

       output_image = Image.new('L', (input_image.width,

                                    input_image.height))

       for x in range(input_image.width):

           for y in range(input_image.height):

               pix = input_image.getpixel((x, y))

               pix = int(pix[0]*0.299 +

                         pix[1]*0.587 +

                         pix[2]*0.114)

               output_image.putpixel((x,y), pix)

   

       return output_image

 

this line checks that we have an input image that is compatible with this function:

 

   if input_image.mode != 'RGB':

       return None 

 

if this is the case, we create a new image in 'L' mode, which means 8 bits per pixel. Then, we iterate over all the input image's pixels, retrieving the pixel's tuple for each position (x,y) visited:

 

       output_image = Image.new('L', (input_image.width,

                                    input_image.height))

       for x in range(input_image.width):

           for y in range(input_image.height):

               pix = input_image.getpixel((x, y)) 

 

we use the tuple's three values (r,g,b) to create the gray pixel, by applying linear approximation. Note that we reuse the pix variable. Next step, we set the output image's pixel:

 

               pix = int(int(pix[0]*0.299)+

                         int(pix[1]*0.587)+

                         int(pix[2]*0.114))

               output_image.putpixel((x,y), pix) 

 

finally, the driver's code:

 

 if __name__ == "__main__":


   root = tk.Tk()

   img = Image.open("retriver.png")


   width = img.width*2+20

   height = img.height

   root.geometry(f'{width}x{height}')

    

   output_im = rgb2gray(img)

   if output_im != None: 

       output_im = ImageTk.PhotoImage(output_im)

       input_im = ImageTk.PhotoImage(img)

    

       canvas = tk.Canvas(root, width=width, height=height, bg="#ffffff")

       canvas.create_image(width/4-1, height/2-1, image=input_im, state="normal")

       canvas.create_image(20+3*img.width/2-1, img.height/2-1, image=output_im, state="normal")

       canvas.place(x=0, y=0)

       canvas.pack()

       

       root.mainloop()

   else:

       print("Input image must be in 'RGB' colour space")

 

below you can see the result after applying the rgb2gray function to the input image:

 

 

The complete code:

from PIL import Image, ImageTk
import tkinter as tk
 
def rgb2gray(input_image):
    if input_image.mode != 'RGB':
        return None
    else:
        output_image = Image.new('L', (input_image.width,
                                     input_image.height))
        for x in range(input_image.width):
            for y in range(input_image.height):
                pix = input_image.getpixel((x, y))
                pix = int(int(pix[0]*0.299)+
                          int(pix[1]*0.587)+
                          int(pix[2]*0.114))
                output_image.putpixel((x,y), pix)
    
        return output_image

if __name__ == "__main__":

    root = tk.Tk()
    img = Image.open("retriver.png")

    width = img.width*2+20
    height = img.height
    root.geometry(f'{width}x{height}')
     
    output_im = rgb2gray(img)
    if output_im != None:
        output_im = ImageTk.PhotoImage(output_im)
        input_im = ImageTk.PhotoImage(img)
     
        canvas = tk.Canvas(root, width=width, height=height, bg="#ffffff")
        canvas.create_image(width/4-1, height/2-1, image=input_im, state="normal")
        canvas.create_image(20+3*img.width/2-1, img.height/2-1, image=output_im, state="normal")
        canvas.place(x=0, y=0)
        canvas.pack()
        
        root.mainloop()
    else:
        print("Input image must be in 'RGB' colour space")


Final words

 

In this first post, we have seen what image processing is and how to implement a simple RGB to the greyscale converter. In the next post, we will study and implement another algorithm to deepen our knowledge about image processing and machine vision.


Thursday, 10 September 2015

Image processing in C: enbossing filter for PPM images

This is the last post of the series about image manipulation for C/C++ beginners. We will see how to create a simple filter to process an image using C. The filter we will implement is named 'embossing filter'. We can find this type of filter image in manipulation programs and image packages like GIMP, Photoshop and ImageMagick.


The embossing filter implementation


 We will write a function that takes a pointer to an image as an argument and returns another pointer to a filtered image. As a start, the function calculates the difference between each pixel channel's value and the value of the channels of its upper left neighbour. Then saves the highest absolute difference, adding to it 128. If there are multiple equal differences with different signs (e.g. -3 and 3), the function will favour the red channel's difference first, then the greens, then the blues. A value above 255 and below 0 will be set to 255 or 0, respectively. 


 The following code is the C implementation of the embossing filter:



Image * emboss_image(Image * img)
{

    //get size of image
    unsigned int w = getImgWidth(img);
    unsigned int h = getImgHeight(img);




    //create a new image to return
    Image * emb_img = new_image(w, h);
    

    Pixel diff;
    Pixel upleft;
    Pixel curr;
    char maxDiff, tmp = 0;
    unsigned char v;

 
    //for each pixel
    int x, y;
    for(y=1; y<h; y++)
    {
        for(x=1; x<w; x++)
        {
                
                 if(maxDiff > tmp)
                      tmp = maxDiff;
                 else
                      maxDiff = tmp;
                
                 //get upper-left pixel value
                 upleft = getPixel(img, x-1,y-1);

                 //get current pixel value
                 curr = getPixel(img, x, y);

                 //find difference between channels
                 diff.r = curr.r - upleft.r;
                 diff.g = curr.g - upleft.g;
                 diff.b = curr.b - upleft.b;


                 //for equal values choose in favor of red
                 if((abs(diff.r)==diff.g && diff.g > diff.b)||(abs(diff.r)==diff.b && diff.b > diff.g))
                      maxDiff = diff.r;


                 //or find the maximum value
                 else
                      maxDiff = max(diff.r,max(diff.g,diff.b));
                 

                 //add 128 to the maximum difference
                 v = 128 + maxDiff;

                 //remove values above 255 or below 0
                 if(v<0)v=0;
                 if(v>255)v=255;


                 //create and set the pixel in emb_img
                 Pixel val2 = {v,v,v};
                 setPixel(x,y,&emb_img,val2);
        }
    }
    return emb_img;
}




Loading a PPM image in memory 


Before we can process an image, we require the capacity to load and retrieve this from our computer's memory. Luckily for us, achieving this with a PPM image, it's simple. Simple as reading a text file. The following C function reads a PPM image's data and stores it inside an Image structure. Then, returns a pointer to this structure:




Image * read_PPM(const char *filename)
{
     char buff[16];
     Image * img;
     int ch;
     int rgb;
   
     //open and read image file
     FILE * fp = fopen(filename, "rb");
     if (!fp) {
          fprintf(stderr, "Unable to open file '%s'\n", filename);
          exit(1);
     }

     //read image format
     if (!fgets(buff, sizeof(buff), fp)) {
          perror(filename);
          fclose(fp);
          exit(1);
     }

    //parse magic number
    if (buff[0] != 'P' || buff[1] != '6') {
         fprintf(stderr, "Invalid image format (must be 'P6')\n");
         fclose(fp);
         exit(1);
    }

    //parse comment
    ch = getc(fp);
    while (ch == '#') {
        while (getc(fp) != '\n') ;
             ch = getc(fp);
    }
    ungetc(ch, fp);
  
    int width, height;
    //parse width and height
    if (fscanf(fp, "%d %d", & width, & height) != 2) {
         fprintf(stderr, "Invalid image size (error loading '%s')\n", filename);
         exit(1);
    }

    //parse max rgb
    if (fscanf(fp, "%d", &rgb) != 1) {
         fprintf(stderr, "Invalid rgb component (error loading '%s')\n", filename);
         fclose(fp);
         exit(1);
    }

    //check max rgb value correctness
    if (rgb !=  RGB) {
         fprintf(stderr, "'%s' is not a 24-bits image.\n", filename);
         fclose(fp);
         exit(1);
    }

   
    //create new image
    img = new_image(width, height);
  
    int bytes;
    //read image data from file and store it inside pixel array
    if ((bytes=fread(img->data, 3 * img->width, img->height, fp)) != img->height) {
         fprintf(stderr, "Error loading image '%s'\n", filename);
         exit(1);
    }
    
    fclose(fp);
    return img;
}


By reading the comment in the code, it should be easy to understand what is happening. We start by opening the PPM image which we want to process. Once we did that, we need to parse the header, which will provide us with all the information's we need to process the image, as the size and the maximum value for the intensity of the colour. Also, we check for comments, as they can be present in a PPM image file, after the magic number. After we finished reading the header, we will create a new image and will store in it the data from the file. At this point, the file can be closed and we can return a pointer to our Image structure. 

If you have read the emboss function, you may have noticed that we call a function called getPixel. With this function we retrieve the RGB value of each pixel stored into the Image structure pointed to by the Image pointer argument :




  
  Pixel  getPixel(Image * i, int x, int y){ 
    return i->data[getImgWidth(i)*y + x]; 
  }


As you can see, the function returns an element of the pixel array of the image passed as an argument. The second and the third parameters represent the coordinates of the pixel in the image plane. If the index used in the array looks weird for you, just remember that if we store the pixels of an image in a 1d array, by using the x y image plane pixel's coordinate, we can access a pixel in the 1d array by indexing this with the integer resulting from the following calculation:

                            
                          pixel position = image_width * y + x   


from this page you can copy or download the complete code for this post. Below you can see the image before and after applying the emboss filter:



                     before                      after

 Conclusions


The image formats implemented in the Netpbm package, and in particular, the PPM format, represent a great and easy media to work with, especially for beginners who wish to start their journey to learning the art of computer graphics, to whom this post series was dedicated.