Placeholder cover image for the example project

Traffic Sign Classification with CNN

A CNN based solution to classify traffic images across 43 different categories.

CNN, Computer Vision, Tensorflow, OpenCV, Python

Problem

I was at Deep Learning, and i thought why not come up with something that can runs on edge computing, and helps autonomous vehicles detect traffic signal, since that’s one of the most important parts of driving, So i started making a Neural Network for this very reason.

Approach

Initially, i was following a 4-5 layers CNN architecutre as that sort of resembles VGGNet[but was not that], but i didn’t give it much of an importance.

Then i started reading some research papers so that i follow more structured approach to this : like VGGNet, LeNet, and GoogLeNet Inception, then i came across their major findings.

VGGNet : This paper was explaining how you could improve your model’s performance by increasing more CNN Layers and making it more deep.

LeNet : This paper was explaining how 1x1 convolutional bottlenecks are really good for performance & more fuller representations.

GoogLeNet : This paper was sort of Competing with VGGNet, saying that Deeper the NN goes, the more computation-heavy it becomes. It very extensively used the findings from LeNet paper, owing to its performance benifits. The paper introduced parallel representations of the same Feature Map/Image, representations from different perspectives : sort of looking at finer details and looking at larger details at the same time.

So i just took their findings as my railings and started building one of my own, and 60% of the times it was just experimentation, what works, what doesn’t, and reasoning behind it. I had to do a lot of Tensorboard logs just to see what wents which way.

I Started Building the project in Step by Step manner :
First i took GTSRB Dataset, Downloaded it. Then as usual, I made a logic to read the repository of dataset, and load it in my device.

Then i moved to preprocessing and augmentation part, since the dataset was highly skewed - there were 7-8 categories having as low as 100-200 images and i didn’t recommend myself augmentation, since either that would make some categories tolerant to variations like rotations, random crop, lighting up and down, while other categories would remain same OR the whole model recieves less quality - less diverse dataset. So i just did preprocessing of images like normalizing pixel values, or some images were over-exposed, i reduced their alpha and vice-versa. And the dataset didn’t require much preprocessing.

Then about balancing dataset : i used scikit-learn weight calculation, so basically you calculate the ratio wc = N/(Nc*k).
where
wc = weight of class c
N = total number of training examples
Nc = Number of examples in category c
k = number of categories
so that classes with less examples are weighed more, consequently its almost the same as increasing learning rate of that particular category by the same ratio, so since you have less training examples, so you’d take longer jumps but in a constrained manner limited by the weight.

Then came building the architecture, during this phase I did changes to batch sizes, learning rates, dataset augmentations, Train-Eval-Test sets partitioning, number of layers, convolutional filters, wider[not deeper] architecture and more. so i verified the findings of those researchers through my own experimentation.

Then I went on to getting the results.

Result

MY KEY FINDINGS THROUGH EXPERIMENTATION :
1. Too Large Batch size worsens learning
2. 1x1 convolutions increase performance[reduced time/step], but upto a certain limit only.
3. Through Dataset augmentation, NN can be made invariant to many things like rotate image, colors, lightings, or many things.
4. You increase more filters in the middle layers, so that we get more abstract representations and classification becomes better.
5. Keep learning Rate low [but not too much] to keep it easy going [so that it doesn’t wobble in the end].
6. Increasing Image Dimensions drastically increases computational requirements, i was at 80x80 and the time/step was around 100ms [or so, i don’t remember].

HERE ARE SOME RESULTING METRICS :

F1 Score

F1 Score

Top-5 Score

Top-5 Score

Precision Score

Precision Score

Recall Score

Recall Score

Accuracy

Accuracy

Note:
1. I didn’t give much importance to accuracy since the dataset was imbalanced [although i was doing weighing of different categories].
2. The model isn’t converged because this project was meant only for learning, not Deploying, with better dataset, i’d love to converge it.

INSIGHT :
Since we have good precision and less recall, it means if model predicts an image to have a category C, then its highly probable that it belongs to C. But it can’t correctly predict all images of the category C in the same category. So the model is like a sharpshooter with short-term-memory loss haha :)