Image Captions with Attention in Tensorflow, Step-by-step
An end-to-end example using Encoder-Decoder with Attention in Keras and Tensorflow 2.0, in Plain English

Now that the images are ready for training, we have to prepare the captions data next.
We wrap the training data in a Tensorflow Dataset object so that it can be efficiently fetched and fed, one batch at a time, to the model during training. The data is fetched lazily so that it doesn't all have to be in memory at the same time. This allows us to support very large datasets.
The dataset loads the pre-processed encoded image vectors that were saved earlier. It uses the image file name to identify the saved file path.
Much of the code for this example has been taken from the Tensorflow Image Caption tutorial.
There is a lot happening during training, and the flow of computations can get a little confusing. So let's go through them step by step.

Conclusion
Image Captioning is an interesting application because it combines techniques of Computer Vision and NLP, and requires working with both images and text.
We walked through an end-to-end example of Image Captions using the Encoder-Decoder architecture with Attention. We saw how Attention is used to boost the ability of the network to predict better captions. This is one of the better-performing architectures from recent years.
And finally, if you liked this article, you might also enjoy my other series on Transformers, Audio Deep Learning, and Geolocation Machine Learning.
Transformers Explained Visually (Part 1): Overview of Functionality
Audio Deep Learning Made Simple (Part 1): State-of-the-Art Techniques
Leveraging Geolocation Data for Machine Learning: Essential Techniques
Let's keep learning!








