admin@publications.scrs.in   
Data Science and Intelligent Computing Techniques

Image Captioning using CNN and Attention Based Transformer

Authors: Deepa Mulimani, Prakashgoud Patil and Nagaraj Chaklabbi


Publishing Date: 13-01-2023

ISBN: 978-81-955020-2-8

DOI: https://doi.org/10.56155/978-81-955020-2-8-14

Abstract

Image captioning is a technique for generating sentences that describe a scenario captured in photos. It can identify objects in a picture and carries out a few processes with the goal of locating the image’s most crucial parts. Algorithms now have the ability to generate text in the context of natural phrases that accurately describe an image. To extract image visual features, this work employs a pre-trained Convolution Neural Network (CNN) viz. EfficientNetB0, and then uses Transformer Encoder and Decoder to construct an appropriate caption. The model is trained using the Flickr8k dataset. The findings back up the model’s capacity to understand and produce text from pictures. The evaluation metric is the BLEU (bilingual evaluation understudy) score. The model obtains the image description, converts into text, and then into a voice. For visually impaired people who are unable to grasp visuals, image description is the ideal approach.

Keywords

Image caption generator, CNN, Transformer, Attention, Decoder, Encoder, Efficient-Net

Cite as

Deepa Mulimani, Prakashgoud Patil and Nagaraj Chaklabbi, "Image Captioning using CNN and Attention Based Transformer", In: Satyasai Jagannath Nanda and Rajendra Prasad Yadav (eds), Data Science and Intelligent Computing Techniques, SCRS, India, 2023, pp. 157-166. https://doi.org/10.56155/978-81-955020-2-8-14

Recent