Hyperparameter Analysis for Image Captioning

Amish Patel,Aravind Varier
DOI: https://doi.org/10.48550/arXiv.2006.10923
2020-06-19
Abstract:In this paper, we perform a thorough sensitivity analysis on state-of-the-art image captioning approaches using two different architectures: CNN+LSTM and CNN+Transformer. Experiments were carried out using the Flickr8k dataset. The biggest takeaway from the experiments is that fine-tuning the CNN encoder outperforms the baseline and all other experiments carried out for both architectures.
Computer Vision and Pattern Recognition,Machine Learning
What problem does this paper attempt to address?