岩月ほか: ′′Listen and Tell: 深層学習を用いた音響シーンのキャプション生成′′, 情報処理学会第81回全国大会, 2019
Lin et al.: ′′Microsoft coco: Common objects in context′′, European conference on computer vision, 740-755, 2014
Plummer et al.: ′′Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models′′, Proceedings of the IEEE international conference on computer vision, 2641-2649,2015
Vinyals et al.: ′′Show and tell: A neural image caption generator′′, Proceedings of the IEEE conference on computer vision and pattern recognition, 3156-3164,2015