DETECTION AND RECOGNITION OF TEXTS FEATURES FROM A TOPOGRAPHIC MAP USING DEEP LEARNING
Keywords:
Digitization, topographic map, faster R-CNN, convolutional recurrent neural network, text detection, text recognitionAbstract
The aim of this research is to digitize texts present on topographic maps. Digitization is critical for information processing, hence texts on topographic maps that are in pixel format are digitized. The above is achieved using deep learning technology. Detection and recognition deep learning models are leveraged for first detecting texts on topographic maps and then recognition the detected texts. Detection model is inspired by Faster Region Based Convolutional Neural Networks (Faster R-CNN) to achieve real time text detection using region proposal networks and recognition model is inspired by Convolutional Recurrent Neural Network (CRNN), which is an end to end trainable neural network to achieve image-based sequence recognition that is suitable for recognizing texts on character level. The models have been trained and validated on self-curated dataset of topographic maps. The digitized information may be very helpful for many GIS applications like automated construction of Digital Elevation Model, Digital Surface Model, etc.
References
Bengio, Y., Simard, P., and Frasconi, P. (1994). Learning long-term dependencies with gradient descent is difficult. IEEE transactions on neural networks, 5:157-166.
Bissacco, A., Cummins, M., Netzer, Y., and Neven, H. (2013). Photoocr: Reading text in uncontrolled conditions. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV); 1-8 December, IEEE Xplore, Sydney, NSW, Australia, p. 785-792.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In: IEEE Conference on Computer Vision and Pattern Recognition; 20-25 June 2009, IEEE, Miami, FL, USA, p. 248-255.
Fu, B., Zhao, X.Y., Li, X., and Ren, Y.G. (2018). A convolutional neural networks denoising approach for salt and pepper noise. Multimedia Tools Appl., 78(21):30,707-30,721.
Gers, F., Schraudolph, N., and Schmidhuber, J. (2002). Learning precise timing with lstm recurrent networks. J. Mach. Learn. Res., 3:115-143.
Girshick, R. (2015). Fast R-CNN. 2015 IEEE International Conference on Computer Vision (ICCV); 7-13 December 2015, IEEE Xplore, Santiago, Chile. p. 1,440-1,448.
Girshick, R., Donahue, J., Darrell, T. and Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. 2014 IEEE Conference on Computer Vision and Pattern Recognition; 23-28 June 2014, IEEE Xplore, Columbus, USA, p. 580-587.
Graves, A., Fernandez, S., and Gomez, F. (2006). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. Proceedings of the International Conference on Machine Learning; 25-29 June 2006, Association for Computing Machinery (ACM), New York, United States, p. 369-376.
Graves, A., Liwicki, M., Fernandez, S., Bertolami, R., Bunke, H., and Schmidhuber, J. (2009). A novel connectionist system for unconstrained handwriting recognition. IEEE Trans. Pattern Anal. Mach. Intell., 31(5):855-868.
Graves, A., Mohamed, A.R., and Hinton, G. (2013). Speech recognition with deep recurrent neural networks. 2013 IEEE International Conference on Acoustics, Speech and Signal Processing; 26-31 May 2013, IEEE, Vancouver, BC, Canada, p. 6,645-6,649.
He, K., Zhang, X., Ren, S., and Sun, J. (2015). Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 27-30 June 2016, IEEE, Las Vegas, NV, USA, p. 770-778.
Hochreiter, S. and Schmidhuber, J. (1997). Long short-term memory. Neural Comp., 9(8):1,735-1,780.
Huang, J., Rathod, V., Sun, C., Zhu, M., Korattikara, A., Fathi, A., Fischer, I., Wojna, Z., Song, Y., Guadarrama, S., and Murphy, K. (2017). Speed/accuracy trade-offs for modern convolutional object detectors. 2017 IEEE Conference on Computer Vision and Pattern Recognition; 21-26 July 2017, IEEE Computer Society, Honolulu, USA, p. 3,296-3,297.
Jaderberg, M., Simonyan, K., Vedaldi, A., and Zisserman, A. (2016). Reading text in the wild with convolutional neural networks. Int. J. Comp. Vis., 116(1):1-20.
Krizhevsky, A., Sutskever, I. and Hinton, G.E. (2012). Imagenet classification with deep convolutional neural networks. 25th International Conference on Neural Information Processing Systems; 3-6 December, 2012 ACM Digital Library, Lake Tahoe, Nevada, p. 1,097–1,105.
Liao, M., Shi, B. and Bai, X. (2018). Textboxes++: A single-shot oriented scene text detector. IEEE Transactions on Image Processing, vol. 27(8):3,676-3,690.
Lin, T.Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017). Feature pyramid networks for object detection. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 21-26 July 2017, IEEE, Honolulu, USA, p. 936-944.
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Scott R., Fu, C.Y., and Berg, A.C. (2016). SSD: Single shot multibox detector. European Conference on Computer Vision: Lecture Notes in Computer Science, 11-14 October 2016, Springer, Amsterdam, The Netherlands, p. 21-37.
Olson, M., Wyner, A., and Berk, R. (2018). Modern neural networks generalize on small data sets. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc.
Redmon, J. and Farhadi A. (2017). Yolo9000: better, faster, stronger. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 21-26 July 2017, IEEE Xplore, Honolulu, USA, p. 6,517-6,525.
Redmon, J. and Farhadi, A. (2018). Yolov3: An incremental improvement, Technical Report on Computer Vision and Pattern Recognition, Cornell University, arxiv: 1804.02767.
Redmon, J., Divvala, S., Girshick, R. and Farhadi, A. (2016). You only look once: Unified, real-time object detection. IEEE Conference on Computer Vision and Pattern Recognition; 27-30 June 2016, IEEE Xplore, Las Vegas, USA, p. 779-788.
Ren, S., He, K., Girshick, B.R., and Sun, J. (2016). Faster r-cnn: Towards real-time object detection with region proposal networks. Technical Report on Computer Vision and Pattern Recognition, Cornell University, earXiv: 1506.01497.
Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R. and LeCun, Y. (2014). Overfeat: Integrated recognition, localization and detection using convolutional networks. Technical Report on Computer Vision and Pattern Recognition, Cornell University, arXiv:1312.6229.
Shi, B., Bai X., and Yao C. (2015). An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. Technical Report on Computer Vision and Pattern Recognition, Cornell University, arXiv:1507.05717.
Su, B. and Lu, S. (2014). Accurate scene text recognition based on recurrent neural network. ACCV 2014: Computer Vision - ACCV 2014; 1-5 November 2014, Springer, Singapore, p 35-48.
Szegedy, C., Ioffe, S., Vanhoucke, V. and Alemi, A. (2016). Inception-v4, inception-resnet and the impact of residual connections on learning. AAAI Conference on Artificial Intelligence (AAAI-16); 12-17 February 2016, Arizona, USA.
Uijlings, J.R.R., van de Sande, K.E.A., Gevers, T., and Smeulders, A.W.M. (2013). Selective search for object recognition. Int. J. Comp. Vis., 104(2):154-171.
Wang, T., Wu, D. J., Coates, A., and Ng., A.Y. (2012). End-to-end text recognition with convolutional neural networks. In: Proceedings of the 21st International Conference on Pattern Recognition (ICPR2012); 11-15 November 2012, Tsukuba Science City, Japan, p. 3,304-3,308.








