AUTOMATIC SYNCHRONIZATION OF AUDIO-VISUAL CONTENTS WITH SUBTITLES
Keywords:
DTW, LCS, PLCS, CGF, Needleman-Wunch, MFCC, HMMAbstract
This paper presents an automated Kannada subtitle generator from Kannada video which is implemented to assist people with auditory problems for watching videos. Subtitle generation has become an important task for supporting such special people and it integrates an audio extraction and a speech recognition module and synchronization. The proposed work involves three phase such as extraction of audio form video, recognition of speech, generating subtitle and synchronization of subtitle with audio and video. An adaptive speech recognition module is implemented AMFCC for feature extraction. For decoder acoustic module, AHMMis used to reduce the computational time and memory usage. The text file from the speech recognition module is rendered to synchronize the missing offset with the video using parallel processing. One amongst the foremost essential factors with audience perception is the quality of the video subtitle for those people who are deaf and difficult for hearing. In this paper we are discussing the technique for Sub-Synchronization tool for automatic synchronization between media substance and subtitles, by taking an advantage of current familiar techniques which are utilized for word sequences alignment. The main objective of this paper is to deal the lack of synchronizing that happens within the subtitles once made during the video that have some improvised parts. In order to synchronize the audio-Video content with subtitle an efficient technique is used for sequence alignment is PLCS algorithm which gives best performance in terms of time and space, accuracy compare to other existing technique for sequence alignment algorithm.
References
Anagnostopoulos CN., Iliou T., and Giannoukos. I. (2015). Features and classifiers for emotion recognition from speech: a survey from 2000 to 2011. Artificial Intelligence Review Springer., 43:155-77.
Aparício, M., Figueiredo, P., Raposo, F., de Matos, DM., Ribeiro, R., and Marujo L. (2106). Summarization of films and documentaries based on subtitles and scripts. Pattern Recognition Letters Elsevier., (73):7-12.
Araújo, de., TM., Ferreira FL., dos Santos Silva DA., Lemos FH., Neto GP., Omaia D., de Souza Filho GL., and Tavares TA (2013). Automatic generation of Brazilian sign language windows for digital TV systems. J. of the Brazilian Computer Society. Springer., 107-25.
Biswas., Astik, PrakashKumarSahu., and Mahesh Chandra. (2015). Multiple cameras in-car audiovisual speech recognition using phonetic and viremic information. Com. and Electrical Eng., (47):35-50.
Cuzco-Calleet, I., Ingavélez, P., Robles-Bykbaev, V., and Calle-López. D. (2018). An interactive system to automatically generate video summaries andperform subtitles synchronization for persons with hearing loss. Proc. IEEE 20th Int. Conf. Electron. Electr. Eng. Compute., 1-4.
Deventer van, M.O., Stokking, H., Hammond, M., Le Feuvre, J., and Cesar., P. (2016) Standards for multi-stream and multi-device media synchronization IEEECommun. Mag., 54(3):16-21.
Disabil. Soc. (2011), The world report on disability. Disabil.Soc., 26(5):655-658.
Gao, J., Zhao, Q., Li, T., and Y, Yan. (2009). Simultaneous synchronization of textand speech for broadcast news subtitling. Advances in Neural Net-works (Lecture Notes in Computer Science)., 5,553:576-585.
Garcia, J.E., Ortega, A., Lleida, E., Lozano, T., Bernues. E.D. (2009). Sanchez Audio and text synchronization for TV news subtitling based on auto-matic speech recognition. In: Proc IEEE Int. Symp. Broadband Multime-dia Syst, Broadcast, p. 1-6
Health Law, Eur. J. (2006) Convention on the rights of persons with disabilities. vol. 14, no. 3, p. 98-281
Hirschberg, D.S. (1975) A linear space algorithm for computing maximal common subsequences. Commun. ACM., 18(6):341-343.
Huang, C., Hsu, W., and Chang, S. (2003). Automatic closed caption alignmentbased on speech recognition transcripts. In: Columbia Univ. NY, USA, Tech. Rep, p 007.
Karpukhin, I.A. and Konushin, A.S., (2017). Constructing a speech audio-videocorpus by aligning long segments of speech and text MoscowUniv.Comput. Math.Cybern., 41(2):97-103.
Kedačićet, D., Herceg, M., Peković, V., and Mihic, V. (2018). Application for testingof video and subtitle synchronization. Proc. Int. Conf. Smart Syst.Technol., 23-27.
Lert wong khanakool, N., Punyabukkana, P., and Suchato, A. (2013). Real-timesynchronization of live speech with its transcription. Proc. 10thInt. Conf. Elect. Eng/Electron, Comput, Telecommun. Inf. Technol., 1-5.
Ley. (2010) General De La ComunicaciónAudiovisual, BOE, Madrid, Spain, p. 2,010
Loni, Deepali Shaila and Shaila Subbaraman. (2015). Singing voice identification using a harmonic spectral envelope. In: Information Processing (ICIP), International Conference on IEEE, p. 119-123.
Montagud, M., Boronat, F., González, J., and Pastor. J. (2017). Web-based platformfor subtitles customization and synchronization in multi-screen scenarios. Proc. Adjunct Publication ACM Int. Conf. Interact. Experiences TVOnlineVideo., 81-82
Orero., D.A.C.J.P. and Remael. A. (2007). Media for All: ‘Subtitling forthe Deaf, Audio Description, and Sign Language’. Amsterdam, The Netherlands: Rodopi. vol. 30.
Rodriguez-Alsina, A., Talavera, G., Orero, P., and J. Carrabina. (2013), Subtitlesynchronization across multiple screens and devices. Sensors., 12(7):8,710-8,731.
Schmidhuber and Jürgen. (2015). Deep learning in neural networks: An overview. Neural networks. 61:85-117.
Stan Adriana, Yoshitaka Mamiya, Junichi Yamagishi, Peter Bell, Oliver Watts, Robert AJ Clark, and Simon King. (2016). ALISA: An automatic lightly supervised speech segmentation and alignment tool. Computer Speech and Language., 35:116-133.
Szeto and Elson. (2015). Community of Inquiry as an instructional approach: What effects of teaching, social and cognitive presences are there in blended synchronous learning and teaching. Computers and Education., 81:191-201.
Vanderplank and Robert. (2016). Some essential themes in building the case for captions in language learning. In: Captioned Media in Foreign Language Learning and Teaching. Palgrave Macmillan, UK, p. 7-41.
Wilkerson J. and Casas. A. (2017). Large-scale computerized text analysis inpolitical science: Opportunities and challenges. In: Annu. Rev. Political Sci., vol. 20, no. 1, p. 529-544.
Wright, SJ., Kanevsky, D., Deng, L., He, X., Heigold, G., and Li, H. (2013). Optimization algorithms and applications for speech and language processing. IEEE Transactions on Audio. speech Language Processing., 2,231-43.
Wunsch, C.D., and Needleman. S.B. (1970). A general method applicable to thesearch for similarities in the amino acid sequence of two proteins. J. Mol.Biol., 48(3):443-453.
Xuemei, W., Yue, J., and Fang., W. (2017), The study of subtitle translation basedon multi-hierarchy semantic segmentation and extraction in digital video. Humanities Social Sci., 5(2):91-96.
Zhang, Y., Tang, Z., Zhang, C., Liu, J., and Lu, H. (2015). Automatic face annota-tion in TV series by video/script alignment. Neurocomputing., 152:316-321.








