QUESTION ANSWERING SYSTEM USING BERT ON CANCER DATA

Authors

  • Deva Kumar Salluri Department of CSE, VFSTR deemed to be University, Andhra Pradesh 522213, India.
  • Jyothi Swaroop Badiginchala Department of CSE, Vignan’s Lara Institute of Technology and Science, Andhra Pradesh 522213, India.
  • Veerendra Bellapukonda Department of CSE, Vignan’s Lara Institute of Technology and Science, Andhra Pradesh 522213, India.
  • Kudaravalli Deepika Department of CSE, Vignan’s Lara Institute of Technology and Science, Andhra Pradesh 522213, India.

Keywords:

Cancer diseases, BERT model, Natural Language Processing, question answering, medical domain, artificial Intelligence

Abstract

Question Answering systems can be used to generate an answer as a response to a question asked by the user, rather than searching from a large pool of files including answers. The main vision of the system in Natural Language Processing is to generate perfect answers for the queries asked by the users. The System needs to have information about each kind of Cancer disease. So, to achieve this there is a need for a system that accepts the questions and paragraphs as input and returns the answers as fast and accurately as possible. To achieve this, authors need a Language Model, which is a probability distribution over sequences of text sentences. The model gives a context understanding between the bag of words and sentences. Among the all-language models, authors chose the Bert-large-uncased-whole-word-masking-fine-tuned-squad model for the question answering system. By implementing this method will be helpful to identify diseases in the early stages which will be useful to the patients.

References

Annamoradnejad, I., Fazli, M., and Habibi, J. (2020). Predicting Subjective features from questions on qa websites using BERT. In: 2020 6th International Conference on Web Research, (ICWR), 2020, p. 240-244. doi: 10.1109/ICWR 49608.2020.9122318.

Athenikos, S.J., Han, H., and Brooks, A.D. (2009). A framework of a logic-based question-answering system for the medical domain (LOQAS-Med).In: Proceedings of the ACM Symposium on Applied Computing, p. 847-851. doi: 10.1145/1529282.1529462.

Bao, Q., Ni, L., and Liu, J. (2020). HHH: An Online Medical Chatbot System Based on Knowledge Graph and Hierarchical Bi-Directional Attention. In: Proceedings of the Australasian Computer Science Week Multiconference (ACSW '20). Association for Computing Machinery, New York, NY, USA, Article no. 32:1-10. https://doi.org/ 10.1145/3373017.3373049.

Burk, H. (1999). Das Hunderttage-Stadion: Entstehungsgeschichte des Bad Nauheimer Kunsteisstadions unter Colonel Paul R. Knight. Stadt Bad Nauheim, 63p.

Chaix, B., Bibault, J., Pienkowski, A., Delamon, G., Guillemassé, A., Nectoux, P., and Brouard. B. (2019). When chatbots meet patients: one-year prospective study of conversations between patients with breast cancer and a chatbot. JMIR Cancer, 5(1):e12856. doi: 10.2196/12856.

Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019). BERT: pre-training of deep bidirectional transformers for language understanding. In: NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 1(Mlm):4,171-4,186.

Djuric, N., Zhou, J., Morris, R., Grbovic, M., Radosavljevic, V., and Bhamidipati, N. (2015). Doc2Vec. In: Proceedings of the 24th International Conference on World Wide Web - WWW ’15 Companion, 32:29-30.

Do, T., Nguyen, B.X., Tjiputra, E., Tran, M., Tran, Q.D., and Nguyen, A. (2021). Multiple meta-model quantifying for medical visual question answering. Comp. Vis. Patt. Recogn., (i):1-15. https://doi.org/10.48550/arXiv.2105. 08913

Evans, H.P., Anastasiou, A., Edwards, A., Hibbert, P., Makeham, M., Luz, S., Sheikh, A., Donaldson, L., and Carson-Stevens, A. (2020). Automated classification of primary care patient safety incident report content and severity using supervised machine learning (ML) approaches. Health Inform. J., 26(4):3,123-3,139. doi: 10.1177/14604 58219833102.

He, J., Fu, M., and Tu, M. (2019). Applying deep matching networks to chinese medical question answering: a study and a dataset. BMC Med. Inform. Decis. Mak., 19(Suppl 2):52. doi: 10.1186/s12911-019-0761-8.

He, P., Liu, X., Gao, J., and Chen, W. (2020a). DeBERTa: decoding-enhanced BERT with disentangled attention. Comp. Langu., p. 1-21. https://doi.org/10.48550/arXiv.2006. 03654.

He, X., Zhang, Y., Mou, L., Xing, E., and Xie, P. (2020b). PathVQA: 30000+ questions for medical visual question answering. Comp. Langu., p. 1-12. https://doi.org/10.48550/ arXiv.2003.10286.

Hoermann, S., McCabe, K.L., Milne, D.N., and Calvo, R.A. (2017). Application of synchronous text-based dialogue systems in mental health interventions: systematic review. J. Med. Int. Res., 19(8):e267. doi: 10.2196/jmir.7023.

Jackson, R.G., Patel, R., Jayatilleke, N., Kolliakou, A., Ball, M., Gorrell, G., Roberts, A., Dobson, R.J., and Stewart, R. (2017). Natural language processing to extract symptoms of severe mental illness from clinical text: the clinical record interactive search comprehensive data extraction (CRIS-CODE) project. BMJ Open 7(1):e012012. doi: 10.1136/bmjopen-2016-012012.

Karri, S.P.R. and Kumar. B.S. (2020). Deep learning techniques for implementation of chatbots. In: 2020 International Conference on Computer Communication and Informatics, (ICCCI) p. 1-5. doi: 10.1109/ICCCI48352.2020.9104143.

Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2019). ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations. Comp. Langu., p. 1-17. https://doi.org/10.48550/arXiv.1909.11942.

Lee, M., Cimino, J., Zhu, H.R., Sable, C., Shanker, V., Ely, J., and Yu, H. (2006). Beyond information retrieval--medical question answering. AMIA ... Annual Symposium Proceedings / AMIA Symposium. AMIA Symposium, p. 469-473.

Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. (2020a). BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 7,871-7,880. doi: 10.18653/v1/2020.acl-main.703.

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. (2020b). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Comp. Langu., p. 1-19. https://doi.org/10.48550/arXiv.2005.11401.

Liang, H., Tsui, B.Y., Ni, H., Valentim, C.C.S., Baxter, S.L., Liu, G., Cai, W., Kermany, D.S., Sun, X., Chen, J., He, L., Zhu, J., Tian, P., Shao, H., Zheng, L., Hou, R., Hewett, S., Li, G., Liang, P., Zang, X., Zhang, Z., Pan, L., Cai, H., Ling, R., Li, S., Cui, Y., Tang, S., Ye, H., Huang, X., He, W., Liang, W., Zhang, Q., Jiang, J., Yu, W., Gao, J., Ou, W.,Deng, Y., Hou, Q., Wang, B., Yao, C., Liang, Y., Zhang, S., Duan, Y., Zhang, R., Gibson, S., Zhang, C.L., Li, O., Zhang, E.D., Karin, G., Nguyen, N., Wu, X., Wen, C., Xu, J., Xu, W., Wang, B., Wang, W., Li, J., Pizzato, B., Bao, C., Xiang, D., He, W., He, S., Zhou, Y., Haw, W., Goldbaum, M., Tremoulet, A., Hsu, C.N., Carter, H., Zhu, L., Zhang, K., and Xia, H. (2019). Evaluation and accurate diagnoses of pediatric diseases using artificial intelligence. Nature Med., 25(3):433-438. DOI: 10.1038/s41591-018-0335-9.

Logeswaran, L. and Lee, H. (2018). An efficient framework for learning sentence representations. In: 6th International Conference on Learning Representations, ICLR 2018 - Conference Track Proceedings, p. 1-16.

Min, S., Chen, D., Zettlemoyer, L., and Hajishirzi, H. (2019). Knowledge guided text retrieval and reading for open domain question answering. Comp. Langu., p. 1-12. https://doi.org/10.48550/arXiv.1911.03868.

Nguyen, B.D., Do, T.T., Nguyen, B.X., Do, T., Tjiputra, E., and Tran, Q.D. (2019). Overcoming data limitation in medical visual question answering. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 11767 LNCS:522-530. doi: 10.1007/978-3-030-32251-9_57.

Omoregbe, N.A.I., Ndaman, I.O., Misra, S., Abayomi-Alli, O.O., and Damaševičius, R. (2020). Text messaging-based medical diagnosis using natural language processing and fuzzy logic. J. Health. Eng., doi: 10.1155/2020/8839524.

Parikh, A.P., Täckström, O., Das, D., and Uszkoreit, J. (2016). A decomposable attention model for natural language inference. In: EMNLP 2016 - Conference on Empirical Methods in Natural Language Processing, Proceedings, p. 2,249- 2,255. doi: 10.18653/v1/d16-1244.

Peters, M.E., Ammar, W., Bhagavatula, C., and Power, R. (2017). Semi-supervised sequence tagging with bidirectional language models. In: ACL 2017 – 55th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (Long Papers), 1:1,756-1,765. doi: 10.18653/v1/P17-1161.

Peters, M.E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018). Deep contextualized word representations. In: NAACL HLT 2018 – 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 1:2,227=2,237. doi: 10.18653/v1/n18-1202.

Putra, F.B., Yusuf, A.A., Yulianus, H., Pratama, Y. P., Humairra, D.S., Erifani, U., Basuki, D.K., Sukaridhoto, S., and Budiarti, R.P.N. (2019). Identification of symptoms based on natural language processing (NLP) for disease diagnosis based on international classification of diseases and related health problems (ICD-11). In: 2019- International Electronics Symposium: The Role of Techno-Intelligence in Creating an Open Energy System Towards Energy Democracy, Proceedings, p. 1-5. Doi: 10.1109/ELECSYM.2019.8901644.

Qiu, X. and Huang. X. (2015). Convolutional neural tensor network architecture for community-based question answering. In: IJCAI International Joint Conference on Artificial Intelligence, 2015-Janua(Ijcai):1,305-1,311.

Sanh, V., Debut, L., Chaumond, J., and Wolf, T. (2019). DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter. Comp. Langu., p. 2-6. https://doi. org/10.48550/arXiv.1910.01108

Zhang, X., Wu, J., He, Z., Liu, X., and Su, Y. (2018). Medical exam question answering with large-scale reading comprehension. In: 32nd AAAI Conference on Artificial Intelligence, AAAI, p. 5,706-5,713. doi: https://doi.org/ 10.1609/aaai.v32i1.11970.

Zhou, X., Wu, B., and Zhou, Q. (2018). A depth evidence score fusion algorithm for chinese medical intelligence question answering system. J. Health. Eng., 2018:1205354. doi: 10.1155/2018/1205354. PMID: 30123438; PMCID: PMC6079581.

Downloads

Published

2026-08-28

How to Cite

Kumar Salluri, D., Swaroop Badiginchala, J., Bellapukonda, V., & Deepika, K. (2026). QUESTION ANSWERING SYSTEM USING BERT ON CANCER DATA. Suranaree Journal of Science and Technology, 29(5), 010161(1–11). retrieved from https://ph04.tci-thaijo.org/index.php/SUJST/article/view/15177

Issue

Section

Research Article