HYBRID SYSTEM FOR VISUAL QUESTION LOCATION GENERATION AND ANSWERING IN GASTRO-INTESTINAL ENDOSCOPY IMAGES

Main Article Content

Rajeswari Jayaraman , Kavitha Srinivasan,, Divyasri Krishnakumar, Cyril Melvin Vincent

Abstract

The Visual Question Location Generation and Answering (VQLGA) is a hybrid system, which combines the Visual Question Answering (VQA) and Visual Location-based Question Answering (VLQA) approaches together for assisting the healthcare professionals. VQLGA system focuses on gastrointestinal (GI) endoscopy images of ImageCLEF med 2023 challenge, detecting and preventing colorectal cancer. The GI endoscopy images generated through endoscopy procedures are complex in nature, therefore analyzing these images are time consuming for medical experts. To address this challenge a hybrid system is designed and implemented using suitable Deep Learning techniques with appropriate quantitative metrics for validation at each stage of process. VQLGA system, leverages VGG16 for image features and LSTM for text features in the VQA model achieving an accuracy of 80%. The VLQA model, driven by DenseNet121 and UNet architectures, attains a remarkable Intersection over Union (IoU) score of 83.4%. By integrating these models, the hybrid VQLGA system achieved an accuracy of 71% and IoU of 81.8% is evidence for enhanced diagnostic efficiency through proposed system. Additionally, the segmented output of the VQLGA system is validated using Explainable AI (XAI) technique to support the segmented region through visualization.

Article Details

Section
Articles