Abstract:Abstract The detailed physiological perspectives captured by medical imaging provides actionable insights to doctors to manage comprehensive care of patients. However, the quality of such diagnostic image modalities is often affected by mismanagement of the image capturing process by poorly trained technicians and older/poorly maintained imaging equipment. Further, a patient is often subjected to scanning at different orientations to capture the frontal, lateral and sagittal views of the affected areas. Due to the large volume of diagnostic scans performed at a modern hospital, adequate documentation of such additional perspectives is mostly overlooked, which is also an essential key element of quality diagnostic systems and predictive analytics systems. Another crucial challenge affecting effective medical image data management is that the diagnostic scans are essentially stored as unstructured data, lacking a well-defined processing methodology for enabling intelligent image data management for supporting applications like similar patient retrieval , automated disease prediction etc. One solution is to incorporate automated diagnostic image descriptions of the observation/findings by leveraging computer vision and natural language processing. In this work, we present multi-task neural models capable of addressing these critical challenges. We propose ESRGAN, an image enhancement technique for improving the quality and visualization of medical chest x-ray images, thereby substantially improving the potential for accurate diagnosis, automatic detection and region-of-interest segmentation. We also propose a CNN-based model called ViewNet for predicting the view orientation of the x-ray image and generating a medical report using Xception net, thus facilitating a robust medical image management system for intelligent diagnosis applications. Experimental results are demonstrated using standard metrics like BRISQUE, PIQE and BLEU scores, indicating that the proposed models achieved excellent performance. Further, the proposed deep learning approaches enable diagnosis in a lesser time and their hybrid architecture shows significant potential for supporting many intelligent diagnosis applications.

MATNet: Exploiting Multi-Modal Features for Radiology Report Generation.

VMEKNet: Visual Memory and External Knowledge Based Network for Medical Report Generation.

Automatic Report Generation Method Based on Multiscale Feature Extraction and Word Attention Network.

An Inclusive Task-Aware Framework for Radiology Report Generation

MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report Generation

Automated Radiographic Report Generation Purely on Transformer: A Multicriteria Supervised Approach

Radiology Report Generation Using Transformers Conditioned with Non-imaging Data

Generating Radiology Reports via Memory-driven Transformer

Deep neural models for automated multi-task diagnostic scan management—quality enhancement, view classification and report generation

Multi-modal transformer architecture for medical image analysis and automated report generation

Generating radiology reports via auxiliary signal guidance and a memory-driven network

Memory-based Cross-modal Semantic Alignment Network for Radiology Report Generation

Radiology Report Generation with a Learned Knowledge Base and Multi-Modal Alignment

A medical report generation method integrating teacher–student model and encoder–decoder network

A Medical Semantic-Assisted Transformer for Radiographic Report Generation

Salad, house dressing, but hold the sulfites.

[Research on automatic generation of multimodal medical image reports based on memory driven]

CSAMDT: Conditional Self Attention Memory-Driven Transformers for Radiology Report Generation from Chest X-Ray

Multi-modality Regional Alignment Network for Covid X-Ray Survival Prediction and Report Generation

A label information fused medical image report generation framework