Abstract:Background: The regulatory affairs (RA) division in a pharmaceutical establishment is the point of contact between regulatory authorities and pharmaceutical companies. They are delegated the crucial and strenuous task of extracting and summarizing relevant information in the most meticulous manner from various search systems. An artificial intelligence (AI)-based intelligent search system that can significantly bring down the manual efforts in the existing processes of the RA department while maintaining and improving the quality of final outcomes is desirable. We proposed a "frequently asked questions" component and its utility in an AI-based intelligent search system in this paper. The scenario is further complicated by the lack of publicly available relevant data sets in the RA domain to train the machine learning models that can facilitate cognitive search systems for regulatory authorities. Objective: In this study, we aimed to use AI-based intelligent computational models to automatically recognize semantically similar question pairs in the RA domain and evaluate the Recognizing Question Entailment-based system. Methods: We used transfer learning techniques and experimented with transformer-based models pretrained on corpora collected from different resources, such as Bidirectional Encoder Representations from Transformers (BERT), Clinical BERT, BioBERT, and BlueBERT. We used a manually labeled data set that contained 150 question pairs in the pharmaceutical regulatory domain to evaluate the performance of our model. Results: The Clinical BERT model performed better than other domain-specific BERT-based models in identifying question similarity from the RA domain. The BERT model had the best ability to learn domain-specific knowledge with transfer learning, which reached the best performance when fine-tuned with sufficient clinical domain question pairs. The top-performing model achieved an accuracy of 90.66% on the test set. Conclusions: This study demonstrates the possibility of using pretrained language models to recognize question similarity in the pharmaceutical regulatory domain. Transformer-based models that are pretrained on clinical notes perform better than models pretrained on biomedical text in recognizing the question's semantic similarity in this domain. We also discuss the challenges of using data augmentation techniques to address the lack of relevant data in this domain. The results of our experiment indicated that increasing the number of training samples using back translation and entity replacement did not enhance the model's performance. This lack of improvement may be attributed to the intricate and specialized nature of texts in the regulatory domain. Our work provides the foundation for further studies that apply state-of-the-art linguistic models to regulatory documents in the pharmaceutical industry.

ALBERT-QM: An ALBERT Based Method for Chinese Health Related Question Matching (Preprint)

A Joint Model For Question-Answering Over Traditional Chinese Medicine

A Question-Answering System over Traditional Chinese Medicine

Domain-specific Cross-Language Relevant Question Retrieval.

HealthQA: A Chinese QA Summary System for Smart Health.

Optimized Biomedical Question-Answering Services with LLM and Multi-BERT Integration

Exploration of text matching methods in Chinese disease Q&A systems: A method using ensemble based on BERT and boosted tree models

A Chinese Question Answering System in Medical Domain

Medical Data Inquiry Using a Question Answering Model.

A medical question answering system using large language models and knowledge graphs

BERT-Based Mixed Question Answering Matching Model

Benchmarking for biomedical natural language processing tasks with a domain specific ALBERT

Match<SUP>2</SUP>: A Matching over Matching Model for Similar Question Identification

Identifying the Question Similarity of Regulatory Documents in the Pharmaceutical Industry by Using the Recognizing Question Entailment System: Evaluation Study

HHH: An Online Medical Chatbot System based on Knowledge Graph and Hierarchical Bi-Directional Attention

Research on semantic matching algorithm of BERT intelligent question answering system

Huatuo-26M, a Large-scale Chinese Medical QA Dataset

CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering

Using Semantic Text Similarity calculation for question matching in a rheumatoid arthritis question-answering system

MAGE: Multi-scale Context-aware Interaction based on Multi-granularity Embedding for Chinese Medical Question Answer Matching