Abstract:BackgroundCardiovascular disease (CVD) has become the leading cause of death in China, and most of the cases can be prevented by controlling risk factors. The goal of this study was to build a corpus of CVD risk factor annotations based on Chinese electronic medical records (CEMRs). This corpus is intended to be used to develop a risk factor information extraction system that, in turn, can be applied as a foundation for the further study of the progress of risk factors and CVD.ResultsWe designed a light annotation task to capture CVD risk factors with indicators, temporal attributes and assertions that were explicitly or implicitly displayed in the records. The task included: 1) preparing data; 2) creating guidelines for capturing annotations (these were created with the help of clinicians); 3) proposing an annotation method including building the guidelines draft, training the annotators and updating the guidelines, and corpus construction. Meanwhile, we proposed some creative annotation guidelines: (1) the under-threshold medical examination values were annotated for our purpose of studying the progress of risk factors and CVD; (2) possible and negative risk factors were concerned for the same reason, and we created assertions for annotations; (3) we added four temporal attributes to CVD risk factors in CEMRs for constructing long term variations. Then, a risk factor annotated corpus based on de-identified discharge summaries and progress notes from 600 patients was developed. Built with the help of clinicians, this corpus has an inter-annotator agreement (IAA) F1-measure of 0.968, indicating a high reliability.ConclusionTo the best of our knowledge, this is the first annotated corpus concerning CVD risk factors in CEMRs and the guidelines for capturing CVD risk factor annotations from CEMRs were proposed. The obtained document-level annotations can be applied in future studies to monitor risk factors and CVD over the long term.

Developing a linguistically annotated corpus of Chinese electronic medical record

Temporal Expression Recognition and Temporal Relationship Extraction from Chinese Narrative Medical Records

Lexical Characteristics Analysis of Chinese Clinical Documents

Building a comprehensive syntactic and semantic corpus of Chinese clinical texts

A unified framework of medical information annotation and extraction for Chinese clinical text

A Fine-Grained Chinese Word Segmentation and Part-of-speech Tagging Corpus for Clinical Text

Construction, evaluation, and application of an electronic medical record corpus for cerebral palsy rehabilitation

LI-EMRSQL: Linking Information Enhanced Text2SQL Parsing on Complex Electronic Medical Records

Constructing a Chinese Electronic Medical Record Corpus for Named Entity Recognition on Resident Admit Notes

Chinese Word Segmentation in Flectronic Medical Record Text via Graph Neural Network-Bidirectional LSTM-CRF Model

Cross-department chunking based on Chinese electronic medical record

Named Entity Recognition in Chinese Electronic Medical Records Based on CRF.

Automatic Conversion of Electronic Medical Record Text for OpenEHR Based on Semantic Analysis

Black-Box Segmentation of Electronic Medical Records

Hybrid Granularity-Based Medical Event Extraction in Chinese Electronic Medical Records

Annotating the Contemporary Chinese Corpus

Developing a cardiovascular disease risk factor annotated corpus of Chinese electronic medical records

Extracting Clinical Entities And Their Assertions From Chinese Electronic Medical Records Based On Machine Learning

Research on Named Entity Recognition in Chinese EMR Based on Semi-Supervised Learning with Dual Selected Strategy.

Development of large-scale TCM corpus using hybrid named entity recognition methods for clinical phenotype detection: An initial study

Overview of CCKS 2018 Task 1 - Named Entity Recognition in Chinese Electronic Medical Records.