Abstract:Corpora are applied to analyze and study the characteristics of the target language. In language education, corpora are playing an increasingly essential role due to their large capacity, authenticity, rapid and accurate retrieval, as well as quick and easy statistics. At present, a great number of universities are trying to apply the textbook corpus to English teaching. However, most of the existing corpora face the issue of poor sharing. In addition, these corpora may be limited to a specific textbook, which leads to the lack of wide coverage of the retrieval and analysis results. As a result, it is quite necessary to develop a set of English corpora that is highly relevant, well shared, and easy to use by fully integrating existing teaching resources according to the characteristics of English subjects in universities. In recent years, the use of corpus-assisted English language teaching has gained widespread attention and exploration as computers have become more and more popular. After all, a corpus-based teaching model can effectively eliminate the various drawbacks of traditional vocabulary teaching. In fact, the corpus has a large amount of authentic corpus. The authenticity and practicality of the corpus facilitate students’ mastery and use of English vocabulary in real contexts. What is more, the new model of corpus-assisted English vocabulary teaching can greatly increase independent learning and cooperative activities, so that students can increase their internal motivation for learning. This study begins with a brief introduction to the concept and characteristics of corpora. To be specific, the advantages of the corpus application in foreign language teaching are explained. At the same time, this research further analyzes the shortcomings of the existing corpus in university English education from the perspective of the current development and application of English corpora as well as clarifies the importance of building a corpus of university English teaching materials. After that, the system’s operating environment and main development techniques are determined according to the specific requirements of the corpus for university English textbooks. In other words, the overall design and detailed design of the corpus and its management system were then carried out on the basis of the chosen technology platform. In addition, the structure of the tables in the database is analyzed and the basic components and operation procedures of the system are introduced. Furthermore, the functional modules of the system are designed. At the same time, the automatic word and sentence separation methods of the original corpus, the corpus entry process, the cross-distance search of the corpus, and the statistical analysis of the search results are discussed in detail. In conclusion, this study is based on English text collection and data cleaning techniques to build an online corpus.

Corpus Analysis with spaCy

MAKING USE OF A ‘SPACY’ MODULE IN THE NATURAL LANGUAGE PROCESSING

Launching into clinical space with medspaCy: a new clinical text processing toolkit in Python

Text Summarizer Using SpaCy in NLP

The Application of NLTK Library for Python Natural Language Processing in Corpus Research

Design and implementation of an open source Greek POS Tagger and Entity Recognizer using spaCy

Online Corpus Construction of English Text Collection, Data Cleaning, and Similarity Analysis

Corpus Linguistics in EFL Language Teaching: Insights From Research and Practice

ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing

Text Analysis in Python for Social Scientists

SpaDeLeF: A Dataset for Hierarchical Classification of Lexical Functions for Collocations in Spanish

MedLexSp - a medical lexicon for Spanish medical natural language processing

What's In My Big Data?

A Corpus-based Analysis of the Terminology of the Social Sciences and Humanities

Multi-Mosaics: Corpus Summarizing and Exploration using multiple Concordance Mosaic Visualisations

Corpus Tools and Technology

A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature

English Corpus Linguistics

ILiAD: An Interactive Corpus for Linguistic Annotated Data from Twitter Posts

Corpus Linguistics Challenging Traditional Grammar