Abstract:Corpora are applied to analyze and study the characteristics of the target language. In language education, corpora are playing an increasingly essential role due to their large capacity, authenticity, rapid and accurate retrieval, as well as quick and easy statistics. At present, a great number of universities are trying to apply the textbook corpus to English teaching. However, most of the existing corpora face the issue of poor sharing. In addition, these corpora may be limited to a specific textbook, which leads to the lack of wide coverage of the retrieval and analysis results. As a result, it is quite necessary to develop a set of English corpora that is highly relevant, well shared, and easy to use by fully integrating existing teaching resources according to the characteristics of English subjects in universities. In recent years, the use of corpus-assisted English language teaching has gained widespread attention and exploration as computers have become more and more popular. After all, a corpus-based teaching model can effectively eliminate the various drawbacks of traditional vocabulary teaching. In fact, the corpus has a large amount of authentic corpus. The authenticity and practicality of the corpus facilitate students’ mastery and use of English vocabulary in real contexts. What is more, the new model of corpus-assisted English vocabulary teaching can greatly increase independent learning and cooperative activities, so that students can increase their internal motivation for learning. This study begins with a brief introduction to the concept and characteristics of corpora. To be specific, the advantages of the corpus application in foreign language teaching are explained. At the same time, this research further analyzes the shortcomings of the existing corpus in university English education from the perspective of the current development and application of English corpora as well as clarifies the importance of building a corpus of university English teaching materials. After that, the system’s operating environment and main development techniques are determined according to the specific requirements of the corpus for university English textbooks. In other words, the overall design and detailed design of the corpus and its management system were then carried out on the basis of the chosen technology platform. In addition, the structure of the tables in the database is analyzed and the basic components and operation procedures of the system are introduced. Furthermore, the functional modules of the system are designed. At the same time, the automatic word and sentence separation methods of the original corpus, the corpus entry process, the cross-distance search of the corpus, and the statistical analysis of the search results are discussed in detail. In conclusion, this study is based on English text collection and data cleaning techniques to build an online corpus.

The Application of NLTK Library for Python Natural Language Processing in Corpus Research

Development and Evaluation of Task-Specific NLP Framework in China.

The Comparative study of Python Libraries for Natural Language Processing (NLP)

Recent Developments in Chinese Corpus Research

Some Ideas on Natural Language Processing

Exploring the Landscape of Natural Language Processing Research

Survey of Natural Language Processing for Education: Taxonomy, Systematic Review, and Future Trends

Online Corpus Construction of English Text Collection, Data Cleaning, and Similarity Analysis

DIY Parallel Corpora for Petroleum Production Engineering and Its Academic Application

Construction of English and American Literature Corpus Based on Machine Learning Algorithm

Deep Learning and Its Applications to Natural Language Processing

Research Status and Current Problems of Corpus Linguistics in China

A Comprehensive Analytical Study of Traditional and Recent Development in Natural Language Processing

From Text to CQL: Bridging Natural Language and Corpus Search Engine

Language Processing and Python

Corpus Analysis with spaCy

Usable Amharic text corpus for natural language processing applications

Digital language resources and NLP tools

Research on Corpus Creation and Development of Chinese Traditional Medicine

A Bibliometric Review of Natural Language Processing Applications in Psychology from 1991 to 2023

Progress in the Application of Natural Language Processing to Information Retrieval Tasks