Abstract:Corpora are applied to analyze and study the characteristics of the target language. In language education, corpora are playing an increasingly essential role due to their large capacity, authenticity, rapid and accurate retrieval, as well as quick and easy statistics. At present, a great number of universities are trying to apply the textbook corpus to English teaching. However, most of the existing corpora face the issue of poor sharing. In addition, these corpora may be limited to a specific textbook, which leads to the lack of wide coverage of the retrieval and analysis results. As a result, it is quite necessary to develop a set of English corpora that is highly relevant, well shared, and easy to use by fully integrating existing teaching resources according to the characteristics of English subjects in universities. In recent years, the use of corpus-assisted English language teaching has gained widespread attention and exploration as computers have become more and more popular. After all, a corpus-based teaching model can effectively eliminate the various drawbacks of traditional vocabulary teaching. In fact, the corpus has a large amount of authentic corpus. The authenticity and practicality of the corpus facilitate students’ mastery and use of English vocabulary in real contexts. What is more, the new model of corpus-assisted English vocabulary teaching can greatly increase independent learning and cooperative activities, so that students can increase their internal motivation for learning. This study begins with a brief introduction to the concept and characteristics of corpora. To be specific, the advantages of the corpus application in foreign language teaching are explained. At the same time, this research further analyzes the shortcomings of the existing corpus in university English education from the perspective of the current development and application of English corpora as well as clarifies the importance of building a corpus of university English teaching materials. After that, the system’s operating environment and main development techniques are determined according to the specific requirements of the corpus for university English textbooks. In other words, the overall design and detailed design of the corpus and its management system were then carried out on the basis of the chosen technology platform. In addition, the structure of the tables in the database is analyzed and the basic components and operation procedures of the system are introduced. Furthermore, the functional modules of the system are designed. At the same time, the automatic word and sentence separation methods of the original corpus, the corpus entry process, the cross-distance search of the corpus, and the statistical analysis of the search results are discussed in detail. In conclusion, this study is based on English text collection and data cleaning techniques to build an online corpus.

The Jinan Chinese Learner Corpus.

YACLC: A Chinese Learner Corpus with Multidimensional Annotation

A Learner Corpus - ESCCL

A learner corpus is born this way: From raw data to processed dataset

The Construction of Chinese Multi-dimensional Learner Corpus:YACLC

LSICC: A Large Scale Informal Chinese Corpus

UM-Corpus: A Large English-Chinese Parallel Corpus for Statistical Machine Translation

Building a Non-native Speech Corpus Featuring Chinese-English Bilingual Children: Compilation and Rationale

CityU corpus of essay drafts of English language learners: a corpus of textual revision in second language writing

Analysis of Chinese Character Writing Norms for Learners of Chinese as A Second Language

On Compiling Chinese Learner English Corpus for Ethnic Minorities

Saudi Learner Translation Corpus: The design and compilation of an English-Arabic learner translation corpus

On the Construction of Longitudinal Chinese Inter-language Corpus

Building a Large Japanese Web Corpus for Large Language Models

Designing and Implementing Online Japanese Intelligent Video Corpus: JV-Finder——Learning Cross-Cultural Communications in Video Contexts

A large-scale database of Chinese characters and words collected from elementary school textbooks

CSL: A Large-scale Chinese Scientific Literature Dataset

Online Corpus Construction of English Text Collection, Data Cleaning, and Similarity Analysis

Sell-Corpus: An Open Source Multiple Accented Chinese-English Speech Corpus For L2 English Learning Assessment

LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech

A Computer Corpus-Based Study of Chinese EFL Learners’ Use of Adverbial Connectors and Its Implications for Building a Language-Based Learning Environment