Ulysses Tesemõ: a new large corpus for Brazilian legal and governmental domain

Felipe A. Siqueira,Douglas Vitório,Ellen Souza,José A. P. Santos,Hidelberg O. Albuquerque,Márcio S. Dias,Nádia F. F. Silva,André C. P. L. F. de Carvalho,Adriano L. I. Oliveira,Carmelo Bastos-Filho
DOI: https://doi.org/10.1007/s10579-024-09762-8
2024-07-20
Language Resources and Evaluation
Abstract:The increasing use of artificial intelligence methods in the legal field has sparked interest in applying Natural Language Processing techniques to handle legal tasks and reduce the workload of these professionals. However, the availability of legal corpora in Portuguese, especially for the Brazilian legal domain, is limited. Existing resources offer some legal data but lack comprehensive coverage. To address this gap, we present Ulysses Tesemõ, a large corpus specifically built for the Brazilian legal domain. The corpus consists of over 3.5 million files, totaling 30.7 GiB of raw text, collected from 159 sources encompassing judicial, legislative, academic, news, and other related data. The data was collected by scraping public information from governmental websites, emphasizing contents generated over the past two decades. We categorized the obtained files into 30 distinct categories, covering various branches of the Brazilian government and different types of texts. The corpus retains the original content with minimal data transformations, addressing the scarcity of Portuguese legal corpora and providing researchers with a valuable resource for advancing in the research area.
computer science, interdisciplinary applications
What problem does this paper attempt to address?