Abstract:Purpose This article presents two Brazilian Portuguese corpora collected from different media concerning public security issues in a specific location. The primary motivation is supporting analyses, so security authorities can make appropriate decisions about their actions. Design/methodology/approach The corpora were obtained through web scraping from a newspaper's website and tweets from a Brazilian metropolitan region. Natural language processing was applied considering: text cleaning, lemmatization, summarization, part-of-speech and dependencies parsing, named entities recognition, and topic modeling. Findings Several results were obtained based on the methodology used, highlighting some: an example of a summarization using an automated process; dependency parsing; the most common topics in each corpus; the forty named entities and the most common slogans were extracted, highlighting those linked to public security. Research limitations/implications Some critical tasks were identified for the research perspective, related to the applied methodology: the treatment of noise from obtaining news on their source websites, passing through textual elements quite present in social network posts such as abbreviations, emojis/emoticons, and even writing errors; the treatment of subjectivity, to eliminate noise from irony and sarcasm; the search for authentic news of issues within the target domain. All these tasks aim to improve the process to enable interested authorities to perform accurate analyses. Practical implications The corpora dedicated to the public security domain enable several analyses, such as mining public opinion on security actions in a given location; understanding criminals' behaviors reported in the news or even on social networks and drawing their attitudes timeline; detecting movements that may cause damage to public property and people welfare through texts from social networks; extracting the history and repercussions of police actions, crossing news with records on social networks; among many other possibilities. Originality/value The work on behalf of the corpora reported in this text represents one of the first initiatives to create textual bases in Portuguese, dedicated to Brazil's specific public security domain.

Ulysses Tesemõ: a new large corpus for Brazilian legal and governmental domain

LegalNLP -- Natural Language Processing methods for the Brazilian Legal Language

Towards corpora creation from social web in Brazilian Portuguese to support public security analyses and decisions

Building a relevance feedback corpus for legal information retrieval in the real-case scenario of the Brazilian Chamber of Deputies

Datasets for Portuguese Legal Semantic Textual Similarity: Comparing weak supervision and an annotation process approaches

Segmenting Brazilian legislative text using weak supervision and active learning

SemClinBr -- a multi institutional and multi specialty semantically annotated corpus for Portuguese clinical NLP tasks

Towards a parallel corpus of Portuguese and the Bantu language Emakhuwa of Mozambique

TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian Portuguese

ACE-2005-PT: Corpus for Event Extraction in Portuguese

Historical Portuguese corpora: a survey

Essay-BR: a Brazilian Corpus of Essays

CDJUR-BR -- A Golden Collection of Legal Document from Brazilian Justice with Fine-Grained Named Entities

Portuguese FAQ for Financial Services

Building a Sentiment Corpus of Tweets in Brazilian Portuguese

Predicting Legal Proceedings Status: Approaches Based on Sequential Text Data

CORAA: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian Portuguese

Analysing similarities between legal court documents using natural language processing approaches based on Transformers

Quati: A Brazilian Portuguese Information Retrieval Dataset from Native Speakers

TuPy-E: detecting hate speech in Brazilian Portuguese social media with a novel dataset and comprehensive analysis of models

A Legal Framework for Natural Language Processing Model Training in Portugal