Abstract:The objective of this paper is to present and make publicly available the NILC-Metrix, a computational system comprising 200 metrics proposed in studies on discourse, psycholinguistics, cognitive and computational linguistics, to assess textual complexity in Brazilian Portuguese (BP). The metrics are relevant for descriptive analysis and the creation of computational models and can be used to extract information from various linguistic levels of written and spoken language. The metrics were developed during the last 13 years, starting in the end of 2007, within the scope of the PorSimples project. Once the PorSimples finished, new metrics were added to the initial 48 metrics of the Coh-Metrix-Port tool. Coh-Metrix-Port adapted some metrics to BP from the Coh-Metrix tool that computes metrics related to cohesion and coherence of texts in English. Given the large number of metrics, we present them following an organisation similar to the metrics of Coh-Metrix v3.0 to facilitate comparisons made with metrics in Portuguese and English, in future studies using both tools. In this paper, we illustrate the potential of the NILC-Metrix by presenting three applications: (i) a descriptive analysis of the differences between children's film subtitles and texts written for Elementary School I (comprises classes from 1st to 5th grade) and II (Final Years) (comprises classes from 6th to 9th grade, in an age group that corresponds to the transition between childhood and adolescence); (ii) a new predictor of textual complexity for the corpus of original and simplified texts of the PorSimples project; (iii) a complexity prediction model for school grades, using transcripts of children's story narratives told by teenagers. For each application, we evaluate which groups of metrics are more discriminative, showing their contribution for each task.

A Portuguese Native Language Identification Dataset

Quati: A Brazilian Portuguese Information Retrieval Dataset from Native Speakers

Language Variety Identification with True Labels

Native Language Identification with Large Language Models

OLID-BR: offensive language identification dataset for Brazilian Portuguese

Turkish Native Language Identification

PORTULAN ExtraGLUE Datasets and Models: Kick-starting a Benchmark for the Neural Processing of Portuguese

IndoNLI: A Natural Language Inference Dataset for Indonesian

SOME EFFECTS OF A PARTIALLY PURIFIED LYMPHOCYTE‐INHIBITING FACTOR FROM CALF THYMUS

Introducing Bode: A Fine-Tuned Large Language Model for Portuguese Prompt-Based Task

NILC-Metrix: assessing the complexity of written and spoken language in Brazilian Portuguese

JamPatoisNLI: A Jamaican Patois Natural Language Inference Dataset

ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification

SLIM-RAFT: A Novel Fine-Tuning Approach to Improve Cross-Linguistic Performance for Mercosur Common Nomenclature

LextPT: A reliable and efficient vocabulary size test for L2 Portuguese proficiency

XNLIeu: a dataset for cross-lingual NLI in Basque

PHOR-in-One: A multilingual lexical database with PHonological, ORthographic and PHonographic word similarity estimates in four languages

OCNLI: Original Chinese Natural Language Inference

GlórIA -- A Generative and Open Large Language Model for Portuguese

NATIVE LANGUAGE IDENTIFICATION FOR RUSSIAN USING ERRORS TYPES

TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian Portuguese