Abstract:Since 2018, when the Transformer architecture was introduced, Natural Language Processing has gained significant momentum with pre-trained Transformer-based models that can be fine-tuned for various tasks. Most models are pre-trained on large English corpora, making them less applicable to other languages, such as Brazilian Portuguese. In our research, we identified two models pre-trained in Brazilian Portuguese (BERTimbau and PTT5) and two multilingual models (mBERT and mT5). BERTimbau and mBERT use only the Encoder module, while PTT5 and mT5 use both the Encoder and Decoder. Our study aimed to evaluate their performance on a financial Named Entity Recognition (NER) task and determine the computational requirements for fine-tuning and inference. To this end, we developed the Brazilian Financial NER (BraFiNER) dataset, comprising sentences from Brazilian banks' earnings calls transcripts annotated using a weakly supervised approach. Additionally, we introduced a novel approach that reframes the token classification task as a text generation problem. After fine-tuning the models, we evaluated them using performance and error metrics. Our findings reveal that BERT-based models consistently outperform T5-based models. While the multilingual models exhibit comparable macro F1-scores, BERTimbau demonstrates superior performance over PTT5. In terms of error metrics, BERTimbau outperforms the other models. We also observed that PTT5 and mT5 generated sentences with changes in monetary and percentage values, highlighting the importance of accuracy and consistency in the financial domain. Our findings provide insights into the differing performance of BERT- and T5-based models for the NER task.

ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language

PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data

Lite Training Strategies for Portuguese-English and English-Portuguese Translation

Tucano: Advancing Neural Text Generation for Portuguese

PORTULAN ExtraGLUE Datasets and Models: Kick-starting a Benchmark for the Neural Processing of Portuguese

Evaluating Named Entity Recognition: A comparative analysis of mono- and multilingual transformer models on a novel Brazilian corporate earnings call transcripts dataset

Advancing Neural Encoding of Portuguese with Transformer Albertina PT-*

Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Towards Effective and Efficient Continual Pre-training of Large Language Models

Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models

Sequence-to-Sequence Spanish Pre-trained Language Models

Evaluating Named Entity Recognition: Comparative Analysis of Mono- and Multilingual Transformer Models on Brazilian Corporate Earnings Call Transcriptions

Cabrita: closing the gap for foreign languages

Continual Mixed-Language Pre-Training for Extremely Low-Resource Neural Machine Translation

Empirical Analysis of Efficient Fine-Tuning Methods for Large Pre-Trained Language Models

Investigating Pre-trained Language Models on Cross-Domain Datasets, a Step Closer to General AI

From Brazilian Portuguese to European Portuguese

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese

SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token Detection

Transformer Models for Brazilian Portuguese Question Generation: An Experimental Study