Abstract:Since 2018, when the Transformer architecture was introduced, Natural Language Processing has gained significant momentum with pre-trained Transformer-based models that can be fine-tuned for various tasks. Most models are pre-trained on large English corpora, making them less applicable to other languages, such as Brazilian Portuguese. In our research, we identified two models pre-trained in Brazilian Portuguese (BERTimbau and PTT5) and two multilingual models (mBERT and mT5). BERTimbau and mBERT use only the Encoder module, while PTT5 and mT5 use both the Encoder and Decoder. Our study aimed to evaluate their performance on a financial Named Entity Recognition (NER) task and determine the computational requirements for fine-tuning and inference. To this end, we developed the Brazilian Financial NER (BraFiNER) dataset, comprising sentences from Brazilian banks' earnings calls transcripts annotated using a weakly supervised approach. Additionally, we introduced a novel approach that reframes the token classification task as a text generation problem. After fine-tuning the models, we evaluated them using performance and error metrics. Our findings reveal that BERT-based models consistently outperform T5-based models. While the multilingual models exhibit comparable macro F1-scores, BERTimbau demonstrates superior performance over PTT5. In terms of error metrics, BERTimbau outperforms the other models. We also observed that PTT5 and mT5 generated sentences with changes in monetary and percentage values, highlighting the importance of accuracy and consistency in the financial domain. Our findings provide insights into the differing performance of BERT- and T5-based models for the NER task.

BERTaú: Itaú BERT for digital customer service

Portuguese FAQ for Financial Services

Augmenting Customer Support with an NLP-based Receptionist

Evaluating Named Entity Recognition: A comparative analysis of mono- and multilingual transformer models on a novel Brazilian corporate earnings call transcripts dataset

Portuguese Named Entity Recognition using BERT-CRF

BERT for Sentiment Analysis: Pre-trained and Fine-Tuned Alternatives

A Financial Service Chatbot based on Deep Bidirectional Transformers

Real Life Application of a Question Answering System Using BERT Language Model

Cabrita: closing the gap for foreign languages

Advancing Neural Encoding of Portuguese with Transformer Albertina PT-*

Transformer Models for Brazilian Portuguese Question Generation: An Experimental Study

Deploying a BERT-based Query-Title Relevance Classifier in a Production System: a View from the Trenches

Performance and sustainability of BERT derivatives in dyadic data

Sentiment Analysis in Portuguese Restaurant Reviews: Application of Transformer Models in Edge Computing

TourBERT: A pretrained language model for the tourism industry

Evaluating Named Entity Recognition: Comparative Analysis of Mono- and Multilingual Transformer Models on Brazilian Corporate Earnings Call Transcriptions

FinBERT: A Pre-trained Financial Language Representation Model for Financial Text Mining

RoBERTuito: a pre-trained language model for social media text in Spanish

CA-BERT: Leveraging Context Awareness for Enhanced Multi-Turn Chat Interaction

SentPT: A customized solution for multi-genre sentiment analysis of Portuguese-language texts

DB-BERT: a Database Tuning Tool that "Reads the Manual"