Abstract:Collecting labeled datasets in finance is challenging due to scarcity of domain experts and higher cost of employing them. While Large Language Models (LLMs) have demonstrated remarkable performance in data annotation tasks on general domain datasets, their effectiveness on domain specific datasets remains underexplored. To address this gap, we investigate the potential of LLMs as efficient data annotators for extracting relations in financial documents. We compare the annotations produced by three LLMs (GPT-4, PaLM 2, and MPT Instruct) against expert annotators and crowdworkers. We demonstrate that the current state-of-the-art LLMs can be sufficient alternatives to non-expert crowdworkers. We analyze models using various prompts and parameter settings and find that customizing the prompts for each relation group by providing specific examples belonging to those groups is paramount. Furthermore, we introduce a reliability index (LLM-RelIndex) used to identify outputs that may require expert attention. Finally, we perform an extensive time, cost and error analysis and provide recommendations for the collection and usage of automated annotations in domain-specific settings.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is in the financial field, how to use large language models (LLMs) as data annotation tools to improve the efficiency and cost - effectiveness of data annotation. Specifically, the paper focuses on the task of extracting relationships in financial documents, which usually requires in - depth financial knowledge to understand complex terms and calculations, and these tasks are challenging for non - professional human annotators and are prone to inconsistent and inaccurate annotation results. Therefore, the paper explores whether LLMs can replace non - professional crowdsourcing workers and compares them with expert annotators to evaluate the effectiveness and reliability of LLMs in this specific area. The main contributions of the paper include: 1. For the first time in the financial field, it demonstrates the ability of LLMs as data annotation tools by comparing their performance with domain experts and non - professional crowdsourcing workers. 2. It compares three different LLMs (GPT - 4, PaLM 2 and MPT Instruct), as well as different parameter settings (such as temperature, random seed and prompt strategy) to determine the most accurate and reliable configuration. 3. It introduces a reliability index (LLM - RelIndex) to identify reliable samples and filter out samples that require human intervention. 4. It shows that LLMs can replace non - professional crowdsourcing workers in a large part of the data set, but for the remaining part, the intervention of experts is still required to ensure the accuracy of annotation. In addition, the paper also provides best - practice suggestions for implementing the LLMs annotation process.

Large Language Models as Financial Data Annotators: A Study on Effectiveness and Efficiency

Data-Centric Financial Large Language Models

Large Language Models for Data Annotation: A Survey

Large Language Model Adaptation for Financial Sentiment Analysis

A Survey of Large Language Models in Finance (FinLLMs)

Evaluating Large Language Models on Financial Report Summarization: An Empirical Study

Large Language Models in Finance: A Survey

Large Language Models as Annotators: Enhancing Generalization of NLP Models at Minimal Cost

Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications

Extracting Financial Data From Unstructured Sources: Leveraging Large Language Models

FinLLMs: A Framework for Financial Reasoning Dataset Generation with Large Language Models

Large Language Models for Data Annotation and Synthesis: A Survey

A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges

AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators

Leveraging Large Language Models to Democratize Access to Costly Financial Datasets for Academic Research

Large Language Model in Financial Regulatory Interpretation

LLMaAA: Making Large Language Models as Active Annotators

The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation

Enhancing Financial Sentiment Analysis via Retrieval Augmented Large Language Models

Large Language Models: An Applied Econometric Framework