AlphaFin: Benchmarking Financial Analysis with Retrieval-Augmented Stock-Chain Framework

Xiang Li,Zhenyu Li,Chen Shi,Yong Xu,Qing Du,Mingkui Tan,Jun Huang,Wei Lin
2024-03-19
Abstract:The task of financial analysis primarily encompasses two key areas: stock trend prediction and the corresponding financial question answering. Currently, machine learning and deep learning algorithms (ML&DL) have been widely applied for stock trend predictions, leading to significant progress. However, these methods fail to provide reasons for predictions, lacking interpretability and reasoning processes. Also, they can not integrate textual information such as financial news or reports. Meanwhile, large language models (LLMs) have remarkable textual understanding and generation ability. But due to the scarcity of financial training datasets and limited integration with real-time knowledge, LLMs still suffer from hallucinations and are unable to keep up with the latest information. To tackle these challenges, we first release AlphaFin datasets, combining traditional research datasets, real-time financial data, and handwritten chain-of-thought (CoT) data. It has a positive impact on training LLMs for completing financial analysis. We then use AlphaFin datasets to benchmark a state-of-the-art method, called Stock-Chain, for effectively tackling the financial analysis task, which integrates retrieval-augmented generation (RAG) techniques. Extensive experiments are conducted to demonstrate the effectiveness of our framework on financial analysis.
Computation and Language
What problem does this paper attempt to address?
The paper aims to address two key issues in financial analysis: stock trend prediction and corresponding Financial Question Answering (FQA). Currently, machine learning (ML) and deep learning (DL) algorithms have made significant progress in stock trend prediction, but these methods lack interpretability and reasoning processes and cannot integrate textual information (such as financial news or reports). Additionally, although large-scale language models (LLMs) perform well in text understanding and generation, they still suffer from hallucination due to the lack of high-quality financial training datasets and real-time knowledge integration, making it difficult for them to keep up with the latest information changes. To address these issues, the authors first released the AlphaFin dataset, which combines traditional research datasets, real-time financial data, and handwritten Chain-of-Thought (CoT) data. These datasets help enhance the capabilities of LLMs in performing financial analysis tasks. Then, the authors benchmarked a method called Stock-Chain using the AlphaFin dataset, which integrates Retrieval-Augmented Generation (RAG) technology. Extensive experiments validated the effectiveness of this framework in financial analysis tasks. Specifically, Stock-Chain not only provides stock trend predictions but also integrates real-time market data and macroeconomic news through RAG, enabling accurate stock analysis during interactions with investors. Experimental results show that Stock-Chain achieves state-of-the-art accuracy in stock trend prediction tasks and exceeds an annualized return rate (ARR) of 30%.