Won: Establishing Best Practices for Korean Financial NLP
₩ON: First open leaderboard for Korean financial LLMs, evaluating 1,119 submissions.
Key Findings
Methodology
The study created a finance benchmark with 5.5k multiple-choice questions covering five key areas: financial markets, finance and accounting, domestic company analysis, financial agents, and stock price prediction. Additionally, it included 100 open-ended QA tasks. Over two months, an open leaderboard evaluated over 1,000 model submissions, resulting in an 80k-instance open instruction dataset.
Key Results
- During the leaderboard operation, 1,119 model submissions were evaluated, with over 600 models publicly available, laying the groundwork for future research.
- In the finance and accounting category, models showed significant improvement, with top scores reaching 0.83.
- In open-ended FinQA tasks, the ₩ON model excelled, demonstrating strong reasoning capabilities.
Significance
This study advances the development of Korean financial LLMs by creating an open leaderboard and dataset. It provides a valuable platform for evaluation and promotes transparency and openness in the field, benefiting both academia and industry.
Technical Contribution
Technical contributions include developing the first open leaderboard focused on the Korean financial domain, providing a comprehensive benchmark dataset, and introducing the ₩ON model, which enhances financial task performance through improved reasoning capabilities.
Novelty
This study is the first to introduce an open leaderboard for evaluating Korean financial LLMs, promoting transparency and development in the field through open datasets and models.
Limitations
- The multiple-choice questions used in the leaderboard may not fully reflect real-world prompts.
- The limited number of open-ended QA tasks may not comprehensively evaluate the models' reasoning capabilities.
Future Work
Future work includes expanding the range of open-ended QA tasks, improving models' reasoning capabilities, and exploring more applications in the financial domain.
AI Executive Summary
The field of natural language processing in Korean finance has been hindered by the closed nature of the industry, limiting model and dataset sharing. To address this gap, the research team created an open leaderboard, evaluating 1,119 model submissions across five key financial areas. Through the leaderboard's operation, a large collection of models and data was gathered, resulting in an 80k-instance open instruction dataset.
The introduction of the ₩ON model marks a significant milestone in the Korean financial LLM domain. By incorporating effective training strategies observed on the leaderboard, ₩ON demonstrates outstanding performance in finance and accounting and open-ended FinQA tasks through enhanced reasoning capabilities.
Despite these achievements, the study faces limitations, such as the multiple-choice questions not fully reflecting real-world scenarios. Future research will focus on expanding the range of open-ended QA tasks and further enhancing the models' reasoning capabilities to better serve practical applications in the financial sector.
Deep Analysis
Background
Natural language processing in the financial sector has seen significant advancements, but the closed nature of the industry limits model and dataset sharing, slowing technological progress and leading to duplicated efforts. Existing evaluation tools like KRX-Bench and KMMLU provide some standards but fall short of covering the broad potential applications in finance.
Core Problem
Large language models in finance face performance instability, potentially leading to financial losses. Developing reliable evaluation systems is crucial, yet the closed nature of the financial industry restricts model and dataset sharing, hindering the development of relevant techniques.
Innovation
The core innovation of this study is the creation of an open leaderboard for evaluating Korean financial LLMs. By collecting and analyzing model submissions on the leaderboard, the research team summarized effective training strategies and introduced the ₩ON model.
Methodology
- �� Created a finance benchmark with 5.5k multiple-choice questions covering five key areas.
- �� Operated a two-month open leaderboard, evaluating 1,119 model submissions.
- �� Collected and filtered 200k instances from competing teams, resulting in a high-quality 80k-instance instruction dataset.
- �� Used Deepseek-R1 to generate responses and trained on these trajectories to release the ₩ON model.
Experiments
The experimental design included evaluating multiple financial domain datasets, such as financial markets, finance and accounting, and domestic company analysis. The benchmark dataset included 5.5k multiple-choice questions and 100 open-ended QA tasks. Evaluation metrics included model accuracy and reasoning capabilities across different tasks.
Results
During the leaderboard operation, 1,119 model submissions were evaluated, with over 600 models publicly available. The ₩ON model excelled in the finance and accounting category, with top scores reaching 0.83. In open-ended FinQA tasks, ₩ON demonstrated strong reasoning capabilities.
Applications
Application scenarios include financial market analysis, financial statement interpretation, and stock price prediction. The model can help financial institutions improve decision-making efficiency and reduce risks.
Limitations & Outlook
Despite significant progress, the multiple-choice questions may not fully reflect real-world scenarios. Additionally, the limited number of open-ended QA tasks may not comprehensively evaluate the models' reasoning capabilities. Future research will focus on expanding the range of open-ended QA tasks and further enhancing the models' reasoning capabilities.
Plain Language Accessible to non-experts
Imagine you're playing a financial market simulation game, and ₩ON is your smart assistant. It helps you analyze market trends, predict stock prices, and answer your questions about financial statements. By learning from a vast amount of financial data and questions, it can quickly provide accurate advice, just like an experienced investment advisor.
ELI14 Explained like you're 14
Imagine you're playing a financial game, and ₩ON is your super assistant! It helps you predict stock prices, analyze market trends, just like a smart robot friend. You ask it questions, and it gives you answers super fast, helping you make the best decisions in the game. Isn't that cool?
Glossary
Multiple-Choice Question (MCQA)
A test format offering multiple options to choose from, typically used to assess knowledge or ability.
Used in the study to evaluate models' financial knowledge.
Open Leaderboard
A public evaluation platform allowing multiple model submissions and performance comparisons.
Used to evaluate Korean financial LLMs' performance.
₩ON Model
An open large language model focused on the Korean financial domain, incorporating best practices from the leaderboard.
The model introduced in the study, demonstrating strong reasoning capabilities.
Deepseek-R1
An algorithm used to generate model responses, aiding in training reasoning capabilities.
Used to generate training data for the ₩ON model.
Reasoning Capability
The ability of a model to understand and process complex problems, typically achieved through multi-step logical deductions.
Demonstrated by the ₩ON model in open-ended FinQA tasks.
Open Questions Unanswered questions from this research
- 1 How to expand the range of open-ended QA tasks to more comprehensively evaluate models' reasoning capabilities?
- 2 How to improve model performance in knowledge-intensive tasks?
Applications
Immediate Applications
Financial Market Analysis
The model can be used to analyze market trends, helping investors make more informed decisions.
Long-term Vision
Intelligent Financial Assistant
In the future, the model is expected to become an intelligent assistant for financial institutions, improving decision-making efficiency and reducing risks.
Abstract
In this work, we present the first open leaderboard for evaluating Korean large language models focused on finance. Operated for about eight weeks, the leaderboard evaluated 1,119 submissions on a closed benchmark covering five MCQA categories: finance and accounting, stock price prediction, domestic company analysis, financial markets, and financial agent tasks and one open-ended qa task. Building on insights from these evaluations, we release an open instruction dataset of 80k instances and summarize widely used training strategies observed among top-performing models. Finally, we introduce Won, a fully open and transparent LLM built using these best practices. We hope our contributions help advance the development of better and safer financial LLMs for Korean and other languages.