Greenback Bears and Fiscal Hawks: Finance is a Jungle and Text Embeddings Must Adapt

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anderson, Peter, Janardhanan, Mano Vikash, He, Jason, Cheng, Wei, Flanagan, Charlie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909384085340160
author Anderson, Peter
Janardhanan, Mano Vikash
He, Jason
Cheng, Wei
Flanagan, Charlie
author_facet Anderson, Peter
Janardhanan, Mano Vikash
He, Jason
Cheng, Wei
Flanagan, Charlie
contents Financial documents are filled with specialized terminology, arcane jargon, and curious acronyms that pose challenges for general-purpose text embeddings. Yet, few text embeddings specialized for finance have been reported in the literature, perhaps in part due to a lack of public datasets and benchmarks. We present BAM embeddings, a set of text embeddings finetuned on a carefully constructed dataset of 14.3M query-passage pairs. Demonstrating the benefits of domain-specific training, BAM embeddings achieve Recall@1 of 62.8% on a held-out test set, vs. only 39.2% for the best general-purpose text embedding from OpenAI. Further, BAM embeddings increase question answering accuracy by 8% on FinanceBench and show increased sensitivity to the finance-specific elements that are found in detailed, forward-looking and company and date-specific queries. To support further research we describe our approach in detail, quantify the importance of hard negative mining and dataset scale.
format Preprint
id arxiv_https___arxiv_org_abs_2411_07142
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Greenback Bears and Fiscal Hawks: Finance is a Jungle and Text Embeddings Must Adapt
Anderson, Peter
Janardhanan, Mano Vikash
He, Jason
Cheng, Wei
Flanagan, Charlie
Computation and Language
Financial documents are filled with specialized terminology, arcane jargon, and curious acronyms that pose challenges for general-purpose text embeddings. Yet, few text embeddings specialized for finance have been reported in the literature, perhaps in part due to a lack of public datasets and benchmarks. We present BAM embeddings, a set of text embeddings finetuned on a carefully constructed dataset of 14.3M query-passage pairs. Demonstrating the benefits of domain-specific training, BAM embeddings achieve Recall@1 of 62.8% on a held-out test set, vs. only 39.2% for the best general-purpose text embedding from OpenAI. Further, BAM embeddings increase question answering accuracy by 8% on FinanceBench and show increased sensitivity to the finance-specific elements that are found in detailed, forward-looking and company and date-specific queries. To support further research we describe our approach in detail, quantify the importance of hard negative mining and dataset scale.
title Greenback Bears and Fiscal Hawks: Finance is a Jungle and Text Embeddings Must Adapt
topic Computation and Language
url https://arxiv.org/abs/2411.07142