News Reporter: A Multi-lingual LLM Framework for Broadcast T.V News

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jain, Tarun, Gao, Yufei, Vanga, Sridhar, Singla, Karan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929578570678272
author Jain, Tarun
Gao, Yufei
Vanga, Sridhar
Singla, Karan
author_facet Jain, Tarun
Gao, Yufei
Vanga, Sridhar
Singla, Karan
contents Large Language Models (LLMs) have fast become an essential tools to many conversational chatbots due to their ability to provide coherent answers for varied queries. Datasets used to train these LLMs are often a mix of generic and synthetic samples, thus lacking the verification needed to provide correct and verifiable answers for T.V. News. We collect and share a large collection of QA pairs extracted from transcripts of news recordings from various news-channels across the United States. Resultant QA pairs are then used to fine-tune an off-the-shelf LLM model. Our model surpasses base models of similar size on several open LLM benchmarks. We further integrate and propose a RAG method to improve contextualization of our answers and also point it to a verifiable news recording.
format Preprint
id arxiv_https___arxiv_org_abs_2410_07520
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle News Reporter: A Multi-lingual LLM Framework for Broadcast T.V News
Jain, Tarun
Gao, Yufei
Vanga, Sridhar
Singla, Karan
Computation and Language
Large Language Models (LLMs) have fast become an essential tools to many conversational chatbots due to their ability to provide coherent answers for varied queries. Datasets used to train these LLMs are often a mix of generic and synthetic samples, thus lacking the verification needed to provide correct and verifiable answers for T.V. News. We collect and share a large collection of QA pairs extracted from transcripts of news recordings from various news-channels across the United States. Resultant QA pairs are then used to fine-tune an off-the-shelf LLM model. Our model surpasses base models of similar size on several open LLM benchmarks. We further integrate and propose a RAG method to improve contextualization of our answers and also point it to a verifiable news recording.
title News Reporter: A Multi-lingual LLM Framework for Broadcast T.V News
topic Computation and Language
url https://arxiv.org/abs/2410.07520