LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kmainasi, Mohamed Bayan, Shahroor, Ali Ezzat, Hasanain, Maram, Laskar, Sahinur Rahman, Hassan, Naeemul, Alam, Firoj
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917938218401792
author Kmainasi, Mohamed Bayan
Shahroor, Ali Ezzat
Hasanain, Maram
Laskar, Sahinur Rahman
Hassan, Naeemul
Alam, Firoj
author_facet Kmainasi, Mohamed Bayan
Shahroor, Ali Ezzat
Hasanain, Maram
Laskar, Sahinur Rahman
Hassan, Naeemul
Alam, Firoj
contents Large Language Models (LLMs) have demonstrated remarkable success as general-purpose task solvers across various fields. However, their capabilities remain limited when addressing domain-specific problems, particularly in downstream NLP tasks. Research has shown that models fine-tuned on instruction-based downstream NLP datasets outperform those that are not fine-tuned. While most efforts in this area have primarily focused on resource-rich languages like English and broad domains, little attention has been given to multilingual settings and specific domains. To address this gap, this study focuses on developing a specialized LLM, LlamaLens, for analyzing news and social media content in a multilingual context. To the best of our knowledge, this is the first attempt to tackle both domain specificity and multilinguality, with a particular focus on news and social media. Our experimental setup includes 18 tasks, represented by 52 datasets covering Arabic, English, and Hindi. We demonstrate that LlamaLens outperforms the current state-of-the-art (SOTA) on 23 testing sets, and achieves comparable performance on 8 sets. We make the models and resources publicly available for the research community (https://huggingface.co/collections/QCRI/llamalens-672f7e0604a0498c6a2f0fe9).
format Preprint
id arxiv_https___arxiv_org_abs_2410_15308
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content
Kmainasi, Mohamed Bayan
Shahroor, Ali Ezzat
Hasanain, Maram
Laskar, Sahinur Rahman
Hassan, Naeemul
Alam, Firoj
Computation and Language
Artificial Intelligence
68T50
F.2.2; I.2.7
Large Language Models (LLMs) have demonstrated remarkable success as general-purpose task solvers across various fields. However, their capabilities remain limited when addressing domain-specific problems, particularly in downstream NLP tasks. Research has shown that models fine-tuned on instruction-based downstream NLP datasets outperform those that are not fine-tuned. While most efforts in this area have primarily focused on resource-rich languages like English and broad domains, little attention has been given to multilingual settings and specific domains. To address this gap, this study focuses on developing a specialized LLM, LlamaLens, for analyzing news and social media content in a multilingual context. To the best of our knowledge, this is the first attempt to tackle both domain specificity and multilinguality, with a particular focus on news and social media. Our experimental setup includes 18 tasks, represented by 52 datasets covering Arabic, English, and Hindi. We demonstrate that LlamaLens outperforms the current state-of-the-art (SOTA) on 23 testing sets, and achieves comparable performance on 8 sets. We make the models and resources publicly available for the research community (https://huggingface.co/collections/QCRI/llamalens-672f7e0604a0498c6a2f0fe9).
title LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content
topic Computation and Language
Artificial Intelligence
68T50
F.2.2; I.2.7
url https://arxiv.org/abs/2410.15308