Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gharami, Kanchon, Muhtaseem, Quazi Sarwar, Gupta, Deepti, Elluri, Lavanya, Moni, Shafika Showkat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Efficient Privacy-preserving Intrusion Detection Scheme for UAV Swarm Networks
von: Gharami, Kanchon, et al.
Veröffentlicht: (2025)
von: Gharami, Kanchon, et al.
Veröffentlicht: (2025)
Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation
von: Gharami, Kanchon, et al.
Veröffentlicht: (2025)
von: Gharami, Kanchon, et al.
Veröffentlicht: (2025)
Blockchain-Enhanced Framework for Secure Third-Party Vendor Risk Management and Vigilant Security Controls
von: Gupta, Deepti, et al.
Veröffentlicht: (2024)
von: Gupta, Deepti, et al.
Veröffentlicht: (2024)
ChatGPT: Excellent Paper! Accept It. Editor: Imposter Found! Review Rejected
von: Gharami, Kanchon, et al.
Veröffentlicht: (2025)
von: Gharami, Kanchon, et al.
Veröffentlicht: (2025)
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
von: Bhowmik, Shimanto, et al.
Veröffentlicht: (2025)
von: Bhowmik, Shimanto, et al.
Veröffentlicht: (2025)
Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation
von: Gupta, Ashray, et al.
Veröffentlicht: (2025)
von: Gupta, Ashray, et al.
Veröffentlicht: (2025)
Low-Resource Counterspeech Generation for Indic Languages: The Case of Bengali and Hindi
von: Das, Mithun, et al.
Veröffentlicht: (2024)
von: Das, Mithun, et al.
Veröffentlicht: (2024)
HindiLLM: Large Language Model for Hindi
von: Chouhan, Sanjay, et al.
Veröffentlicht: (2024)
von: Chouhan, Sanjay, et al.
Veröffentlicht: (2024)
TriNER: A Series of Named Entity Recognition Models For Hindi, Bengali & Marathi
von: Dhamaskar, Mohammed Amaan, et al.
Veröffentlicht: (2025)
von: Dhamaskar, Mohammed Amaan, et al.
Veröffentlicht: (2025)
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
Microbenchmarking Cloud Cryptographic Workloads for Privacy-Preserving Healthcare IoT
von: Webb, Jeremiah L., et al.
Veröffentlicht: (2026)
von: Webb, Jeremiah L., et al.
Veröffentlicht: (2026)
Privacy-Preserving Data Sharing in Agriculture: Enforcing Policy Rules for Secure and Confidential Data Synthesis
von: Kotal, Anantaa, et al.
Veröffentlicht: (2023)
von: Kotal, Anantaa, et al.
Veröffentlicht: (2023)
BTPD: A Multilingual Hand-curated Dataset of Bengali Transnational Political Discourse Across Online Communities
von: Das, Dipto, et al.
Veröffentlicht: (2025)
von: Das, Dipto, et al.
Veröffentlicht: (2025)
Building pre-train LLM Dataset for the INDIC Languages: a case study on Hindi
von: Parida, Shantipriya, et al.
Veröffentlicht: (2024)
von: Parida, Shantipriya, et al.
Veröffentlicht: (2024)
PrivComp-KG : Leveraging Knowledge Graph and Large Language Models for Privacy Policy Compliance Verification
von: Garza, Leon, et al.
Veröffentlicht: (2024)
von: Garza, Leon, et al.
Veröffentlicht: (2024)
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
von: Karthika, N J, et al.
Veröffentlicht: (2026)
von: Karthika, N J, et al.
Veröffentlicht: (2026)
Bengali Fake Reviews: A Benchmark Dataset and Detection System
von: Shahariar, G. M., et al.
Veröffentlicht: (2023)
von: Shahariar, G. M., et al.
Veröffentlicht: (2023)
Airavata: Introducing Hindi Instruction-tuned LLM
von: Gala, Jay, et al.
Veröffentlicht: (2024)
von: Gala, Jay, et al.
Veröffentlicht: (2024)
Breaking Language Barriers: A Question Answering Dataset for Hindi and Marathi
von: Sabane, Maithili, et al.
Veröffentlicht: (2023)
von: Sabane, Maithili, et al.
Veröffentlicht: (2023)
TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2025)
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2025)
BNLI: A Linguistically-Refined Bengali Dataset for Natural Language Inference
von: Haque, Farah Binta, et al.
Veröffentlicht: (2025)
von: Haque, Farah Binta, et al.
Veröffentlicht: (2025)
Romanization Encoding For Multilingual ASR
von: Ding, Wen, et al.
Veröffentlicht: (2024)
von: Ding, Wen, et al.
Veröffentlicht: (2024)
BengaliSent140: A Large-Scale Bengali Binary Sentiment Dataset for Hate and Non-Hate Speech Classification
von: Islam, Akif, et al.
Veröffentlicht: (2026)
von: Islam, Akif, et al.
Veröffentlicht: (2026)
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans
von: Sakai, Yusuke, et al.
Veröffentlicht: (2024)
von: Sakai, Yusuke, et al.
Veröffentlicht: (2024)
Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi
von: Acharya, Arkadeep, et al.
Veröffentlicht: (2024)
von: Acharya, Arkadeep, et al.
Veröffentlicht: (2024)
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5
von: Acharya, Arkadeep, et al.
Veröffentlicht: (2024)
von: Acharya, Arkadeep, et al.
Veröffentlicht: (2024)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
von: Patwary, Firoj Ahmmed, et al.
Veröffentlicht: (2025)
von: Patwary, Firoj Ahmmed, et al.
Veröffentlicht: (2025)
HiMed: Incentivizing Hindi Reasoning in Medical LLMs
von: Jiang, Dingfeng, et al.
Veröffentlicht: (2026)
von: Jiang, Dingfeng, et al.
Veröffentlicht: (2026)
Motamot: A Dataset for Revealing the Supremacy of Large Language Models over Transformer Models in Bengali Political Sentiment Analysis
von: Faria, Fatema Tuj Johora, et al.
Veröffentlicht: (2024)
von: Faria, Fatema Tuj Johora, et al.
Veröffentlicht: (2024)
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
von: Ridoy, Shahriyar Zaman, et al.
Veröffentlicht: (2025)
von: Ridoy, Shahriyar Zaman, et al.
Veröffentlicht: (2025)
Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement
von: Hasan, Mohammed Rakibul, et al.
Veröffentlicht: (2025)
von: Hasan, Mohammed Rakibul, et al.
Veröffentlicht: (2025)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
von: Kabir, Daeen, et al.
Veröffentlicht: (2025)
von: Kabir, Daeen, et al.
Veröffentlicht: (2025)
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
von: Kamath, Anusha, et al.
Veröffentlicht: (2025)
von: Kamath, Anusha, et al.
Veröffentlicht: (2025)
Leveraging LLM For Synchronizing Information Across Multilingual Tables
von: Khincha, Siddharth, et al.
Veröffentlicht: (2025)
von: Khincha, Siddharth, et al.
Veröffentlicht: (2025)
IPA Transcription of Bengali Texts
von: Fatema, Kanij, et al.
Veröffentlicht: (2024)
von: Fatema, Kanij, et al.
Veröffentlicht: (2024)
Systematic Evaluation of Long-Context LLMs on Financial Concepts
von: Gupta, Lavanya, et al.
Veröffentlicht: (2024)
von: Gupta, Lavanya, et al.
Veröffentlicht: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
GHTM: A Graph-based Hybrid Topic Modeling Approach with a Benchmark Dataset for the Low-Resource Bengali Language
von: Haque, Farhana, et al.
Veröffentlicht: (2025)
von: Haque, Farhana, et al.
Veröffentlicht: (2025)
BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis
von: Alam, Sadia, et al.
Veröffentlicht: (2024)
von: Alam, Sadia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
An Efficient Privacy-preserving Intrusion Detection Scheme for UAV Swarm Networks
von: Gharami, Kanchon, et al.
Veröffentlicht: (2025) -
Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation
von: Gharami, Kanchon, et al.
Veröffentlicht: (2025) -
Blockchain-Enhanced Framework for Secure Third-Party Vendor Risk Management and Vigilant Security Controls
von: Gupta, Deepti, et al.
Veröffentlicht: (2024) -
ChatGPT: Excellent Paper! Accept It. Editor: Imposter Found! Review Rejected
von: Gharami, Kanchon, et al.
Veröffentlicht: (2025) -
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
von: Bhowmik, Shimanto, et al.
Veröffentlicht: (2025)