Saved in:
| Main Authors: | Okocha, Chibuzor, Ezema, Kelechi, Grant, Christan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.21554 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can large audio language models understand child stuttering speech? speech summarization, and source separation
by: Okocha, Chibuzor, et al.
Published: (2025)
by: Okocha, Chibuzor, et al.
Published: (2025)
RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
AfroBench: How Good are Large Language Models on African Languages?
by: Ojo, Jessica, et al.
Published: (2023)
by: Ojo, Jessica, et al.
Published: (2023)
MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data
by: Driggers-Ellis, Christopher, et al.
Published: (2025)
by: Driggers-Ellis, Christopher, et al.
Published: (2025)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
by: Pattnaik, Pulkit, et al.
Published: (2024)
by: Pattnaik, Pulkit, et al.
Published: (2024)
On Creating an English-Thai Code-switched Machine Translation in Medical Domain
by: Pengpun, Parinthapat, et al.
Published: (2024)
by: Pengpun, Parinthapat, et al.
Published: (2024)
Label-semantics Aware Generative Approach for Domain-Agnostic Multilabel Classification
by: Khatuya, Subhendu, et al.
Published: (2025)
by: Khatuya, Subhendu, et al.
Published: (2025)
Improving Self-supervised Pre-training using Accent-Specific Codebooks
by: Prabhu, Darshan, et al.
Published: (2024)
by: Prabhu, Darshan, et al.
Published: (2024)
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
M3: A Multi-Task Mixed-Objective Learning Framework for Open-Domain Multi-Hop Dense Sentence Retrieval
by: Bai, Yang, et al.
Published: (2024)
by: Bai, Yang, et al.
Published: (2024)
ASR-Synchronized Speaker-Role Diarization
by: Ghosh, Arindam, et al.
Published: (2025)
by: Ghosh, Arindam, et al.
Published: (2025)
Domain-Shift-Aware Conformal Prediction for Large Language Models
by: Lin, Zhexiao, et al.
Published: (2025)
by: Lin, Zhexiao, et al.
Published: (2025)
Double Jeopardy and Climate Impact in the Use of Large Language Models: Socio-economic Disparities and Reduced Utility for Non-English Speakers
by: Solatorio, Aivin V., et al.
Published: (2024)
by: Solatorio, Aivin V., et al.
Published: (2024)
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
by: Goffinet, Etienne, et al.
Published: (2025)
by: Goffinet, Etienne, et al.
Published: (2025)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
Edeflip: Supervised Word Translation between English and Yoruba
by: Abioye, Ikeoluwa, et al.
Published: (2025)
by: Abioye, Ikeoluwa, et al.
Published: (2025)
Translating Hanja Historical Documents to Contemporary Korean and English
by: Son, Juhee, et al.
Published: (2022)
by: Son, Juhee, et al.
Published: (2022)
Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-training
by: Hassan, Muhammad Taimoor, et al.
Published: (2026)
by: Hassan, Muhammad Taimoor, et al.
Published: (2026)
Many-to-English Machine Translation Tools, Data, and Pretrained Models
by: Gowda, Thamme, et al.
Published: (2021)
by: Gowda, Thamme, et al.
Published: (2021)
Algorithmic Fairness Generalization under Covariate and Dependence Shifts Simultaneously
by: Zhao, Chen, et al.
Published: (2023)
by: Zhao, Chen, et al.
Published: (2023)
Lugha-Llama: Adapting Large Language Models for African Languages
by: Buzaaba, Happy, et al.
Published: (2025)
by: Buzaaba, Happy, et al.
Published: (2025)
IITK at SemEval-2024 Task 10: Who is the speaker? Improving Emotion Recognition and Flip Reasoning in Conversations via Speaker Embeddings
by: Patel, Shubham, et al.
Published: (2024)
by: Patel, Shubham, et al.
Published: (2024)
Keyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content
by: Rizvi, F. A., et al.
Published: (2025)
by: Rizvi, F. A., et al.
Published: (2025)
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
Enhancing Multilingual Sentiment Analysis with Explainability for Sinhala, English, and Code-Mixed Content
by: Rizvi, Azmarah, et al.
Published: (2025)
by: Rizvi, Azmarah, et al.
Published: (2025)
BgGPT 1.0: Extending English-centric LLMs to other languages
by: Alexandrov, Anton, et al.
Published: (2024)
by: Alexandrov, Anton, et al.
Published: (2024)
When Domains Interact: Asymmetric and Order-Sensitive Cross-Domain Effects in Reinforcement Learning for Reasoning
by: Yang, Wang, et al.
Published: (2026)
by: Yang, Wang, et al.
Published: (2026)
Noise-Aware Training of Layout-Aware Language Models
by: Sarkhel, Ritesh, et al.
Published: (2024)
by: Sarkhel, Ritesh, et al.
Published: (2024)
The Saturation Point of Backtranslation in High Quality Low Resource English Gujarati Machine Translation
by: Arif, Arwa
Published: (2025)
by: Arif, Arwa
Published: (2025)
Can General-Purpose Large Language Models Generalize to English-Thai Machine Translation ?
by: Chiaranaipanich, Jirat, et al.
Published: (2024)
by: Chiaranaipanich, Jirat, et al.
Published: (2024)
Automated Multi-Language to English Machine Translation Using Generative Pre-Trained Transformers
by: Pelofske, Elijah, et al.
Published: (2024)
by: Pelofske, Elijah, et al.
Published: (2024)
Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language
by: Sharma, Yash, et al.
Published: (2024)
by: Sharma, Yash, et al.
Published: (2024)
FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
by: Sarwar, Nobin, et al.
Published: (2025)
by: Sarwar, Nobin, et al.
Published: (2025)
Evaluating Embedding Frameworks for Scientific Domain
by: Ahmed, Nouman, et al.
Published: (2025)
by: Ahmed, Nouman, et al.
Published: (2025)
Utilizing GPT to Enhance Text Summarization: A Strategy to Minimize Hallucinations
by: Shakil, Hassan, et al.
Published: (2024)
by: Shakil, Hassan, et al.
Published: (2024)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
by: Rashidi, Sina, et al.
Published: (2025)
by: Rashidi, Sina, et al.
Published: (2025)
Constructing Synthetic Instruction Datasets for Improving Reasoning in Domain-Specific LLMs: A Case Study in the Japanese Financial Domain
by: Okochi, Yuma, et al.
Published: (2026)
by: Okochi, Yuma, et al.
Published: (2026)
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023)
by: Fan, Simin, et al.
Published: (2023)
Finetuning Large Language Models for Automated Depression Screening in Nigerian Pidgin English: GENSCORE Pilot Study
by: Olufadewa, Isaac Iyinoluwa, et al.
Published: (2025)
by: Olufadewa, Isaac Iyinoluwa, et al.
Published: (2025)
Similar Items
-
Can large audio language models understand child stuttering speech? speech summarization, and source separation
by: Okocha, Chibuzor, et al.
Published: (2025) -
RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering
by: Bai, Yang, et al.
Published: (2025) -
Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization
by: Li, Xiang, et al.
Published: (2024) -
AfroBench: How Good are Large Language Models on African Languages?
by: Ojo, Jessica, et al.
Published: (2023) -
MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data
by: Driggers-Ellis, Christopher, et al.
Published: (2025)