FLEURS-Kobani: Extending the FLEURS Dataset for Northern Kurdish
Fuente:
arXiv
Salvato in:
| Autori principali: | Jaff, Daban Q., Mohammadamini, Mohammad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization
di: Mohammadamini, Mohammad, et al.
Pubblicazione: (2026)
di: Mohammadamini, Mohammad, et al.
Pubblicazione: (2026)
From Consensus to Split Decisions: ABC-Stratified Sentiment in Holocaust Oral Histories
di: Jaff, Daban Q.
Pubblicazione: (2026)
di: Jaff, Daban Q.
Pubblicazione: (2026)
Language and Speech Technology for Central Kurdish Varieties
di: Ahmadi, Sina, et al.
Pubblicazione: (2024)
di: Ahmadi, Sina, et al.
Pubblicazione: (2024)
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
di: Yan, Brian, et al.
Pubblicazione: (2025)
di: Yan, Brian, et al.
Pubblicazione: (2025)
FLEURS-ASL: Including American Sign Language in Massively Multilingual Multitask Evaluation
di: Tanzer, Garrett
Pubblicazione: (2024)
di: Tanzer, Garrett
Pubblicazione: (2024)
FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks
di: Ma, Min, et al.
Pubblicazione: (2024)
di: Ma, Min, et al.
Pubblicazione: (2024)
Named Entity Recognition for the Kurdish Sorani Language: Dataset Creation and Comparative Analysis
di: Abdalla, Bakhtawar, et al.
Pubblicazione: (2025)
di: Abdalla, Bakhtawar, et al.
Pubblicazione: (2025)
LES FLEURS DU MAL ANTES DE AS FLORES DO MAL: OS PRIMEIRÍSSIMOS BAUDELAIRIANOS
di: Ricardo Meirelles
Pubblicazione: (2018)
di: Ricardo Meirelles
Pubblicazione: (2018)
End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2024)
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2024)
Idiom Detection in Sorani Kurdish Texts
di: Omer, Skala Kamaran, et al.
Pubblicazione: (2025)
di: Omer, Skala Kamaran, et al.
Pubblicazione: (2025)
Subword Tokenization Strategies for Kurdish Word Embeddings
di: Salehi, Ali, et al.
Pubblicazione: (2025)
di: Salehi, Ali, et al.
Pubblicazione: (2025)
KurdSTS: The Kurdish Semantic Textual Similarity
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2025)
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2025)
Automatic Text Summarization (ATS) for Research Documents in Sorani Kurdish
di: Abdulrahman, Rondik Hadi, et al.
Pubblicazione: (2025)
di: Abdulrahman, Rondik Hadi, et al.
Pubblicazione: (2025)
Domain-Specific Machine Translation to Translate Medicine Brochures in English to Sorani Kurdish
di: Shamal, Mariam, et al.
Pubblicazione: (2025)
di: Shamal, Mariam, et al.
Pubblicazione: (2025)
Making Old Kurdish Publications Processable by Augmenting Available Optical Character Recognition Engines
di: Yaseen, Blnd, et al.
Pubblicazione: (2024)
di: Yaseen, Blnd, et al.
Pubblicazione: (2024)
An In-Depth Investigation of Data Collection in LLM App Ecosystems
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
A Comprehensive Part-of-Speech Tagging to Standardize Central-Kurdish Language: A Research Guide for Kurdish Natural Language Processing Tasks
di: Sabr, Shadan Shukr, et al.
Pubblicazione: (2025)
di: Sabr, Shadan Shukr, et al.
Pubblicazione: (2025)
KuBERT: Central Kurdish BERT Model and Its Application for Sentiment Analysis
di: Awlla, Kozhin muhealddin, et al.
Pubblicazione: (2025)
di: Awlla, Kozhin muhealddin, et al.
Pubblicazione: (2025)
Which one Performs Better? Wav2Vec or Whisper? Applying both in Badini Kurdish Speech to Text (BKSTT)
di: Adnan, Renas, et al.
Pubblicazione: (2025)
di: Adnan, Renas, et al.
Pubblicazione: (2025)
Where Are You From? Let Me Guess! Subdialect Recognition of Speeches in Sorani Kurdish
di: Isam, Sana, et al.
Pubblicazione: (2024)
di: Isam, Sana, et al.
Pubblicazione: (2024)
GRDD+: An Extended Greek Dialectal Dataset with Cross-Architecture Fine-tuning Evaluation
di: Chatzikyriakidis, Stergios, et al.
Pubblicazione: (2025)
di: Chatzikyriakidis, Stergios, et al.
Pubblicazione: (2025)
Extending Czech Aspect-Based Sentiment Analysis with Opinion Terms: Dataset and LLM Benchmarks
di: Šmíd, Jakub, et al.
Pubblicazione: (2026)
di: Šmíd, Jakub, et al.
Pubblicazione: (2026)
Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2024)
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2024)
The Kurdish Child in America: A Handbook for Educators.
di: Rytterager, Linda
Pubblicazione: (1993)
di: Rytterager, Linda
Pubblicazione: (1993)
Extended Japanese Commonsense Morality Dataset with Masked Token and Label Enhancement
di: Ohashi, Takumi, et al.
Pubblicazione: (2024)
di: Ohashi, Takumi, et al.
Pubblicazione: (2024)
Mining Mental Health Signals: A Comparative Study of Four Machine Learning Methods for Depression Detection from Social Media Posts in Sorani Kurdish
di: Mohammed, Idrees, et al.
Pubblicazione: (2025)
di: Mohammed, Idrees, et al.
Pubblicazione: (2025)
Persian Abstract Meaning Representation: Annotation Guidelines and Gold Standard Dataset
di: Takhshid, Reza, et al.
Pubblicazione: (2022)
di: Takhshid, Reza, et al.
Pubblicazione: (2022)
Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and Competence
di: Ayesh, Mutaz, et al.
Pubblicazione: (2026)
di: Ayesh, Mutaz, et al.
Pubblicazione: (2026)
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets
di: Zakizadeh, Mahdi, et al.
Pubblicazione: (2025)
di: Zakizadeh, Mahdi, et al.
Pubblicazione: (2025)
Bengali Fake Reviews: A Benchmark Dataset and Detection System
di: Shahariar, G. M., et al.
Pubblicazione: (2023)
di: Shahariar, G. M., et al.
Pubblicazione: (2023)
Noor-Ghateh: A Benchmark Dataset for Evaluating Arabic Word Segmenters in Hadith Domain
di: AlShuhayeb, Huda, et al.
Pubblicazione: (2023)
di: AlShuhayeb, Huda, et al.
Pubblicazione: (2023)
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
HamRaz: A Culture-Based Persian Conversation Dataset for Person-Centered Therapy Using LLM Agents
di: Abbasi, Mohammad Amin, et al.
Pubblicazione: (2025)
di: Abbasi, Mohammad Amin, et al.
Pubblicazione: (2025)
ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development
di: Huy, Ta Duc, et al.
Pubblicazione: (2023)
di: Huy, Ta Duc, et al.
Pubblicazione: (2023)
Building a Multivariate Time Series Benchmarking Datasets Inspired by Natural Language Processing (NLP)
di: Mustafa, Mohammad Asif Ibna, et al.
Pubblicazione: (2024)
di: Mustafa, Mohammad Asif Ibna, et al.
Pubblicazione: (2024)
Zero-Shot Commonsense Validation and Reasoning with Large Language Models: An Evaluation on SemEval-2020 Task 4 Dataset
di: Alfugaha, Rawand, et al.
Pubblicazione: (2025)
di: Alfugaha, Rawand, et al.
Pubblicazione: (2025)
A Benchmark Dataset and Evaluation Framework for Vietnamese Large Language Models in Customer Support
di: Nguyen, Long S. T., et al.
Pubblicazione: (2025)
di: Nguyen, Long S. T., et al.
Pubblicazione: (2025)
On Extending Direct Preference Optimization to Accommodate Ties
di: Chen, Jinghong, et al.
Pubblicazione: (2024)
di: Chen, Jinghong, et al.
Pubblicazione: (2024)
Extending LLMs' Context Window with 100 Samples
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
di: Su, Dan, et al.
Pubblicazione: (2024)
di: Su, Dan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization
di: Mohammadamini, Mohammad, et al.
Pubblicazione: (2026) -
From Consensus to Split Decisions: ABC-Stratified Sentiment in Holocaust Oral Histories
di: Jaff, Daban Q.
Pubblicazione: (2026) -
Language and Speech Technology for Central Kurdish Varieties
di: Ahmadi, Sina, et al.
Pubblicazione: (2024) -
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
di: Yan, Brian, et al.
Pubblicazione: (2025) -
FLEURS-ASL: Including American Sign Language in Massively Multilingual Multitask Evaluation
di: Tanzer, Garrett
Pubblicazione: (2024)