Phonetically rich corpus construction for a low-resourced language
Fuente:
arXiv
Guardado en:
| Autores principales: | Amadeus, Marcellus, Castañeda, William Alberto Cruz, Lobato, Wilmer, Aquino, Niasche |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluation Metrics for Text Data Augmentation in NLP
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
Amadeus-Verbo Technical Report: The powerful Qwen2.5 family models trained in Portuguese
por: Cruz-Castañeda, William Alberto, et al.
Publicado: (2025)
por: Cruz-Castañeda, William Alberto, et al.
Publicado: (2025)
Image captioning for Brazilian Portuguese using GRIT model
por: de Alencar, Rafael Silva, et al.
Publicado: (2024)
por: de Alencar, Rafael Silva, et al.
Publicado: (2024)
From Pampas to Pixels: Fine-Tuning Diffusion Models for Gaúcho Heritage
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
An Inpainting-Infused Pipeline for Attire and Background Replacement
por: Perche-Mahlow, Felipe Rodrigues, et al.
Publicado: (2024)
por: Perche-Mahlow, Felipe Rodrigues, et al.
Publicado: (2024)
Jabuticaba: The largest commercial corpus for LLMs in Portuguese
por: Amadeus, Marcellus, et al.
Publicado: (2025)
por: Amadeus, Marcellus, et al.
Publicado: (2025)
Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
por: Burda-Lassen, Olena
Publicado: (2024)
por: Burda-Lassen, Olena
Publicado: (2024)
Leveraging LLMs for MT in Crisis Scenarios: a blueprint for low-resource languages
por: Lankford, Séamus, et al.
Publicado: (2024)
por: Lankford, Séamus, et al.
Publicado: (2024)
Multilingual jailbreaking of LLMs using low-resource languages
por: Marx, Dylan, et al.
Publicado: (2026)
por: Marx, Dylan, et al.
Publicado: (2026)
NER- RoBERTa: Fine-Tuning RoBERTa for Named Entity Recognition (NER) within low-resource languages
por: Abdullah, Abdulhady Abas, et al.
Publicado: (2024)
por: Abdullah, Abdulhady Abas, et al.
Publicado: (2024)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
por: Aswal, Darpan, et al.
Publicado: (2025)
por: Aswal, Darpan, et al.
Publicado: (2025)
Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation
por: Han, Lifeng, et al.
Publicado: (2020)
por: Han, Lifeng, et al.
Publicado: (2020)
Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers
por: Dhakal, Prakash, et al.
Publicado: (2024)
por: Dhakal, Prakash, et al.
Publicado: (2024)
Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
por: Fucci, Dennis, et al.
Publicado: (2025)
por: Fucci, Dennis, et al.
Publicado: (2025)
PERCORE: A Deep Learning-Based Framework for Persian Spelling Correction with Phonetic Analysis
por: Dashti, Seyed Mohammad Sadegh, et al.
Publicado: (2024)
por: Dashti, Seyed Mohammad Sadegh, et al.
Publicado: (2024)
LLMs Know More Than Words: A Genre Study with Syntax, Metaphor & Phonetics
por: Shi, Weiye, et al.
Publicado: (2025)
por: Shi, Weiye, et al.
Publicado: (2025)
Assessment and manipulation of latent constructs in pre-trained language models using psychometric scales
por: Reuben, Maor, et al.
Publicado: (2024)
por: Reuben, Maor, et al.
Publicado: (2024)
Speak & Spell: LLM-Driven Controllable Phonetic Error Augmentation for Robust Dialogue State Tracking
por: Lee, Jihyun, et al.
Publicado: (2024)
por: Lee, Jihyun, et al.
Publicado: (2024)
Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context
por: Trinh, Viet Anh, et al.
Publicado: (2025)
por: Trinh, Viet Anh, et al.
Publicado: (2025)
Using Large Language Models for education managements in Vietnamese with low resources
por: Minh, Duc Do, et al.
Publicado: (2025)
por: Minh, Duc Do, et al.
Publicado: (2025)
Do LLM hallucination detectors suffer from low-resource effect?
por: Datta, Debtanu, et al.
Publicado: (2026)
por: Datta, Debtanu, et al.
Publicado: (2026)
Normalization through Fine-tuning: Understanding Wav2vec 2.0 Embeddings for Phonetic Analysis
por: Wang, Yiming, et al.
Publicado: (2025)
por: Wang, Yiming, et al.
Publicado: (2025)
Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service
por: Merx, Raphael, et al.
Publicado: (2024)
por: Merx, Raphael, et al.
Publicado: (2024)
The Lucie-7B LLM and the Lucie Training Dataset: Open resources for multilingual language generation
por: Gouvert, Olivier, et al.
Publicado: (2025)
por: Gouvert, Olivier, et al.
Publicado: (2025)
Phonetically-Augmented Discriminative Rescoring for Voice Search Error Correction
por: Van Gysel, Christophe, et al.
Publicado: (2025)
por: Van Gysel, Christophe, et al.
Publicado: (2025)
From RAGs to rich parameters: Probing how language models utilize external knowledge over parametric information for factual queries
por: Wadhwa, Hitesh, et al.
Publicado: (2024)
por: Wadhwa, Hitesh, et al.
Publicado: (2024)
Bringing legal knowledge to the public by constructing a legal question bank using large-scale pre-trained language model
por: Yuan, Mingruo, et al.
Publicado: (2025)
por: Yuan, Mingruo, et al.
Publicado: (2025)
TechGPT-2.0: A large language model project to solve the task of knowledge graph construction
por: Wang, Jiaqi, et al.
Publicado: (2024)
por: Wang, Jiaqi, et al.
Publicado: (2024)
Othering and low status framing of immigrant cuisines in US restaurant reviews and large language models
por: Luo, Yiwei, et al.
Publicado: (2023)
por: Luo, Yiwei, et al.
Publicado: (2023)
Tenyidie Syllabification corpus creation and deep learning applications
por: Angami, Teisovi, et al.
Publicado: (2025)
por: Angami, Teisovi, et al.
Publicado: (2025)
$π$-yalli: un nouveau corpus pour le nahuatl
por: Torres-Moreno, Juan-Manuel, et al.
Publicado: (2024)
por: Torres-Moreno, Juan-Manuel, et al.
Publicado: (2024)
Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
por: Gogoi, Parismita, et al.
Publicado: (2025)
por: Gogoi, Parismita, et al.
Publicado: (2025)
Natural language processing for African languages
por: Adelani, David Ifeoluwa
Publicado: (2025)
por: Adelani, David Ifeoluwa
Publicado: (2025)
Pre-training LLMs using human-like development data corpus
por: Bhardwaj, Khushi, et al.
Publicado: (2023)
por: Bhardwaj, Khushi, et al.
Publicado: (2023)
Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation
por: Roh, Jaechul, et al.
Publicado: (2025)
por: Roh, Jaechul, et al.
Publicado: (2025)
Dissociating language and thought in large language models
por: Mahowald, Kyle, et al.
Publicado: (2023)
por: Mahowald, Kyle, et al.
Publicado: (2023)
ESG-FTSE: A corpus of news articles with ESG relevance labels and use cases
por: Pavlova, Mariya, et al.
Publicado: (2024)
por: Pavlova, Mariya, et al.
Publicado: (2024)
Do language models practice what they preach? Examining language ideologies about gendered language reform encoded in LLMs
por: Watson, Julia, et al.
Publicado: (2024)
por: Watson, Julia, et al.
Publicado: (2024)
Assessing the potential of LLM-assisted annotation for corpus-based pragmatics and discourse analysis: The case of apology
por: Yu, Danni, et al.
Publicado: (2023)
por: Yu, Danni, et al.
Publicado: (2023)
FRACCO: A gold-standard annotated corpus of oncological entities with ICD-O-3.1 normalisation
por: Pignat, Johann, et al.
Publicado: (2025)
por: Pignat, Johann, et al.
Publicado: (2025)
Ejemplares similares
-
Evaluation Metrics for Text Data Augmentation in NLP
por: Amadeus, Marcellus, et al.
Publicado: (2024) -
Amadeus-Verbo Technical Report: The powerful Qwen2.5 family models trained in Portuguese
por: Cruz-Castañeda, William Alberto, et al.
Publicado: (2025) -
Image captioning for Brazilian Portuguese using GRIT model
por: de Alencar, Rafael Silva, et al.
Publicado: (2024) -
From Pampas to Pixels: Fine-Tuning Diffusion Models for Gaúcho Heritage
por: Amadeus, Marcellus, et al.
Publicado: (2024) -
An Inpainting-Infused Pipeline for Attire and Background Replacement
por: Perche-Mahlow, Felipe Rodrigues, et al.
Publicado: (2024)