Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Burda-Lassen, Olena |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Phonetically rich corpus construction for a low-resourced language
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
Tenyidie Syllabification corpus creation and deep learning applications
von: Angami, Teisovi, et al.
Veröffentlicht: (2025)
von: Angami, Teisovi, et al.
Veröffentlicht: (2025)
How Culturally Aware are Vision-Language Models?
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
Large language models for folktale type automation based on motifs: Cinderella case study
von: Arčon, Tjaša, et al.
Veröffentlicht: (2025)
von: Arčon, Tjaša, et al.
Veröffentlicht: (2025)
Automated evaluation of LLMs for effective machine translation of Mandarin Chinese to English
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation
von: Han, Lifeng, et al.
Veröffentlicht: (2020)
von: Han, Lifeng, et al.
Veröffentlicht: (2020)
$π$-yalli: un nouveau corpus pour le nahuatl
von: Torres-Moreno, Juan-Manuel, et al.
Veröffentlicht: (2024)
von: Torres-Moreno, Juan-Manuel, et al.
Veröffentlicht: (2024)
Pre-training LLMs using human-like development data corpus
von: Bhardwaj, Khushi, et al.
Veröffentlicht: (2023)
von: Bhardwaj, Khushi, et al.
Veröffentlicht: (2023)
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service
von: Merx, Raphael, et al.
Veröffentlicht: (2024)
von: Merx, Raphael, et al.
Veröffentlicht: (2024)
ESG-FTSE: A corpus of news articles with ESG relevance labels and use cases
von: Pavlova, Mariya, et al.
Veröffentlicht: (2024)
von: Pavlova, Mariya, et al.
Veröffentlicht: (2024)
Leveraging LLMs for MT in Crisis Scenarios: a blueprint for low-resource languages
von: Lankford, Séamus, et al.
Veröffentlicht: (2024)
von: Lankford, Séamus, et al.
Veröffentlicht: (2024)
Assessing the potential of LLM-assisted annotation for corpus-based pragmatics and discourse analysis: The case of apology
von: Yu, Danni, et al.
Veröffentlicht: (2023)
von: Yu, Danni, et al.
Veröffentlicht: (2023)
FRACCO: A gold-standard annotated corpus of oncological entities with ICD-O-3.1 normalisation
von: Pignat, Johann, et al.
Veröffentlicht: (2025)
von: Pignat, Johann, et al.
Veröffentlicht: (2025)
Multilingual jailbreaking of LLMs using low-resource languages
von: Marx, Dylan, et al.
Veröffentlicht: (2026)
von: Marx, Dylan, et al.
Veröffentlicht: (2026)
Retrieval-augmented reasoning with lean language models
von: Chan, Ryan Sze-Yin, et al.
Veröffentlicht: (2025)
von: Chan, Ryan Sze-Yin, et al.
Veröffentlicht: (2025)
Can professional translators identify machine-generated text?
von: Farrell, Michael
Veröffentlicht: (2026)
von: Farrell, Michael
Veröffentlicht: (2026)
NeuroVoz: a Castillian Spanish corpus of parkinsonian speech
von: Mendes-Laureano, Janaína, et al.
Veröffentlicht: (2024)
von: Mendes-Laureano, Janaína, et al.
Veröffentlicht: (2024)
Can postgraduate translation students identify machine-generated text?
von: Farrell, Michael
Veröffentlicht: (2025)
von: Farrell, Michael
Veröffentlicht: (2025)
Toward domain-specific machine translation and quality estimation systems
von: Sharami, Javad Pourmostafa Roshan
Veröffentlicht: (2026)
von: Sharami, Javad Pourmostafa Roshan
Veröffentlicht: (2026)
NER- RoBERTa: Fine-Tuning RoBERTa for Named Entity Recognition (NER) within low-resource languages
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
AI-assisted cultural heritage dissemination: Comparing NMT and glossary-augmented LLM translation in rock art documents
von: Briva-Iglesias, Vicent, et al.
Veröffentlicht: (2026)
von: Briva-Iglesias, Vicent, et al.
Veröffentlicht: (2026)
Enhancing textual textbook question answering with large language models and retrieval augmented generation
von: Alawwad, Hessa Abdulrahman, et al.
Veröffentlicht: (2024)
von: Alawwad, Hessa Abdulrahman, et al.
Veröffentlicht: (2024)
Investigating the potential of Sparse Mixtures-of-Experts for multi-domain neural machine translation
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
Child vs. machine language learning: Can the logical structure of human language unleash LLMs?
von: Sauerland, Uli, et al.
Veröffentlicht: (2025)
von: Sauerland, Uli, et al.
Veröffentlicht: (2025)
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations
von: Shukla, Divyaksh, et al.
Veröffentlicht: (2025)
von: Shukla, Divyaksh, et al.
Veröffentlicht: (2025)
Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
von: Daouadi, Kheir Eddine, et al.
Veröffentlicht: (2024)
von: Daouadi, Kheir Eddine, et al.
Veröffentlicht: (2024)
Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers
von: Dhakal, Prakash, et al.
Veröffentlicht: (2024)
von: Dhakal, Prakash, et al.
Veröffentlicht: (2024)
An evaluation of LLMs and Google Translate for translation of selected Indian languages via sentiment and semantic analyses
von: Chandra, Rohitash, et al.
Veröffentlicht: (2025)
von: Chandra, Rohitash, et al.
Veröffentlicht: (2025)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model
von: Soong, David, et al.
Veröffentlicht: (2023)
von: Soong, David, et al.
Veröffentlicht: (2023)
Cross-lingual Text Classification Transfer: The Case of Ukrainian
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
To what extent is ChatGPT useful for language teacher lesson plan creation?
von: Dornburg, Alex, et al.
Veröffentlicht: (2024)
von: Dornburg, Alex, et al.
Veröffentlicht: (2024)
Predicate-Argument Structure Divergences in Chinese and English Parallel Sentences and their Impact on Language Transfer
von: Tripodi, Rocco, et al.
Veröffentlicht: (2025)
von: Tripodi, Rocco, et al.
Veröffentlicht: (2025)
Retrieval augmented generation based dynamic prompting for few-shot biomedical named entity recognition using large language models
von: Ge, Yao, et al.
Veröffentlicht: (2025)
von: Ge, Yao, et al.
Veröffentlicht: (2025)
Using Large Language Models for education managements in Vietnamese with low resources
von: Minh, Duc Do, et al.
Veröffentlicht: (2025)
von: Minh, Duc Do, et al.
Veröffentlicht: (2025)
Do LLM hallucination detectors suffer from low-resource effect?
von: Datta, Debtanu, et al.
Veröffentlicht: (2026)
von: Datta, Debtanu, et al.
Veröffentlicht: (2026)
A technical curriculum on language-oriented artificial intelligence in translation and specialised communication
von: Krüger, Ralph
Veröffentlicht: (2026)
von: Krüger, Ralph
Veröffentlicht: (2026)
The "LLM World of Words" English free association norms generated by large language models
von: Abramski, Katherine, et al.
Veröffentlicht: (2024)
von: Abramski, Katherine, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Phonetically rich corpus construction for a low-resourced language
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024) -
Tenyidie Syllabification corpus creation and deep learning applications
von: Angami, Teisovi, et al.
Veröffentlicht: (2025) -
How Culturally Aware are Vision-Language Models?
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024) -
Large language models for folktale type automation based on motifs: Cinderella case study
von: Arčon, Tjaša, et al.
Veröffentlicht: (2025) -
Automated evaluation of LLMs for effective machine translation of Mandarin Chinese to English
von: Zhang, Yue, et al.
Veröffentlicht: (2026)