SinFoS: A Parallel Dataset for Translating Sinhala Figures of Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sofalas, Johan, Pavithra, Dilushri, Jayatilleke, Nevidu, Weerasinghe, Ruvan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Hybrid Architecture with Efficient Fine Tuning for Abstractive Patent Document Summarization
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025)
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025)
Advancements in Natural Language Processing for Automatic Text Summarization
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025)
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025)
SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2026)
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2026)
SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts
von: Lasandi, Minduli, et al.
Veröffentlicht: (2026)
von: Lasandi, Minduli, et al.
Veröffentlicht: (2026)
SiDiaC: Sinhala Diachronic Corpus
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025)
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025)
SiPaKosa: A Comprehensive Corpus of Canonical and Classical Buddhist Texts in Sinhala and Pali
von: Gurusinghe, Ranidu, et al.
Veröffentlicht: (2026)
von: Gurusinghe, Ranidu, et al.
Veröffentlicht: (2026)
Zero-shot OCR Accuracy of Low-Resourced Languages: A Comparative Analysis on Sinhala and Tamil
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025)
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025)
Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala
von: Rajapakse, Minuri, et al.
Veröffentlicht: (2026)
von: Rajapakse, Minuri, et al.
Veröffentlicht: (2026)
SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
von: Pramodya, Ashmari, et al.
Veröffentlicht: (2025)
von: Pramodya, Ashmari, et al.
Veröffentlicht: (2025)
Swa-bhasha Resource Hub: Romanized Sinhala to Sinhala Transliteration Systems and Data Resources
von: Sumanathilaka, Deshan, et al.
Veröffentlicht: (2025)
von: Sumanathilaka, Deshan, et al.
Veröffentlicht: (2025)
SinLlama -- A Large Language Model for Sinhala
von: Aravinda, H. W. K., et al.
Veröffentlicht: (2025)
von: Aravinda, H. W. K., et al.
Veröffentlicht: (2025)
IndoNLP 2025: Shared Task on Real-Time Reverse Transliteration for Romanized Indo-Aryan languages
von: Sumanathilaka, Deshan, et al.
Veröffentlicht: (2025)
von: Sumanathilaka, Deshan, et al.
Veröffentlicht: (2025)
Linguistic Analysis of Sinhala YouTube Comments on Sinhala Music Videos: A Dataset Study
von: De Mel, W. M. Yomal, et al.
Veröffentlicht: (2025)
von: De Mel, W. M. Yomal, et al.
Veröffentlicht: (2025)
SiTSE: Sinhala Text Simplification Dataset and Evaluation
von: Ranathunga, Surangika, et al.
Veröffentlicht: (2024)
von: Ranathunga, Surangika, et al.
Veröffentlicht: (2024)
A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala
von: Ranathunga, Surangika, et al.
Veröffentlicht: (2024)
von: Ranathunga, Surangika, et al.
Veröffentlicht: (2024)
Sinhala Physical Common Sense Reasoning Dataset for Global PIQA
von: de Silva, Nisansa, et al.
Veröffentlicht: (2026)
von: de Silva, Nisansa, et al.
Veröffentlicht: (2026)
SOLD: Sinhala Offensive Language Dataset
von: Ranasinghe, Tharindu, et al.
Veröffentlicht: (2022)
von: Ranasinghe, Tharindu, et al.
Veröffentlicht: (2022)
A Low-Resource Speech-Driven NLP Pipeline for Sinhala Dyslexia Assistance
von: Perera, Peshala, et al.
Veröffentlicht: (2025)
von: Perera, Peshala, et al.
Veröffentlicht: (2025)
Textless Speech-to-Speech Translation With Limited Parallel Data
von: Diwan, Anuj, et al.
Veröffentlicht: (2023)
von: Diwan, Anuj, et al.
Veröffentlicht: (2023)
FoQA: A Faroese Question-Answering Dataset
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
von: Karthika, N J, et al.
Veröffentlicht: (2026)
von: Karthika, N J, et al.
Veröffentlicht: (2026)
Sinhala-English Word Embedding Alignment: Introducing Datasets and Benchmark for a Low Resource Language
von: Wickramasinghe, Kasun, et al.
Veröffentlicht: (2023)
von: Wickramasinghe, Kasun, et al.
Veröffentlicht: (2023)
ACADATA: Parallel Dataset of Academic Data for Machine Translation
von: Lacunza, Iñaki, et al.
Veröffentlicht: (2025)
von: Lacunza, Iñaki, et al.
Veröffentlicht: (2025)
SPACER: A Parallel Dataset of Speech Production And Comprehension of Error Repairs
von: Upadhye, Shiva, et al.
Veröffentlicht: (2025)
von: Upadhye, Shiva, et al.
Veröffentlicht: (2025)
NSINA: A News Corpus for Sinhala
von: Hettiarachchi, Hansi, et al.
Veröffentlicht: (2024)
von: Hettiarachchi, Hansi, et al.
Veröffentlicht: (2024)
RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speech
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
MELD-ST: An Emotion-aware Speech Translation Dataset
von: Chen, Sirou, et al.
Veröffentlicht: (2024)
von: Chen, Sirou, et al.
Veröffentlicht: (2024)
Swa Bhasha: Message-Based Singlish to Sinhala Transliteration
von: Athukorala, Maneesha U., et al.
Veröffentlicht: (2024)
von: Athukorala, Maneesha U., et al.
Veröffentlicht: (2024)
CBR-RAG: Case-Based Reasoning for Retrieval Augmented Generation in LLMs for Legal Question Answering
von: Wiratunga, Nirmalie, et al.
Veröffentlicht: (2024)
von: Wiratunga, Nirmalie, et al.
Veröffentlicht: (2024)
XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics
von: Liu, Jingxuan, et al.
Veröffentlicht: (2026)
von: Liu, Jingxuan, et al.
Veröffentlicht: (2026)
Stylomech: Unveiling Authorship via Computational Stylometry in English and Romanized Sinhala
von: Faumi, Nabeelah, et al.
Veröffentlicht: (2025)
von: Faumi, Nabeelah, et al.
Veröffentlicht: (2025)
Survey on Publicly Available Sinhala Natural Language Processing Tools and Research
von: de Silva, Nisansa
Veröffentlicht: (2019)
von: de Silva, Nisansa
Veröffentlicht: (2019)
Identifying False Content and Hate Speech in Sinhala YouTube Videos by Analyzing the Audio
von: Wickramaarachchi, W. A. K. M., et al.
Veröffentlicht: (2024)
von: Wickramaarachchi, W. A. K. M., et al.
Veröffentlicht: (2024)
Subasa - Adapting Language Models for Low-resourced Offensive Language Detection in Sinhala
von: Haturusinghe, Shanilka, et al.
Veröffentlicht: (2025)
von: Haturusinghe, Shanilka, et al.
Veröffentlicht: (2025)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
von: Rashidi, Sina, et al.
Veröffentlicht: (2025)
von: Rashidi, Sina, et al.
Veröffentlicht: (2025)
EmoScan: Automatic Screening of Depression Symptoms in Romanized Sinhala Tweets
von: Hewapathirana, Jayathi, et al.
Veröffentlicht: (2024)
von: Hewapathirana, Jayathi, et al.
Veröffentlicht: (2024)
GeeSanBhava: Sentiment Tagged Sinhala Music Video Comment Data Set
von: De Mel, Yomal, et al.
Veröffentlicht: (2025)
von: De Mel, Yomal, et al.
Veröffentlicht: (2025)
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
von: Min, Anna, et al.
Veröffentlicht: (2025)
von: Min, Anna, et al.
Veröffentlicht: (2025)
Granary: Speech Recognition and Translation Dataset in 25 European Languages
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2025)
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2025)
Semi-Synthetic Parallel Data for Translation Quality Estimation: A Case Study of Dataset Building for an Under-Resourced Language Pair
von: Siani, Assaf, et al.
Veröffentlicht: (2026)
von: Siani, Assaf, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Hybrid Architecture with Efficient Fine Tuning for Abstractive Patent Document Summarization
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025) -
Advancements in Natural Language Processing for Automatic Text Summarization
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025) -
SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2026) -
SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts
von: Lasandi, Minduli, et al.
Veröffentlicht: (2026) -
SiDiaC: Sinhala Diachronic Corpus
von: Jayatilleke, Nevidu, et al.
Veröffentlicht: (2025)