SiTSE: Sinhala Text Simplification Dataset and Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Ranathunga, Surangika, Sirithunga, Rumesh, Rathnayake, Himashi, De Silva, Lahiru, Aluthwala, Thamindu, Peramuna, Saman, Shekhar, Ravi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sinhala Physical Common Sense Reasoning Dataset for Global PIQA
by: de Silva, Nisansa, et al.
Published: (2026)
by: de Silva, Nisansa, et al.
Published: (2026)
Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches
by: De Mel, Yomal, et al.
Published: (2024)
by: De Mel, Yomal, et al.
Published: (2024)
SinLlama -- A Large Language Model for Sinhala
by: Aravinda, H. W. K., et al.
Published: (2025)
by: Aravinda, H. W. K., et al.
Published: (2025)
OasisSimp: An Open-source Asian-English Sentence Simplification Dataset
by: Liu, Hannah, et al.
Published: (2026)
by: Liu, Hannah, et al.
Published: (2026)
A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala
by: Ranathunga, Surangika, et al.
Published: (2024)
by: Ranathunga, Surangika, et al.
Published: (2024)
Improving the quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristics
by: Fernando, Aloka, et al.
Published: (2025)
by: Fernando, Aloka, et al.
Published: (2025)
Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora
by: Ranathunga, Surangika, et al.
Published: (2024)
by: Ranathunga, Surangika, et al.
Published: (2024)
Linguistic Entity Masking to Improve Cross-Lingual Representation of Multilingual Language Models for Low-Resource Languages
by: Fernando, Aloka, et al.
Published: (2025)
by: Fernando, Aloka, et al.
Published: (2025)
Unsupervised Bilingual Lexicon Induction for Low Resource Languages
by: Rathnayake, Charitha, et al.
Published: (2024)
by: Rathnayake, Charitha, et al.
Published: (2024)
MuTSE: A Human-in-the-Loop Multi-use Text Simplification Evaluator
by: Roscan, Rares-Alexandru, et al.
Published: (2026)
by: Roscan, Rares-Alexandru, et al.
Published: (2026)
Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
by: Puranegedara, Imalsha, et al.
Published: (2025)
by: Puranegedara, Imalsha, et al.
Published: (2025)
Linguistic Analysis of Sinhala YouTube Comments on Sinhala Music Videos: A Dataset Study
by: De Mel, W. M. Yomal, et al.
Published: (2025)
by: De Mel, W. M. Yomal, et al.
Published: (2025)
Shoulders of Giants: A Look at the Degree and Utility of Openness in NLP Research
by: Ranathunga, Surangika, et al.
Published: (2024)
by: Ranathunga, Surangika, et al.
Published: (2024)
SiDiaC: Sinhala Diachronic Corpus
by: Jayatilleke, Nevidu, et al.
Published: (2025)
by: Jayatilleke, Nevidu, et al.
Published: (2025)
Enhancing Multilingual Sentiment Analysis with Explainability for Sinhala, English, and Code-Mixed Content
by: Rizvi, Azmarah, et al.
Published: (2025)
by: Rizvi, Azmarah, et al.
Published: (2025)
LMSpell: Neural Spell Checking for Low-Resource Languages
by: Gunathilake, Akesh, et al.
Published: (2025)
by: Gunathilake, Akesh, et al.
Published: (2025)
Large Language Models for Ingredient Substitution in Food Recipes using Supervised Fine-tuning and Direct Preference Optimization
by: Senath, Thevin, et al.
Published: (2024)
by: Senath, Thevin, et al.
Published: (2024)
SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
by: Pramodya, Ashmari, et al.
Published: (2025)
by: Pramodya, Ashmari, et al.
Published: (2025)
SiPaKosa: A Comprehensive Corpus of Canonical and Classical Buddhist Texts in Sinhala and Pali
by: Gurusinghe, Ranidu, et al.
Published: (2026)
by: Gurusinghe, Ranidu, et al.
Published: (2026)
Beyond Vanilla Fine-Tuning: Leveraging Multistage, Multilingual, and Domain-Specific Methods for Low-Resource Machine Translation
by: Thillainathan, Sarubi, et al.
Published: (2025)
by: Thillainathan, Sarubi, et al.
Published: (2025)
SOLD: Sinhala Offensive Language Dataset
by: Ranasinghe, Tharindu, et al.
Published: (2022)
by: Ranasinghe, Tharindu, et al.
Published: (2022)
Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
by: Hameed, Sameeah Noreen, et al.
Published: (2025)
by: Hameed, Sameeah Noreen, et al.
Published: (2025)
MCTS: A Multi-Reference Chinese Text Simplification Dataset
by: Chong, Ruining, et al.
Published: (2023)
by: Chong, Ruining, et al.
Published: (2023)
Sinhala-English Word Embedding Alignment: Introducing Datasets and Benchmark for a Low Resource Language
by: Wickramasinghe, Kasun, et al.
Published: (2023)
by: Wickramasinghe, Kasun, et al.
Published: (2023)
SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0
by: Jayatilleke, Nevidu, et al.
Published: (2026)
by: Jayatilleke, Nevidu, et al.
Published: (2026)
Unlocking Parameter-Efficient Fine-Tuning for Low-Resource Language Translation
by: Su, Tong, et al.
Published: (2024)
by: Su, Tong, et al.
Published: (2024)
SinFoS: A Parallel Dataset for Translating Sinhala Figures of Speech
by: Sofalas, Johan, et al.
Published: (2026)
by: Sofalas, Johan, et al.
Published: (2026)
GeeSanBhava: Sentiment Tagged Sinhala Music Video Comment Data Set
by: De Mel, Yomal, et al.
Published: (2025)
by: De Mel, Yomal, et al.
Published: (2025)
Swa-bhasha Resource Hub: Romanized Sinhala to Sinhala Transliteration Systems and Data Resources
by: Sumanathilaka, Deshan, et al.
Published: (2025)
by: Sumanathilaka, Deshan, et al.
Published: (2025)
Evaluating LLMs for Targeted Concept Simplification for Domain-Specific Texts
by: Asthana, Sumit, et al.
Published: (2024)
by: Asthana, Sumit, et al.
Published: (2024)
Improving Estonian Text Simplification through Pretrained Language Models and Custom Datasets
by: Barbu, Eduard, et al.
Published: (2025)
by: Barbu, Eduard, et al.
Published: (2025)
Text Simplification with Sentence Embeddings
by: Shardlow, Matthew
Published: (2025)
by: Shardlow, Matthew
Published: (2025)
A Framework to Assess Multilingual Vulnerabilities of LLMs
by: Tang, Likai, et al.
Published: (2025)
by: Tang, Likai, et al.
Published: (2025)
Survey on Publicly Available Sinhala Natural Language Processing Tools and Research
by: de Silva, Nisansa
Published: (2019)
by: de Silva, Nisansa
Published: (2019)
The SAMER Arabic Text Simplification Corpus
by: Alhafni, Bashar, et al.
Published: (2024)
by: Alhafni, Bashar, et al.
Published: (2024)
SinhaLegal: A Benchmark Corpus for Information Extraction and Analysis in Sinhala Legislative Texts
by: Lasandi, Minduli, et al.
Published: (2026)
by: Lasandi, Minduli, et al.
Published: (2026)
Benchmarking Automated Clinical Language Simplification: Dataset, Algorithm, and Evaluation
by: Luo, Junyu, et al.
Published: (2020)
by: Luo, Junyu, et al.
Published: (2020)
Evaluation Under Imperfect Benchmarks and Ratings: A Case Study in Text Simplification
by: Liu, Joseph, et al.
Published: (2025)
by: Liu, Joseph, et al.
Published: (2025)
Evaluating Small Decoder-Only Language Models for Grammar Correction and Text Simplification
by: Lamelas, Anthony
Published: (2026)
by: Lamelas, Anthony
Published: (2026)
Role of Dependency Distance in Text Simplification: A Human vs ChatGPT Simplification Comparison
by: Lee, Sumi, et al.
Published: (2024)
by: Lee, Sumi, et al.
Published: (2024)
Similar Items
-
Sinhala Physical Common Sense Reasoning Dataset for Global PIQA
by: de Silva, Nisansa, et al.
Published: (2026) -
Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches
by: De Mel, Yomal, et al.
Published: (2024) -
SinLlama -- A Large Language Model for Sinhala
by: Aravinda, H. W. K., et al.
Published: (2025) -
OasisSimp: An Open-source Asian-English Sentence Simplification Dataset
by: Liu, Hannah, et al.
Published: (2026) -
A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala
by: Ranathunga, Surangika, et al.
Published: (2024)