NSINA: A News Corpus for Sinhala
Fuente:
arXiv
Saved in:
| Main Authors: | Hettiarachchi, Hansi, Premasiri, Damith, Uyangodage, Lasitha, Ranasinghe, Tharindu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SOLD: Sinhala Offensive Language Dataset
by: Ranasinghe, Tharindu, et al.
Published: (2022)
by: Ranasinghe, Tharindu, et al.
Published: (2022)
Overview of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)
by: Hettiarachchi, Hansi, et al.
Published: (2024)
by: Hettiarachchi, Hansi, et al.
Published: (2024)
A Federated Learning Approach to Privacy Preserving Offensive Language Identification
by: Zampieri, Marcos, et al.
Published: (2024)
by: Zampieri, Marcos, et al.
Published: (2024)
LLM-based Embedders for Prior Case Retrieval
by: Premasiri, Damith, et al.
Published: (2025)
by: Premasiri, Damith, et al.
Published: (2025)
AHaSIS: Shared Task on Sentiment Analysis for Arabic Dialects
by: Alharbi, Maram, et al.
Published: (2025)
by: Alharbi, Maram, et al.
Published: (2025)
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification
by: North, Kai, et al.
Published: (2022)
by: North, Kai, et al.
Published: (2022)
Keyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content
by: Rizvi, F. A., et al.
Published: (2025)
by: Rizvi, F. A., et al.
Published: (2025)
Enhancing Multilingual Sentiment Analysis with Explainability for Sinhala, English, and Code-Mixed Content
by: Rizvi, Azmarah, et al.
Published: (2025)
by: Rizvi, Azmarah, et al.
Published: (2025)
Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New Language
by: Plum, Alistair, et al.
Published: (2024)
by: Plum, Alistair, et al.
Published: (2024)
AlbNews: A Corpus of Headlines for Topic Modeling in Albanian
by: Çano, Erion, et al.
Published: (2024)
by: Çano, Erion, et al.
Published: (2024)
Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
by: Plum, Alistair, et al.
Published: (2026)
by: Plum, Alistair, et al.
Published: (2026)
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
by: Kesgin, H. Toprak, et al.
Published: (2024)
by: Kesgin, H. Toprak, et al.
Published: (2024)
Graph-Eq: Discovering Mathematical Equations using Graph Generative Models
by: Ranasinghe, Nisal, et al.
Published: (2025)
by: Ranasinghe, Nisal, et al.
Published: (2025)
MUNIChus: Multilingual News Image Captioning Benchmark
by: Chen, Yuji, et al.
Published: (2026)
by: Chen, Yuji, et al.
Published: (2026)
HLDC: Hindi Legal Documents Corpus
by: Kapoor, Arnav, et al.
Published: (2022)
by: Kapoor, Arnav, et al.
Published: (2022)
MultiLS: A Multi-task Lexical Simplification Framework
by: North, Kai, et al.
Published: (2024)
by: North, Kai, et al.
Published: (2024)
IDIAPers @ Causal News Corpus 2022: Efficient Causal Relation Identification Through a Prompt-based Few-shot Approach
by: Burdisso, Sergio, et al.
Published: (2022)
by: Burdisso, Sergio, et al.
Published: (2022)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026)
by: Asonitis, Antonis, et al.
Published: (2026)
LPI-RIT at LeWiDi-2025: Improving Distributional Predictions via Metadata and Loss Reweighting with DisCo
by: Sawkar, Mandira, et al.
Published: (2025)
by: Sawkar, Mandira, et al.
Published: (2025)
Identifying False Content and Hate Speech in Sinhala YouTube Videos by Analyzing the Audio
by: Wickramaarachchi, W. A. K. M., et al.
Published: (2024)
by: Wickramaarachchi, W. A. K. M., et al.
Published: (2024)
Cross-lingual Named Entity Corpus for Slavic Languages
by: Piskorski, Jakub, et al.
Published: (2024)
by: Piskorski, Jakub, et al.
Published: (2024)
Towards Generalized Offensive Language Identification
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
GINN-LP: A Growing Interpretable Neural Network for Discovering Multivariate Laurent Polynomial Equations
by: Ranasinghe, Nisal, et al.
Published: (2023)
by: Ranasinghe, Nisal, et al.
Published: (2023)
Meta4XNLI: A Crosslingual Parallel Corpus for Metaphor Detection and Interpretation
by: Sanchez-Bayona, Elisa, et al.
Published: (2024)
by: Sanchez-Bayona, Elisa, et al.
Published: (2024)
NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
by: Ghimire, Rupak Raj, et al.
Published: (2026)
by: Ghimire, Rupak Raj, et al.
Published: (2026)
Retrieve, Then Classify: Corpus-Grounded Automation of Clinical Value Set Authoring
by: Mukherjee, Sumit, et al.
Published: (2026)
by: Mukherjee, Sumit, et al.
Published: (2026)
PMOA-TTS: Introducing the PubMed Open Access Textual Times Series Corpus
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
ProRefine: Inference-Time Prompt Refinement with Textual Feedback
by: Pandita, Deepak, et al.
Published: (2025)
by: Pandita, Deepak, et al.
Published: (2025)
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
by: Ziegler, Ingo, et al.
Published: (2024)
by: Ziegler, Ingo, et al.
Published: (2024)
Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning
by: Rahman, Md Mushfiqur, et al.
Published: (2024)
by: Rahman, Md Mushfiqur, et al.
Published: (2024)
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
MathBridge: A Large Corpus Dataset for Translating Spoken Mathematical Expressions into $LaTeX$ Formulas for Improved Readability
by: Jung, Kyudan, et al.
Published: (2024)
by: Jung, Kyudan, et al.
Published: (2024)
A Novel Cartography-Based Curriculum Learning Method Applied on RoNLI: The First Romanian Natural Language Inference Corpus
by: Poesina, Eduard, et al.
Published: (2024)
by: Poesina, Eduard, et al.
Published: (2024)
Filtered Corpus Training (FiCT) Shows that Language Models can Generalize from Indirect Evidence
by: Patil, Abhinav, et al.
Published: (2024)
by: Patil, Abhinav, et al.
Published: (2024)
Reconstructing Sepsis Trajectories from Clinical Case Reports using LLMs: the Textual Time Series Corpus for Sepsis
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
UQA: Corpus for Urdu Question Answering
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
Exploring the Performance of Large Language Models on Subjective Span Identification Tasks
by: Dmonte, Alphaeus, et al.
Published: (2026)
by: Dmonte, Alphaeus, et al.
Published: (2026)
DORE: A Dataset For Portuguese Definition Generation
by: Furtado, Anna Beatriz Dimas, et al.
Published: (2024)
by: Furtado, Anna Beatriz Dimas, et al.
Published: (2024)
IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
by: Paul, Shounak, et al.
Published: (2025)
by: Paul, Shounak, et al.
Published: (2025)
Similar Items
-
SOLD: Sinhala Offensive Language Dataset
by: Ranasinghe, Tharindu, et al.
Published: (2022) -
Overview of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)
by: Hettiarachchi, Hansi, et al.
Published: (2024) -
A Federated Learning Approach to Privacy Preserving Offensive Language Identification
by: Zampieri, Marcos, et al.
Published: (2024) -
LLM-based Embedders for Prior Case Retrieval
by: Premasiri, Damith, et al.
Published: (2025) -
AHaSIS: Shared Task on Sentiment Analysis for Arabic Dialects
by: Alharbi, Maram, et al.
Published: (2025)