Tracking Semantic Change in Slovene: A Novel Dataset and Optimal Transport-Based Distance
Fuente:
arXiv
Saved in:
| Main Authors: | Pranjić, Marko, Dobrovoljc, Kaja, Pollak, Senja, Martinc, Matej |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QFS-Composer: Query-focused summarization pipeline for less resourced languages
by: Đuranović, Vuk, et al.
Published: (2026)
by: Đuranović, Vuk, et al.
Published: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Efficient Aspect-Based Summarization of Climate Change Reports with Small Language Models
by: Ghinassi, Iacopo, et al.
Published: (2024)
by: Ghinassi, Iacopo, et al.
Published: (2024)
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
by: Lee, Lung-Hao, et al.
Published: (2026)
by: Lee, Lung-Hao, et al.
Published: (2026)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
by: Bothwell, Stephen, et al.
Published: (2024)
by: Bothwell, Stephen, et al.
Published: (2024)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
by: Bouchekif, Abdessalam, et al.
Published: (2026)
by: Bouchekif, Abdessalam, et al.
Published: (2026)
ML-Promise: A Multilingual Dataset for Corporate Promise Verification
by: Seki, Yohei, et al.
Published: (2024)
by: Seki, Yohei, et al.
Published: (2024)
EMO-KNOW: A Large Scale Dataset on Emotion and Emotion-cause
by: Nguyen, Mia Huong, et al.
Published: (2024)
by: Nguyen, Mia Huong, et al.
Published: (2024)
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
by: Bose, Joy
Published: (2026)
by: Bose, Joy
Published: (2026)
Locations of Characters in Narratives: Andersen and Persuasion Datasets
by: Ozyurt, Batuhan, et al.
Published: (2025)
by: Ozyurt, Batuhan, et al.
Published: (2025)
DimStance: Multilingual Datasets for Dimensional Stance Analysis
by: Becker, Jonas, et al.
Published: (2026)
by: Becker, Jonas, et al.
Published: (2026)
I run as fast as a rabbit, can you? A Multilingual Simile Dialogue Dataset
by: Ma, Longxuan, et al.
Published: (2023)
by: Ma, Longxuan, et al.
Published: (2023)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
by: Wu, Yuexin, et al.
Published: (2025)
by: Wu, Yuexin, et al.
Published: (2025)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
by: Sileo, Damien
Published: (2024)
by: Sileo, Damien
Published: (2024)
A Novel Word Pair-based Gaussian Sentence Similarity Algorithm For Bengali Extractive Text Summarization
by: Morshed, Fahim, et al.
Published: (2024)
by: Morshed, Fahim, et al.
Published: (2024)
DESS: DeBERTa Enhanced Syntactic-Semantic Aspect Sentiment Triplet Extraction
by: Thenuwara, Vishal, et al.
Published: (2025)
by: Thenuwara, Vishal, et al.
Published: (2025)
Algorithm for Semantic Network Generation from Texts of Low Resource Languages Such as Kiswahili
by: Wanjawa, Barack Wamkaya, et al.
Published: (2025)
by: Wanjawa, Barack Wamkaya, et al.
Published: (2025)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
UM_FHS at the CLEF 2025 SimpleText Track: Comparing No-Context and Fine-Tune Approaches for GPT-4.1 Models in Sentence and Document-Level Text Simplification
by: Kocbek, Primoz, et al.
Published: (2025)
by: Kocbek, Primoz, et al.
Published: (2025)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
A Benchmark of French ASR Systems Based on Error Severity
by: Tholly, Antoine, et al.
Published: (2025)
by: Tholly, Antoine, et al.
Published: (2025)
A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
by: Peláez-González, Carlos, et al.
Published: (2025)
by: Peláez-González, Carlos, et al.
Published: (2025)
Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression
by: Baxi, Rahul
Published: (2025)
by: Baxi, Rahul
Published: (2025)
CR-LT-KGQA: A Knowledge Graph Question Answering Dataset Requiring Commonsense Reasoning and Long-Tail Knowledge
by: Guo, Willis, et al.
Published: (2024)
by: Guo, Willis, et al.
Published: (2024)
A RoBERTa-Based Functional Syntax Annotation Model for Chinese Texts
by: Xiaohui, Han, et al.
Published: (2025)
by: Xiaohui, Han, et al.
Published: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
by: Kugler, Kai
Published: (2025)
by: Kugler, Kai
Published: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
Blocks Architecture (BloArk): Efficient, Cost-Effective, and Incremental Dataset Architecture for Wikipedia Revision History
by: Li, Lingxi, et al.
Published: (2024)
by: Li, Lingxi, et al.
Published: (2024)
Sarcasm Detection in a Less-Resourced Language
by: Đoković, Lazar, et al.
Published: (2024)
by: Đoković, Lazar, et al.
Published: (2024)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
by: Cui, Hyang
Published: (2025)
by: Cui, Hyang
Published: (2025)
Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus
by: Boratyn, Daria, et al.
Published: (2026)
by: Boratyn, Daria, et al.
Published: (2026)
SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
by: Chaybouti, Sofian, et al.
Published: (2021)
by: Chaybouti, Sofian, et al.
Published: (2021)
SemEval-2026 Task 3: Dimensional Aspect-Based Sentiment Analysis (DimABSA)
by: Yu, Liang-Chih, et al.
Published: (2026)
by: Yu, Liang-Chih, et al.
Published: (2026)
Similar Items
-
QFS-Composer: Query-focused summarization pipeline for less resourced languages
by: Đuranović, Vuk, et al.
Published: (2026) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025) -
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025) -
Efficient Aspect-Based Summarization of Climate Change Reports with Small Language Models
by: Ghinassi, Iacopo, et al.
Published: (2024)