A Subword Embedding Approach for Variation Detection in Luxembourgish User Comments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lutgen, Anne-Marie, Plum, Alistair, Purschke, Christoph |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neural Text Normalization for Luxembourgish using Real-Life Variation Data
von: Lutgen, Anne-Marie, et al.
Veröffentlicht: (2024)
von: Lutgen, Anne-Marie, et al.
Veröffentlicht: (2024)
LuxBank: The First Universal Dependency Treebank for Luxembourgish
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
Language Ideologies in a Multilingual Society: An LLM-based Analysis of Luxembourgish News Comments
von: Milano, Emilia, et al.
Veröffentlicht: (2026)
von: Milano, Emilia, et al.
Veröffentlicht: (2026)
Variation is the Norm: Embracing Sociolinguistics in NLP
von: Lutgen, Anne-Marie, et al.
Veröffentlicht: (2026)
von: Lutgen, Anne-Marie, et al.
Veröffentlicht: (2026)
Text Generation Models for Luxembourgish with Limited Data: A Balanced Multilingual Strategy
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
Identity-Aware Large Language Models require Cultural Reasoning
von: Plum, Alistair, et al.
Veröffentlicht: (2025)
von: Plum, Alistair, et al.
Veröffentlicht: (2025)
ltzGLUE: Luxembourgish General Language Understanding Evaluation
von: Plum, Alistair, et al.
Veröffentlicht: (2026)
von: Plum, Alistair, et al.
Veröffentlicht: (2026)
Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New Language
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
von: Plum, Alistair, et al.
Veröffentlicht: (2024)
Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction
von: Schwager, Nils, et al.
Veröffentlicht: (2026)
von: Schwager, Nils, et al.
Veröffentlicht: (2026)
Adapting Multilingual Embedding Models to Historical Luxembourgish
von: Michail, Andrianos, et al.
Veröffentlicht: (2025)
von: Michail, Andrianos, et al.
Veröffentlicht: (2025)
LuxEmbedder: A Cross-Lingual Approach to Enhanced Luxembourgish Sentence Embeddings
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
Subword Tokenization Strategies for Kurdish Word Embeddings
von: Salehi, Ali, et al.
Veröffentlicht: (2025)
von: Salehi, Ali, et al.
Veröffentlicht: (2025)
Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
von: Plum, Alistair, et al.
Veröffentlicht: (2026)
von: Plum, Alistair, et al.
Veröffentlicht: (2026)
LuxBorrow: From Pompier to Pompjee, Tracing Borrowing in Luxembourgish
von: Hosseini-Kivanani, Nina, et al.
Veröffentlicht: (2026)
von: Hosseini-Kivanani, Nina, et al.
Veröffentlicht: (2026)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
Distributional Properties of Subword Regularization
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
Lexically Grounded Subword Segmentation
von: Libovický, Jindřich, et al.
Veröffentlicht: (2024)
von: Libovický, Jindřich, et al.
Veröffentlicht: (2024)
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
von: Liu, Yihong, et al.
Veröffentlicht: (2023)
von: Liu, Yihong, et al.
Veröffentlicht: (2023)
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
von: Teklehaymanot, Hailay, et al.
Veröffentlicht: (2026)
von: Teklehaymanot, Hailay, et al.
Veröffentlicht: (2026)
LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data
von: Valline, Julian, et al.
Veröffentlicht: (2025)
von: Valline, Julian, et al.
Veröffentlicht: (2025)
What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects
von: Blaschke, Verena, et al.
Veröffentlicht: (2024)
von: Blaschke, Verena, et al.
Veröffentlicht: (2024)
LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
Exploiting User Comments for Early Detection of Fake News Prior to Users' Commenting
von: Nan, Qiong, et al.
Veröffentlicht: (2023)
von: Nan, Qiong, et al.
Veröffentlicht: (2023)
ByteSpan: Information-Driven Subword Tokenisation
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
Subword models struggle with word learning, but surprisal hides it
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2025)
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2025)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
von: Meyer, Francois, et al.
Veröffentlicht: (2025)
von: Meyer, Francois, et al.
Veröffentlicht: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
von: Parra, Iñigo
Veröffentlicht: (2024)
von: Parra, Iñigo
Veröffentlicht: (2024)
Testing Low-Resource Language Support in LLMs Using Language Proficiency Exams: the Case of Luxembourgish
von: Lothritz, Cedric, et al.
Veröffentlicht: (2025)
von: Lothritz, Cedric, et al.
Veröffentlicht: (2025)
A Systematic Analysis of Subwords and Cross-Lingual Transfer in Multilingual Translation
von: Meyer, Francois, et al.
Veröffentlicht: (2024)
von: Meyer, Francois, et al.
Veröffentlicht: (2024)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
Learning Mutually Informed Representations for Characters and Subwords
von: Wang, Yilin, et al.
Veröffentlicht: (2023)
von: Wang, Yilin, et al.
Veröffentlicht: (2023)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
von: Zouhar, Vilém
Veröffentlicht: (2024)
von: Zouhar, Vilém
Veröffentlicht: (2024)
Tokenization Falling Short: On Subword Robustness in Large Language Models
von: Chai, Yekun, et al.
Veröffentlicht: (2024)
von: Chai, Yekun, et al.
Veröffentlicht: (2024)
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
von: Sims, Anya, et al.
Veröffentlicht: (2025)
von: Sims, Anya, et al.
Veröffentlicht: (2025)
Do LLMs Know What Luxembourgish Borrows? Probing Lexical Neology in Low-Resource Multilingual Models
von: Hosseini-Kivanani, Nina
Veröffentlicht: (2026)
von: Hosseini-Kivanani, Nina
Veröffentlicht: (2026)
Do Large Language Models Grasp The Grammar? Evidence from Grammar-Book-Guided Probing in Luxembourgish
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
von: Patwary, Firoj Ahmmed, et al.
Veröffentlicht: (2025)
von: Patwary, Firoj Ahmmed, et al.
Veröffentlicht: (2025)
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
von: Tsuji, Kohei, et al.
Veröffentlicht: (2024)
von: Tsuji, Kohei, et al.
Veröffentlicht: (2024)
Assessing the Importance of Frequency versus Compositionality for Subword-based Tokenization in NMT
von: Wolleb, Benoist, et al.
Veröffentlicht: (2023)
von: Wolleb, Benoist, et al.
Veröffentlicht: (2023)
Token Alignment via Character Matching for Subword Completion
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Neural Text Normalization for Luxembourgish using Real-Life Variation Data
von: Lutgen, Anne-Marie, et al.
Veröffentlicht: (2024) -
LuxBank: The First Universal Dependency Treebank for Luxembourgish
von: Plum, Alistair, et al.
Veröffentlicht: (2024) -
Language Ideologies in a Multilingual Society: An LLM-based Analysis of Luxembourgish News Comments
von: Milano, Emilia, et al.
Veröffentlicht: (2026) -
Variation is the Norm: Embracing Sociolinguistics in NLP
von: Lutgen, Anne-Marie, et al.
Veröffentlicht: (2026) -
Text Generation Models for Luxembourgish with Limited Data: A Balanced Multilingual Strategy
von: Plum, Alistair, et al.
Veröffentlicht: (2024)