The Effect of Data Partitioning Strategy on Model Generalizability: A Case Study of Morphological Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zoey, Dorr, Bonnie J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
von: Kocbek, Primož, et al.
Veröffentlicht: (2025)
von: Kocbek, Primož, et al.
Veröffentlicht: (2025)
Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
von: Klerings, Alina, et al.
Veröffentlicht: (2025)
von: Klerings, Alina, et al.
Veröffentlicht: (2025)
Dialect Normalization using Large Language Models and Morphological Rules
von: Dimakis, Antonios, et al.
Veröffentlicht: (2025)
von: Dimakis, Antonios, et al.
Veröffentlicht: (2025)
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
von: Dang, Thao Anh, et al.
Veröffentlicht: (2024)
von: Dang, Thao Anh, et al.
Veröffentlicht: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Comonadic Morphophonology: A Compositional Framework for Context-Dependent Morphological Rules in Finnish
von: Jang, Yongseok
Veröffentlicht: (2026)
von: Jang, Yongseok
Veröffentlicht: (2026)
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
von: Akavarapu, V. S. D. S. Mahesh, et al.
Veröffentlicht: (2025)
von: Akavarapu, V. S. D. S. Mahesh, et al.
Veröffentlicht: (2025)
Strategy Adaptation in Large Language Model Werewolf Agents
von: Nakamori, Fuya, et al.
Veröffentlicht: (2025)
von: Nakamori, Fuya, et al.
Veröffentlicht: (2025)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2025)
Morphological Analysis for the Maltese Language: The Challenges of a Hybrid System
von: Borg, Claudia, et al.
Veröffentlicht: (2017)
von: Borg, Claudia, et al.
Veröffentlicht: (2017)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders
von: Sauter, Adrian, et al.
Veröffentlicht: (2025)
von: Sauter, Adrian, et al.
Veröffentlicht: (2025)
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
von: Chen, Yanbing, et al.
Veröffentlicht: (2024)
von: Chen, Yanbing, et al.
Veröffentlicht: (2024)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
Socially Responsible Data for Large Multilingual Language Models
von: Smart, Andrew, et al.
Veröffentlicht: (2024)
von: Smart, Andrew, et al.
Veröffentlicht: (2024)
Recent Trends in Linear Text Segmentation: a Survey
von: Ghinassi, Iacopo, et al.
Veröffentlicht: (2024)
von: Ghinassi, Iacopo, et al.
Veröffentlicht: (2024)
Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task
von: Evelo, Bart, et al.
Veröffentlicht: (2026)
von: Evelo, Bart, et al.
Veröffentlicht: (2026)
Decoding-Free Sampling Strategies for LLM Marginalization
von: Pohl, David, et al.
Veröffentlicht: (2025)
von: Pohl, David, et al.
Veröffentlicht: (2025)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
von: Klöser, Lars, et al.
Veröffentlicht: (2024)
von: Klöser, Lars, et al.
Veröffentlicht: (2024)
Lisbon Computational Linguists at SemEval-2024 Task 2: Using A Mistral 7B Model and Data Augmentation
von: Guimarães, Artur, et al.
Veröffentlicht: (2024)
von: Guimarães, Artur, et al.
Veröffentlicht: (2024)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification
von: Kang, Migyeong, et al.
Veröffentlicht: (2026)
von: Kang, Migyeong, et al.
Veröffentlicht: (2026)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
von: Liu, Han, et al.
Veröffentlicht: (2026)
von: Liu, Han, et al.
Veröffentlicht: (2026)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
von: Nzeyimana, Antoine, et al.
Veröffentlicht: (2025)
von: Nzeyimana, Antoine, et al.
Veröffentlicht: (2025)
Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models
von: Wen, Zhihao, et al.
Veröffentlicht: (2025)
von: Wen, Zhihao, et al.
Veröffentlicht: (2025)
RUQuant: Towards Refining Uniform Quantization for Large Language Models
von: Liu, Han, et al.
Veröffentlicht: (2026)
von: Liu, Han, et al.
Veröffentlicht: (2026)
Curating Grounded Synthetic Data with Global Perspectives for Equitable AI
von: Törnquist, Elin, et al.
Veröffentlicht: (2024)
von: Törnquist, Elin, et al.
Veröffentlicht: (2024)
Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
von: Bouthors, Maxime, et al.
Veröffentlicht: (2025)
von: Bouthors, Maxime, et al.
Veröffentlicht: (2025)
Synthetic Voice Data for Automatic Speech Recognition in African Languages
von: DeRenzi, Brian, et al.
Veröffentlicht: (2025)
von: DeRenzi, Brian, et al.
Veröffentlicht: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
Synthia: Scalable Grounded Persona Generation from Social Media Data
von: Rahimzadeh, Vahid, et al.
Veröffentlicht: (2025)
von: Rahimzadeh, Vahid, et al.
Veröffentlicht: (2025)
ConPET: Continual Parameter-Efficient Tuning for Large Language Models
von: Song, Chenyang, et al.
Veröffentlicht: (2023)
von: Song, Chenyang, et al.
Veröffentlicht: (2023)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning
von: Zou, Run, et al.
Veröffentlicht: (2026)
von: Zou, Run, et al.
Veröffentlicht: (2026)
Vocabulary Transfer for Biomedical Texts: Add Tokens if You Can Not Add Data
von: Singh, Priyanka, et al.
Veröffentlicht: (2022)
von: Singh, Priyanka, et al.
Veröffentlicht: (2022)
Qomhra: A Bilingual Irish and English Large Language Model
von: McInerney, Joseph, et al.
Veröffentlicht: (2025)
von: McInerney, Joseph, et al.
Veröffentlicht: (2025)
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data
von: Lajčinová, Bibiána, et al.
Veröffentlicht: (2024)
von: Lajčinová, Bibiána, et al.
Veröffentlicht: (2024)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
von: Stacey, Joe, et al.
Veröffentlicht: (2025)
von: Stacey, Joe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
von: Kocbek, Primož, et al.
Veröffentlicht: (2025) -
Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
von: Klerings, Alina, et al.
Veröffentlicht: (2025) -
Dialect Normalization using Large Language Models and Morphological Rules
von: Dimakis, Antonios, et al.
Veröffentlicht: (2025) -
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
von: Dang, Thao Anh, et al.
Veröffentlicht: (2024) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)