The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Yanzhu, Shang, Guokan, Vazirgiannis, Michalis, Clavel, Chloé |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Linguistic Diversity of Large Language Models
von: Guo, Yanzhu, et al.
Veröffentlicht: (2024)
von: Guo, Yanzhu, et al.
Veröffentlicht: (2024)
Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
Fast-Decoding Diffusion Language Models via Progress-Aware Confidence Schedules
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
LLM as a Broken Telephone: Iterative Generation Distorts Information
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
Leveraging Discourse Structure for Extractive Meeting Summarization
von: Rennard, Virgile, et al.
Veröffentlicht: (2024)
von: Rennard, Virgile, et al.
Veröffentlicht: (2024)
Markovian Generation Chains in Large Language Models
von: Geng, Mingmeng, et al.
Veröffentlicht: (2026)
von: Geng, Mingmeng, et al.
Veröffentlicht: (2026)
GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
Graph Linearization Methods for Reasoning on Graphs with Large Language Models
von: Xypolopoulos, Christos, et al.
Veröffentlicht: (2024)
von: Xypolopoulos, Christos, et al.
Veröffentlicht: (2024)
Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts
von: Shang, Guokan, et al.
Veröffentlicht: (2025)
von: Shang, Guokan, et al.
Veröffentlicht: (2025)
Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dialect
von: Shang, Guokan, et al.
Veröffentlicht: (2024)
von: Shang, Guokan, et al.
Veröffentlicht: (2024)
Prot2Text: Multimodal Protein's Function Generation with GNNs and Transformers
von: Abdine, Hadi, et al.
Veröffentlicht: (2023)
von: Abdine, Hadi, et al.
Veröffentlicht: (2023)
CARTE: A Benchmark for Mapping Language Model Knowledge Across France
von: Carneiro, Sarah Almeida, et al.
Veröffentlicht: (2026)
von: Carneiro, Sarah Almeida, et al.
Veröffentlicht: (2026)
Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation
von: Chhun, Cyril, et al.
Veröffentlicht: (2024)
von: Chhun, Cyril, et al.
Veröffentlicht: (2024)
Bias in the Mirror: Are LLMs opinions robust to their own adversarial attacks ?
von: Rennard, Virgile, et al.
Veröffentlicht: (2024)
von: Rennard, Virgile, et al.
Veröffentlicht: (2024)
The Anatomy of Speech Persuasion: Linguistic Shifts in LLM-Modified Speeches
von: Barkar, Alisa, et al.
Veröffentlicht: (2025)
von: Barkar, Alisa, et al.
Veröffentlicht: (2025)
Graph Neural Networks on Discriminative Graphs of Words
von: Abbahaddou, Yassine, et al.
Veröffentlicht: (2024)
von: Abbahaddou, Yassine, et al.
Veröffentlicht: (2024)
Word Sense Induction with Hierarchical Clustering and Mutual Information Maximization
von: Abdine, Hadi, et al.
Veröffentlicht: (2022)
von: Abdine, Hadi, et al.
Veröffentlicht: (2022)
EmoDynamiX: Emotional Support Dialogue Strategy Prediction by Modelling MiXed Emotions and Discourse Dynamics
von: Wan, Chenwei, et al.
Veröffentlicht: (2024)
von: Wan, Chenwei, et al.
Veröffentlicht: (2024)
On the Diversity of Synthetic Data and its Impact on Training Large Language Models
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
The Curious Case of Representational Alignment: Unravelling Visio-Linguistic Tasks in Emergent Communication
von: Kouwenhoven, Tom, et al.
Veröffentlicht: (2024)
von: Kouwenhoven, Tom, et al.
Veröffentlicht: (2024)
Towards Trust Calibration in Socially Interactive Agents: Investigating Gendered Multimodal Behaviors Generation with LLMs
von: Galland, Lucie, et al.
Veröffentlicht: (2026)
von: Galland, Lucie, et al.
Veröffentlicht: (2026)
The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations
von: Soler, Aina Garí, et al.
Veröffentlicht: (2024)
von: Soler, Aina Garí, et al.
Veröffentlicht: (2024)
The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models
von: Sourati, Zhivar, et al.
Veröffentlicht: (2025)
von: Sourati, Zhivar, et al.
Veröffentlicht: (2025)
Automatic Analysis of Collaboration Through Human Conversational Data Resources: A Review
von: Yu, Yi, et al.
Veröffentlicht: (2026)
von: Yu, Yi, et al.
Veröffentlicht: (2026)
Linguistic and Argument Diversity in Synthetic Data for Function-Calling Agents
von: Greenstein, Dan, et al.
Veröffentlicht: (2026)
von: Greenstein, Dan, et al.
Veröffentlicht: (2026)
Training a Generally Curious Agent
von: Tajwar, Fahim, et al.
Veröffentlicht: (2025)
von: Tajwar, Fahim, et al.
Veröffentlicht: (2025)
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
von: Guo, Xu, et al.
Veröffentlicht: (2026)
von: Guo, Xu, et al.
Veröffentlicht: (2026)
The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models
von: Lee, Taewhoo, et al.
Veröffentlicht: (2025)
von: Lee, Taewhoo, et al.
Veröffentlicht: (2025)
A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts
von: Tripto, Nafis Irtiza, et al.
Veröffentlicht: (2023)
von: Tripto, Nafis Irtiza, et al.
Veröffentlicht: (2023)
Decoding Linguistic Nuances in Mental Health Text Classification Using Expressive Narrative Stories
von: Tang, Jinwen, et al.
Veröffentlicht: (2024)
von: Tang, Jinwen, et al.
Veröffentlicht: (2024)
"Mm, Wat?" Detecting Other-initiated Repair Requests in Dialogue
von: Ngo, Anh, et al.
Veröffentlicht: (2025)
von: Ngo, Anh, et al.
Veröffentlicht: (2025)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
A Greek Government Decisions Dataset for Public-Sector Analysis and Insight
von: Antoniou, Giorgos, et al.
Veröffentlicht: (2025)
von: Antoniou, Giorgos, et al.
Veröffentlicht: (2025)
The Curious Language Model: Strategic Test-Time Information Acquisition
von: Cooper, Michael, et al.
Veröffentlicht: (2025)
von: Cooper, Michael, et al.
Veröffentlicht: (2025)
The Curious Case of Nonverbal Abstract Reasoning with Multi-Modal Large Language Models
von: Ahrabian, Kian, et al.
Veröffentlicht: (2024)
von: Ahrabian, Kian, et al.
Veröffentlicht: (2024)
The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders
von: Sauter, Adrian, et al.
Veröffentlicht: (2025)
von: Sauter, Adrian, et al.
Veröffentlicht: (2025)
Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
von: Guo, Yanzhu, et al.
Veröffentlicht: (2024)
von: Guo, Yanzhu, et al.
Veröffentlicht: (2024)
Using Natural Language Processing and Networks to Automate Structured Literature Reviews: An Application to Farmers Climate Change Adaptation
von: Gil-Clavel, Sofia, et al.
Veröffentlicht: (2023)
von: Gil-Clavel, Sofia, et al.
Veröffentlicht: (2023)
Investigating Large Language Models' Linguistic Abilities for Text Preprocessing
von: Braga, Marco, et al.
Veröffentlicht: (2025)
von: Braga, Marco, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Benchmarking Linguistic Diversity of Large Language Models
von: Guo, Yanzhu, et al.
Veröffentlicht: (2024) -
Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text
von: Mohamed, Amr, et al.
Veröffentlicht: (2025) -
Fast-Decoding Diffusion Language Models via Progress-Aware Confidence Schedules
von: Mohamed, Amr, et al.
Veröffentlicht: (2025) -
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
von: Zhang, Yang, et al.
Veröffentlicht: (2025) -
LLM as a Broken Telephone: Iterative Generation Distorts Information
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)