Pretraining Language Models for Diachronic Linguistic Change Discovery
Fuente:
arXiv
Saved in:
| Main Authors: | Fittschen, Elisabeth, Li, Sabrina, Lippincott, Tom, Choshen, Leshem, Messner, Craig |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus
by: Messner, Craig, et al.
Published: (2024)
by: Messner, Craig, et al.
Published: (2024)
Transferring Extreme Subword Style Using Ngram Model-Based Logit Scaling
by: Messner, Craig, et al.
Published: (2025)
by: Messner, Craig, et al.
Published: (2025)
Pairing Orthographically Variant Literary Words to Standard Equivalents Using Neural Edit Distance Models
by: Messner, Craig, et al.
Published: (2024)
by: Messner, Craig, et al.
Published: (2024)
Dynamic Embedded Topic Models: properties and recommendations based on diverse corpora
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Detecting Structured Language Alternations in Historical Documents by Combining Language Identification with Fourier Analysis
by: Sirin, Hale, et al.
Published: (2024)
by: Sirin, Hale, et al.
Published: (2024)
Instructions Shape Production of Language, not Processing
by: Waldis, Andreas, et al.
Published: (2026)
by: Waldis, Andreas, et al.
Published: (2026)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)
by: Zaman, Kerem, et al.
Published: (2023)
Can Gradient Descent Simulate Prompting?
by: Zhang, Eric, et al.
Published: (2025)
by: Zhang, Eric, et al.
Published: (2025)
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
by: Akyürek, Afra Feyza, et al.
Published: (2024)
by: Akyürek, Afra Feyza, et al.
Published: (2024)
Naturally Occurring Feedback is Common, Extractable and Useful
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
A Hitchhiker's Guide to Scaling Law Estimation
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
Computational Discovery of Chiasmus in Ancient Religious Text
by: McGovern, Hope, et al.
Published: (2025)
by: McGovern, Hope, et al.
Published: (2025)
Graph-Convolutional Autoencoder Ensembles for the Humanities, Illustrated with a Study of the American Slave Trade
by: Lippincott, Tom
Published: (2024)
by: Lippincott, Tom
Published: (2024)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
by: Din, Alexander Yom, et al.
Published: (2023)
by: Din, Alexander Yom, et al.
Published: (2023)
Mediocrity is the key for LLM as a Judge Anchor Selection
by: Don-Yehiya, Shachar, et al.
Published: (2026)
by: Don-Yehiya, Shachar, et al.
Published: (2026)
Will it Merge? On The Causes of Model Mergeability
by: Rahamim, Adir, et al.
Published: (2026)
by: Rahamim, Adir, et al.
Published: (2026)
Dynamic embedded topic models and change-point detection for exploring literary-historical hypotheses
by: Sirin, Hale, et al.
Published: (2024)
by: Sirin, Hale, et al.
Published: (2024)
Characterizing Linguistic Shifts in Croatian News via Diachronic Word Embeddings
by: Dukić, David, et al.
Published: (2025)
by: Dukić, David, et al.
Published: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
by: Ramesh, Pratik, et al.
Published: (2026)
by: Ramesh, Pratik, et al.
Published: (2026)
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Warstadt, Alex, et al.
Published: (2025)
by: Warstadt, Alex, et al.
Published: (2025)
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Hu, Michael Y., et al.
Published: (2024)
by: Hu, Michael Y., et al.
Published: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning
by: Schwartz, Eli, et al.
Published: (2024)
by: Schwartz, Eli, et al.
Published: (2024)
Do LLMs Benefit From Their Own Words?
by: Huang, Jenny Y., et al.
Published: (2026)
by: Huang, Jenny Y., et al.
Published: (2026)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
Evaluating the Effectiveness of Linguistic Knowledge in Pretrained Language Models: A Case Study of Universal Dependencies
by: Li, Wenxi
Published: (2025)
by: Li, Wenxi
Published: (2025)
Transformer-Enabled Diachronic Analysis of Vedic Sanskrit: Neural Methods for Quantifying Types of Language Change
by: Hariharan, Ananth, et al.
Published: (2025)
by: Hariharan, Ananth, et al.
Published: (2025)
Label-Efficient Model Selection for Text Generation
by: Ashury-Tahan, Shir, et al.
Published: (2024)
by: Ashury-Tahan, Shir, et al.
Published: (2024)
Exploring Diachronic and Diatopic Changes in Dialect Continua: Tasks, Datasets and Challenges
by: Çelikkol, Melis, et al.
Published: (2024)
by: Çelikkol, Melis, et al.
Published: (2024)
Turkronicles: Diachronic Resources for the Fast Evolving Turkish Language
by: Yazar, Togay, et al.
Published: (2024)
by: Yazar, Togay, et al.
Published: (2024)
An Embedded Diachronic Sense Change Model with a Case Study from Ancient Greek
by: Zafar, Schyan, et al.
Published: (2023)
by: Zafar, Schyan, et al.
Published: (2023)
Modelling the Diachronic Emergence of Phoneme Frequency Distributions
by: Martín, Fermín Moscoso del Prado, et al.
Published: (2026)
by: Martín, Fermín Moscoso del Prado, et al.
Published: (2026)
Characterizing the Effects of Translation on Intertextuality using Multilingual Embedding Spaces
by: McGovern, Hope, et al.
Published: (2025)
by: McGovern, Hope, et al.
Published: (2025)
CRISP: Complex Reasoning with Interpretable Step-based Plans
by: Vetzler, Matan, et al.
Published: (2025)
by: Vetzler, Matan, et al.
Published: (2025)
Efficient Benchmarking of Language Models
by: Perlitz, Yotam, et al.
Published: (2023)
by: Perlitz, Yotam, et al.
Published: (2023)
Unforgettable Generalization in Language Models
by: Zhang, Eric, et al.
Published: (2024)
by: Zhang, Eric, et al.
Published: (2024)
Similar Items
-
Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus
by: Messner, Craig, et al.
Published: (2024) -
Transferring Extreme Subword Style Using Ngram Model-Based Logit Scaling
by: Messner, Craig, et al.
Published: (2025) -
Pairing Orthographically Variant Literary Words to Standard Equivalents Using Neural Edit Distance Models
by: Messner, Craig, et al.
Published: (2024) -
Dynamic Embedded Topic Models: properties and recommendations based on diverse corpora
by: Fittschen, Elisabeth, et al.
Published: (2025) -
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
by: Waldis, Andreas, et al.
Published: (2024)