From 124 Million Tokens to 1,021 Neologisms: A Large-Scale Pipeline for Automatic Neologism Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Rossini, Diego, van der Plas, Lonneke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Binary Token-Level Classification with DeBERTa for All-Type MWE Identification: A Lightweight Approach with Linguistic Enhancement
by: Rossini, Diego, et al.
Published: (2026)
by: Rossini, Diego, et al.
Published: (2026)
Neologism Learning for Controllability and Self-Verbalization
by: Hewitt, John, et al.
Published: (2025)
by: Hewitt, John, et al.
Published: (2025)
NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms
by: Zheng, Jonathan, et al.
Published: (2024)
by: Zheng, Jonathan, et al.
Published: (2024)
Understanding the effects of language-specific class imbalance in multilingual fine-tuning
by: Jung, Vincent, et al.
Published: (2024)
by: Jung, Vincent, et al.
Published: (2024)
Can language models learn analogical reasoning? Investigating training objectives and comparisons to human performance
by: Petersen, Molly R., et al.
Published: (2023)
by: Petersen, Molly R., et al.
Published: (2023)
Evaluating Creative Short Story Generation in Humans and Large Language Models
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
NeoN: A Tool for Automated Detection, Linguistic and LLM-Driven Analysis of Neologisms in Polish
by: Tomaszewska, Aleksandra, et al.
Published: (2025)
by: Tomaszewska, Aleksandra, et al.
Published: (2025)
Lexicography of Coronavirus-related Neologisms
Published: (2023)
Published: (2023)
NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning
by: Miao, Zhongtao, et al.
Published: (2026)
by: Miao, Zhongtao, et al.
Published: (2026)
Neologism Learning as a Parameter-Efficient Alternative to Fine-Tuning for Model Steering
by: Park, Sungjoon, et al.
Published: (2025)
by: Park, Sungjoon, et al.
Published: (2025)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
Do LLMs Know What Luxembourgish Borrows? Probing Lexical Neology in Low-Resource Multilingual Models
by: Hosseini-Kivanani, Nina
Published: (2026)
by: Hosseini-Kivanani, Nina
Published: (2026)
Modelling Analogies and Analogical Reasoning: Connecting Cognitive Science Theory and NLP Research
by: Petersen, Molly R, et al.
Published: (2025)
by: Petersen, Molly R, et al.
Published: (2025)
Creativity in AI: Progresses and Challenges
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
by: Ismayilzada, Mete, et al.
Published: (2026)
by: Ismayilzada, Mete, et al.
Published: (2026)
Neologisms in Hungarian terms of quality assurance
by: Réka Sólyom
Published: (2020)
by: Réka Sólyom
Published: (2020)
A Morphological Analysis of Neologisms in Social Media English: A Study of Word Formation Processes
by: Naila Umar, et al.
Published: (2026)
by: Naila Umar, et al.
Published: (2026)
Evaluating Morphological Compositional Generalization in Large Language Models
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
Creative Preference Optimization
by: Ismayilzada, Mete, et al.
Published: (2025)
by: Ismayilzada, Mete, et al.
Published: (2025)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
The Language of Cryptocurrencies: Frequent Words, Neologisms, Acronyms, and Metaphors
by: Ricardo Casañ-Pitarch
Published: (2023)
by: Ricardo Casañ-Pitarch
Published: (2023)
Specialized Terminology in the Video Game Industry: Neologisms and their Translation
by: Ramón Méndez González
Published: (2019)
by: Ramón Méndez González
Published: (2019)
An Autistic “Linguatype”? Neologisms, New Words, and New Insights
by: Emily Zane, et al.
Published: (2025)
by: Emily Zane, et al.
Published: (2025)
Large Language Models Align with the Human Brain during Creative Thinking
by: Ismayilzada, Mete, et al.
Published: (2026)
by: Ismayilzada, Mete, et al.
Published: (2026)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
by: He, Linda, et al.
Published: (2025)
by: He, Linda, et al.
Published: (2025)
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
by: Land, Sander, et al.
Published: (2024)
by: Land, Sander, et al.
Published: (2024)
HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic Claims
by: van der Meer, Michiel, et al.
Published: (2025)
by: van der Meer, Michiel, et al.
Published: (2025)
Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs
by: Synk, Ryan, et al.
Published: (2025)
by: Synk, Ryan, et al.
Published: (2025)
LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
by: Ding, Yiran, et al.
Published: (2024)
by: Ding, Yiran, et al.
Published: (2024)
Loose and Tight: Creative Formation but Rigid Use of Nominal Compounds in Conspiracist Texts
by: Alessandro Miani, et al.
Published: (2024)
by: Alessandro Miani, et al.
Published: (2024)
KARRIEREWEGE: A Large Scale Career Path Prediction Dataset
by: Senger, Elena, et al.
Published: (2024)
by: Senger, Elena, et al.
Published: (2024)
Large Language Models in the Abuse Detection Pipeline
by: Kath, Suraj, et al.
Published: (2026)
by: Kath, Suraj, et al.
Published: (2026)
Refract ICL: Rethinking Example Selection in the Era of Million-Token Models
by: Akula, Arjun R., et al.
Published: (2025)
by: Akula, Arjun R., et al.
Published: (2025)
Out of the Memory Barrier: A Highly Memory Efficient Training System for LLMs with Million-Token Contexts
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
by: Lu, Zhiyuan, et al.
Published: (2026)
by: Lu, Zhiyuan, et al.
Published: (2026)
Constructing a BPE Tokenization DFA
by: Berglund, Martin, et al.
Published: (2024)
by: Berglund, Martin, et al.
Published: (2024)
Automatic Construction of a Legal Citation Graph from 100 Million Ukrainian Court Decisions: Large-Scale Extraction, Topological Analysis, and Ontology-Driven Clustering
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
TableRAG: Million-Token Table Understanding with Language Models
by: Chen, Si-An, et al.
Published: (2024)
by: Chen, Si-An, et al.
Published: (2024)
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
by: Lab, Mind, et al.
Published: (2026)
by: Lab, Mind, et al.
Published: (2026)
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
by: Tavakoli, Mohammad, et al.
Published: (2025)
by: Tavakoli, Mohammad, et al.
Published: (2025)
Similar Items
-
Binary Token-Level Classification with DeBERTa for All-Type MWE Identification: A Lightweight Approach with Linguistic Enhancement
by: Rossini, Diego, et al.
Published: (2026) -
Neologism Learning for Controllability and Self-Verbalization
by: Hewitt, John, et al.
Published: (2025) -
NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms
by: Zheng, Jonathan, et al.
Published: (2024) -
Understanding the effects of language-specific class imbalance in multilingual fine-tuning
by: Jung, Vincent, et al.
Published: (2024) -
Can language models learn analogical reasoning? Investigating training objectives and comparisons to human performance
by: Petersen, Molly R., et al.
Published: (2023)