Transferring Extreme Subword Style Using Ngram Model-Based Logit Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Messner, Craig, Lippincott, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus
by: Messner, Craig, et al.
Published: (2024)
by: Messner, Craig, et al.
Published: (2024)
Pairing Orthographically Variant Literary Words to Standard Equivalents Using Neural Edit Distance Models
by: Messner, Craig, et al.
Published: (2024)
by: Messner, Craig, et al.
Published: (2024)
Pretraining Language Models for Diachronic Linguistic Change Discovery
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
Graph-Convolutional Autoencoder Ensembles for the Humanities, Illustrated with a Study of the American Slave Trade
by: Lippincott, Tom
Published: (2024)
by: Lippincott, Tom
Published: (2024)
Dynamic embedded topic models and change-point detection for exploring literary-historical hypotheses
by: Sirin, Hale, et al.
Published: (2024)
by: Sirin, Hale, et al.
Published: (2024)
Detecting Structured Language Alternations in Historical Documents by Combining Language Identification with Fourier Analysis
by: Sirin, Hale, et al.
Published: (2024)
by: Sirin, Hale, et al.
Published: (2024)
A Systematic Analysis of Subwords and Cross-Lingual Transfer in Multilingual Translation
by: Meyer, Francois, et al.
Published: (2024)
by: Meyer, Francois, et al.
Published: (2024)
Characterizing the Effects of Translation on Intertextuality using Multilingual Embedding Spaces
by: McGovern, Hope, et al.
Published: (2025)
by: McGovern, Hope, et al.
Published: (2025)
Computational Discovery of Chiasmus in Ancient Religious Text
by: McGovern, Hope, et al.
Published: (2025)
by: McGovern, Hope, et al.
Published: (2025)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
by: Batsuren, Khuyagbaatar, et al.
Published: (2024)
by: Batsuren, Khuyagbaatar, et al.
Published: (2024)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Lexically Grounded Subword Segmentation
by: Libovický, Jindřich, et al.
Published: (2024)
by: Libovický, Jindřich, et al.
Published: (2024)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
Limits of n-gram Style Control for LLMs via Logit-Space Injection
by: Ahmed, Sami-ul
Published: (2026)
by: Ahmed, Sami-ul
Published: (2026)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
by: Zouhar, Vilém
Published: (2024)
by: Zouhar, Vilém
Published: (2024)
Tokenization Falling Short: On Subword Robustness in Large Language Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
ByteSpan: Information-Driven Subword Tokenisation
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
Subword Tokenization Strategies for Kurdish Word Embeddings
by: Salehi, Ali, et al.
Published: (2025)
by: Salehi, Ali, et al.
Published: (2025)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
by: McGovern, Hope, et al.
Published: (2026)
by: McGovern, Hope, et al.
Published: (2026)
Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?
by: Zhang, Crystina, et al.
Published: (2024)
by: Zhang, Crystina, et al.
Published: (2024)
Dynamic Embedded Topic Models: properties and recommendations based on diverse corpora
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
Understanding Subword Compositionality of Large Language Models
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay
by: Altinok, Duygu
Published: (2026)
by: Altinok, Duygu
Published: (2026)
Subword models struggle with word learning, but surprisal hides it
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
by: Meyer, Francois, et al.
Published: (2025)
by: Meyer, Francois, et al.
Published: (2025)
On the Effect of (Near) Duplicate Subwords in Language Modelling
by: Schäfer, Anton, et al.
Published: (2024)
by: Schäfer, Anton, et al.
Published: (2024)
Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
by: Gigant, Théo, et al.
Published: (2026)
by: Gigant, Théo, et al.
Published: (2026)
Learning Mutually Informed Representations for Characters and Subwords
by: Wang, Yilin, et al.
Published: (2023)
by: Wang, Yilin, et al.
Published: (2023)
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
by: Sims, Anya, et al.
Published: (2025)
by: Sims, Anya, et al.
Published: (2025)
Replacing Language Model for Style Transfer
by: Cheng, Pengyu, et al.
Published: (2022)
by: Cheng, Pengyu, et al.
Published: (2022)
VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
by: Shao, Jintian, et al.
Published: (2025)
by: Shao, Jintian, et al.
Published: (2025)
LogitLens4LLMs: Extending Logit Lens Analysis to Modern Large Language Models
by: Wang, Zhenyu
Published: (2025)
by: Wang, Zhenyu
Published: (2025)
Style-Specific Neurons for Steering LLMs in Text Style Transfer
by: Lai, Wen, et al.
Published: (2024)
by: Lai, Wen, et al.
Published: (2024)
Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
by: Oh, Byung-Doh, et al.
Published: (2024)
by: Oh, Byung-Doh, et al.
Published: (2024)
Can Pretrained Language Models Derive Correct Semantics from Corrupt Subwords under Noise?
by: Li, Xinzhe, et al.
Published: (2023)
by: Li, Xinzhe, et al.
Published: (2023)
Improving Text Style Transfer using Masked Diffusion Language Models with Inference-time Scaling
by: Padole, Tejomay Kishor, et al.
Published: (2025)
by: Padole, Tejomay Kishor, et al.
Published: (2025)
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
by: Tsuji, Kohei, et al.
Published: (2024)
by: Tsuji, Kohei, et al.
Published: (2024)
A Subword Embedding Approach for Variation Detection in Luxembourgish User Comments
by: Lutgen, Anne-Marie, et al.
Published: (2026)
by: Lutgen, Anne-Marie, et al.
Published: (2026)
Assessing the Importance of Frequency versus Compositionality for Subword-based Tokenization in NMT
by: Wolleb, Benoist, et al.
Published: (2023)
by: Wolleb, Benoist, et al.
Published: (2023)
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
Similar Items
-
Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus
by: Messner, Craig, et al.
Published: (2024) -
Pairing Orthographically Variant Literary Words to Standard Equivalents Using Neural Edit Distance Models
by: Messner, Craig, et al.
Published: (2024) -
Pretraining Language Models for Diachronic Linguistic Change Discovery
by: Fittschen, Elisabeth, et al.
Published: (2025) -
Graph-Convolutional Autoencoder Ensembles for the Humanities, Illustrated with a Study of the American Slave Trade
by: Lippincott, Tom
Published: (2024) -
Dynamic embedded topic models and change-point detection for exploring literary-historical hypotheses
by: Sirin, Hale, et al.
Published: (2024)