Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Feucht, Sheridan, Atkinson, David, Wallace, Byron, Bau, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Dual-Route Model of Induction
by: Feucht, Sheridan, et al.
Published: (2025)
by: Feucht, Sheridan, et al.
Published: (2025)
In-Context Algebra
by: Todd, Eric, et al.
Published: (2025)
by: Todd, Eric, et al.
Published: (2025)
Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure
by: Mahmood, Syed Naveed, et al.
Published: (2026)
by: Mahmood, Syed Naveed, et al.
Published: (2026)
LLM Vocabulary Compression for Low-Compute Environments
by: Vennam, Sreeram, et al.
Published: (2024)
by: Vennam, Sreeram, et al.
Published: (2024)
HYPEROFA: Expanding LLM Vocabulary to New Languages via Hypernetwork-Based Embedding Initialization
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
Position-aware Automatic Circuit Discovery
by: Haklay, Tal, et al.
Published: (2025)
by: Haklay, Tal, et al.
Published: (2025)
Language Models Are Implicitly Continuous
by: Marro, Samuele, et al.
Published: (2025)
by: Marro, Samuele, et al.
Published: (2025)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
by: Walker, Nicholas
Published: (2024)
by: Walker, Nicholas
Published: (2024)
Tethered Reasoning: Decoupling Entropy from Hallucination in Quantized LLMs via Manifold Steering
by: Atkinson, Craig
Published: (2026)
by: Atkinson, Craig
Published: (2026)
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
by: Fagnou, Erwan, et al.
Published: (2026)
by: Fagnou, Erwan, et al.
Published: (2026)
Ambiguity in LLMs is a concept missing problem
by: Hu, Zhibo, et al.
Published: (2025)
by: Hu, Zhibo, et al.
Published: (2025)
RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base
by: Shravan, Rohan
Published: (2026)
by: Shravan, Rohan
Published: (2026)
Reinforcement Learning for Latent-Space Thinking in LLMs
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly
by: Hosseini, Peyman, et al.
Published: (2024)
by: Hosseini, Peyman, et al.
Published: (2024)
The Open Source Advantage in Large Language Models (LLMs)
by: Manchanda, Jiya, et al.
Published: (2024)
by: Manchanda, Jiya, et al.
Published: (2024)
A Framework for Fine-Tuning LLMs using Heterogeneous Feedback
by: Aponte, Ryan, et al.
Published: (2024)
by: Aponte, Ryan, et al.
Published: (2024)
PowLU: An Activation Function for Stable Pre-Training of LLMs
by: Jiang, Peijie, et al.
Published: (2026)
by: Jiang, Peijie, et al.
Published: (2026)
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
by: Zhang, Zhaowei, et al.
Published: (2025)
by: Zhang, Zhaowei, et al.
Published: (2025)
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
by: Sun, Luoyang, et al.
Published: (2026)
by: Sun, Luoyang, et al.
Published: (2026)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
by: He, Yanjin, et al.
Published: (2025)
by: He, Yanjin, et al.
Published: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
by: Shravan, Rohan
Published: (2026)
by: Shravan, Rohan
Published: (2026)
Vocabulary Transfer for Biomedical Texts: Add Tokens if You Can Not Add Data
by: Singh, Priyanka, et al.
Published: (2022)
by: Singh, Priyanka, et al.
Published: (2022)
Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts
by: Chourasia, Shivam, et al.
Published: (2026)
by: Chourasia, Shivam, et al.
Published: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
PersonalLLM: Tailoring LLMs to Individual Preferences
by: Zollo, Thomas P., et al.
Published: (2024)
by: Zollo, Thomas P., et al.
Published: (2024)
Enhancing Traffic Accident Classifications: Application of NLP Methods for City Safety
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
GPT Editors, Not Authors: The Stylistic Footprint of LLMs in Academic Preprints
by: DeHaan, Soren, et al.
Published: (2025)
by: DeHaan, Soren, et al.
Published: (2025)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
by: Young, Richard J., et al.
Published: (2025)
by: Young, Richard J., et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Enhancing In-Context Learning via Implicit Demonstration Augmentation
by: Zhou, Xiaoling, et al.
Published: (2024)
by: Zhou, Xiaoling, et al.
Published: (2024)
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
by: Mamidanna, Siddarth, et al.
Published: (2025)
by: Mamidanna, Siddarth, et al.
Published: (2025)
Engineering A Large Language Model From Scratch
by: Oketunji, Abiodun Finbarrs
Published: (2024)
by: Oketunji, Abiodun Finbarrs
Published: (2024)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026)
by: Vieira, Inês, et al.
Published: (2026)
Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER
by: Ewais, Ahmed, et al.
Published: (2026)
by: Ewais, Ahmed, et al.
Published: (2026)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
by: Quevedo, Ernesto, et al.
Published: (2024)
by: Quevedo, Ernesto, et al.
Published: (2024)
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
by: He, Langzhou, et al.
Published: (2026)
by: He, Langzhou, et al.
Published: (2026)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
by: Hanna, Michael, et al.
Published: (2024)
by: Hanna, Michael, et al.
Published: (2024)
Similar Items
-
The Dual-Route Model of Induction
by: Feucht, Sheridan, et al.
Published: (2025) -
In-Context Algebra
by: Todd, Eric, et al.
Published: (2025) -
Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure
by: Mahmood, Syed Naveed, et al.
Published: (2026) -
LLM Vocabulary Compression for Low-Compute Environments
by: Vennam, Sreeram, et al.
Published: (2024) -
HYPEROFA: Expanding LLM Vocabulary to New Languages via Hypernetwork-Based Embedding Initialization
by: Özeren, Enes, et al.
Published: (2025)