Weight Tying Biases Token Embeddings Towards the Output Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lopardo, Antonio, Harish, Avyukth, Arnett, Catherine, Gupta, Akshat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Morphological Alignment of Tokenizers in 70 Languages
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates
von: Gu, Jian, et al.
Veröffentlicht: (2026)
von: Gu, Jian, et al.
Veröffentlicht: (2026)
Explaining and Mitigating Crosslingual Tokenizer Inequities
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
von: Chizhov, Pavel, et al.
Veröffentlicht: (2024)
von: Chizhov, Pavel, et al.
Veröffentlicht: (2024)
Different Tokenization Schemes Lead to Comparable Performance in Spanish Number Agreement
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
von: Long, Do Xuan, et al.
Veröffentlicht: (2024)
von: Long, Do Xuan, et al.
Veröffentlicht: (2024)
BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization
von: Land, Sander, et al.
Veröffentlicht: (2025)
von: Land, Sander, et al.
Veröffentlicht: (2025)
Lost in Space: Finding the Right Tokens for Structured Output
von: Hamilton, Sil, et al.
Veröffentlicht: (2025)
von: Hamilton, Sil, et al.
Veröffentlicht: (2025)
Understanding Token Probability Encoding in Output Embeddings
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
Why do language models perform worse for morphologically complex languages?
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
von: Michaelov, James A., et al.
Veröffentlicht: (2025)
von: Michaelov, James A., et al.
Veröffentlicht: (2025)
Investigating the Effects of Cognitive Biases in Prompts on Large Language Model Outputs
von: Sun, Yan, et al.
Veröffentlicht: (2025)
von: Sun, Yan, et al.
Veröffentlicht: (2025)
Revenge of the Fallen? Recurrent Models Match Transformers at Predicting Human Language Comprehension Metrics
von: Michaelov, James A., et al.
Veröffentlicht: (2024)
von: Michaelov, James A., et al.
Veröffentlicht: (2024)
A Bit of a Problem: Measurement Disparities in Dataset Sizes Across Languages
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues
von: Kuo, Martin, et al.
Veröffentlicht: (2025)
von: Kuo, Martin, et al.
Veröffentlicht: (2025)
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
Self-Assessment Tests are Unreliable Measures of LLM Personality
von: Gupta, Akshat, et al.
Veröffentlicht: (2023)
von: Gupta, Akshat, et al.
Veröffentlicht: (2023)
DWA-KD: Dual-Space Weighting and Time-Warped Alignment for Cross-Tokenizer Knowledge Distillation
von: Vu, Duc Trung, et al.
Veröffentlicht: (2026)
von: Vu, Duc Trung, et al.
Veröffentlicht: (2026)
Topic Modelling: Going Beyond Token Outputs
von: Williams, Lowri, et al.
Veröffentlicht: (2024)
von: Williams, Lowri, et al.
Veröffentlicht: (2024)
Goldfish: Monolingual Language Models for 350 Languages
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
Toxicity of the Commons: Curating Open-Source Pre-Training Data
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
DEPART: DEcomposing PARiTy across Multilingual LLMs
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
On the Acquisition of Shared Grammatical Representations in Bilingual Language Models
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
von: Arnett, Catherine, et al.
Veröffentlicht: (2025)
Output-Space Search: Targeting LLM Generations in a Frozen Encoder-Defined Output Space
von: Materzok, Tobias
Veröffentlicht: (2026)
von: Materzok, Tobias
Veröffentlicht: (2026)
Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks
von: Kanjirangat, Vani, et al.
Veröffentlicht: (2025)
von: Kanjirangat, Vani, et al.
Veröffentlicht: (2025)
A Unified Framework for Model Editing
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
von: Yoon, Junsang, et al.
Veröffentlicht: (2024)
von: Yoon, Junsang, et al.
Veröffentlicht: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
Token Weighting for Long-Range Language Modeling
von: Helm, Falko, et al.
Veröffentlicht: (2025)
von: Helm, Falko, et al.
Veröffentlicht: (2025)
Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2026)
von: Xu, Zihao, et al.
Veröffentlicht: (2026)
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
von: Liu, Jiacheng, et al.
Veröffentlicht: (2025)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2025)
Quantifying Gender Biases Towards Politicians on Reddit
von: Marjanovic, Sara, et al.
Veröffentlicht: (2021)
von: Marjanovic, Sara, et al.
Veröffentlicht: (2021)
Subword Tokenization Strategies for Kurdish Word Embeddings
von: Salehi, Ali, et al.
Veröffentlicht: (2025)
von: Salehi, Ali, et al.
Veröffentlicht: (2025)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
Adaptive Token Biaser: Knowledge Editing via Biasing Key Entities
von: Bi, Baolong, et al.
Veröffentlicht: (2024)
von: Bi, Baolong, et al.
Veröffentlicht: (2024)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2022)
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2022)
How Do LLMs Use Their Depth?
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating Morphological Alignment of Tokenizers in 70 Languages
von: Arnett, Catherine, et al.
Veröffentlicht: (2025) -
Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates
von: Gu, Jian, et al.
Veröffentlicht: (2026) -
Explaining and Mitigating Crosslingual Tokenizer Inequities
von: Arnett, Catherine, et al.
Veröffentlicht: (2025) -
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
von: Chizhov, Pavel, et al.
Veröffentlicht: (2024) -
Different Tokenization Schemes Lead to Comparable Performance in Spanish Number Agreement
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)