Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Bochkov, A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate
von: Bochkov, A.
Veröffentlicht: (2025)
von: Bochkov, A.
Veröffentlicht: (2025)
Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes
von: Bochkov, A.
Veröffentlicht: (2026)
von: Bochkov, A.
Veröffentlicht: (2026)
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
von: Chen, Daiwei, et al.
Veröffentlicht: (2026)
von: Chen, Daiwei, et al.
Veröffentlicht: (2026)
Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala
von: Rajapakse, Minuri, et al.
Veröffentlicht: (2026)
von: Rajapakse, Minuri, et al.
Veröffentlicht: (2026)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
von: Damani, Mehul, et al.
Veröffentlicht: (2025)
von: Damani, Mehul, et al.
Veröffentlicht: (2025)
Emergent Representations of Program Semantics in Language Models Trained on Programs
von: Jin, Charles, et al.
Veröffentlicht: (2023)
von: Jin, Charles, et al.
Veröffentlicht: (2023)
Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon
von: Prashanth, USVSN Sai, et al.
Veröffentlicht: (2024)
von: Prashanth, USVSN Sai, et al.
Veröffentlicht: (2024)
Moving Beyond Next-Token Prediction: Transformers are Context-Sensitive Language Generators
von: Rhee, Phill Kyu
Veröffentlicht: (2025)
von: Rhee, Phill Kyu
Veröffentlicht: (2025)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
von: Pang, Ziqi, et al.
Veröffentlicht: (2023)
von: Pang, Ziqi, et al.
Veröffentlicht: (2023)
Static Word Embeddings for Sentence Semantic Representation
von: Wada, Takashi, et al.
Veröffentlicht: (2025)
von: Wada, Takashi, et al.
Veröffentlicht: (2025)
From Tokens to Lattices: Emergent Lattice Structures in Language Models
von: Xiong, Bo, et al.
Veröffentlicht: (2025)
von: Xiong, Bo, et al.
Veröffentlicht: (2025)
SBERT studies Meaning Representations: Decomposing Sentence Embeddings into Explainable Semantic Features
von: Opitz, Juri, et al.
Veröffentlicht: (2022)
von: Opitz, Juri, et al.
Veröffentlicht: (2022)
Great Memory, Shallow Reasoning: Limits of $k$NN-LMs
von: Geng, Shangyi, et al.
Veröffentlicht: (2024)
von: Geng, Shangyi, et al.
Veröffentlicht: (2024)
Semantic-Driven Topic Modeling Using Transformer-Based Embeddings and Clustering Algorithms
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2024)
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2024)
Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs
von: Chatterjee, Sagnik, et al.
Veröffentlicht: (2026)
von: Chatterjee, Sagnik, et al.
Veröffentlicht: (2026)
Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMs
von: Trott, Sean, et al.
Veröffentlicht: (2026)
von: Trott, Sean, et al.
Veröffentlicht: (2026)
Semantic Containment as a Fundamental Property of Emergent Misalignment
von: Saxena, Rohan
Veröffentlicht: (2026)
von: Saxena, Rohan
Veröffentlicht: (2026)
Semantic Tokens in Retrieval Augmented Generation
von: Suro, Joel
Veröffentlicht: (2024)
von: Suro, Joel
Veröffentlicht: (2024)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
von: Hyeon, Sieun, et al.
Veröffentlicht: (2024)
von: Hyeon, Sieun, et al.
Veröffentlicht: (2024)
Epicure: Navigating the Emergent Geometry of Food Ingredient Embeddings
von: Radzikowski, Jakub, et al.
Veröffentlicht: (2026)
von: Radzikowski, Jakub, et al.
Veröffentlicht: (2026)
Beyond the Needle's Illusion: Decoupled Evaluation of Evidence Access and Use under Semantic Interference at 326M-Token Scale
von: Lin, Tianwei, et al.
Veröffentlicht: (2026)
von: Lin, Tianwei, et al.
Veröffentlicht: (2026)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
von: Beltoft, Stine Lyngsø, et al.
Veröffentlicht: (2026)
von: Beltoft, Stine Lyngsø, et al.
Veröffentlicht: (2026)
Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
von: Bajaj, Anooshka, et al.
Veröffentlicht: (2025)
von: Bajaj, Anooshka, et al.
Veröffentlicht: (2025)
DSL-LLaDA: Scaling Continuous Denoising to 8B Masked Diffusion LMs
von: Yu, Longxuan, et al.
Veröffentlicht: (2026)
von: Yu, Longxuan, et al.
Veröffentlicht: (2026)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
von: Wang, Duo, et al.
Veröffentlicht: (2024)
von: Wang, Duo, et al.
Veröffentlicht: (2024)
Text Generation Beyond Discrete Token Sampling
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
Interpretable Emergent Language Using Inter-Agent Transformers
von: Bhardwaj, Mannan
Veröffentlicht: (2025)
von: Bhardwaj, Mannan
Veröffentlicht: (2025)
Lost in Translation: The Algorithmic Gap Between LMs and the Brain
von: Tosato, Tommaso, et al.
Veröffentlicht: (2024)
von: Tosato, Tommaso, et al.
Veröffentlicht: (2024)
TexIm FAST: Text-to-Image Representation for Semantic Similarity Evaluation using Transformers
von: Ansar, Wazib, et al.
Veröffentlicht: (2024)
von: Ansar, Wazib, et al.
Veröffentlicht: (2024)
Learning Evidence Highlighting for Frozen LLMs
von: Li, Shaoang, et al.
Veröffentlicht: (2026)
von: Li, Shaoang, et al.
Veröffentlicht: (2026)
Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models
von: Chen, Kexin, et al.
Veröffentlicht: (2025)
von: Chen, Kexin, et al.
Veröffentlicht: (2025)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
von: Lansiaux, Edouard, et al.
Veröffentlicht: (2025)
von: Lansiaux, Edouard, et al.
Veröffentlicht: (2025)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
von: Aneja, Krishak, et al.
Veröffentlicht: (2026)
von: Aneja, Krishak, et al.
Veröffentlicht: (2026)
VISTA: Visualization of Token Attribution via Efficient Analysis
von: Ahmed, Syed, et al.
Veröffentlicht: (2026)
von: Ahmed, Syed, et al.
Veröffentlicht: (2026)
Where Should I Study? Biased Language Models Decide! Evaluating Fairness in LMs for Academic Recommendations
von: Shailya, Krithi, et al.
Veröffentlicht: (2025)
von: Shailya, Krithi, et al.
Veröffentlicht: (2025)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
von: Trhlik, Filip, et al.
Veröffentlicht: (2026)
von: Trhlik, Filip, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate
von: Bochkov, A.
Veröffentlicht: (2025) -
Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes
von: Bochkov, A.
Veröffentlicht: (2026) -
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
von: Chen, Daiwei, et al.
Veröffentlicht: (2026) -
Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala
von: Rajapakse, Minuri, et al.
Veröffentlicht: (2026) -
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
von: Damani, Mehul, et al.
Veröffentlicht: (2025)