Measuring Intrinsic Dimension of Token Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Kataiwa, Takuya, Hakaze, Cho, Ohki, Tetsushi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Token Probability Encoding in Output Embeddings
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
Token-based Decision Criteria Are Suboptimal in In-context Learning
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
Affinity and Diversity: A Unified Metric for Demonstration Selection via Internal Representations
by: Kato, Mariko, et al.
Published: (2025)
by: Kato, Mariko, et al.
Published: (2025)
Mechanism of Task-oriented Information Removal in In-context Learning
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Revisiting In-context Learning Inference Circuit in Large Language Models
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
Mechanistic Fine-tuning for In-context Learning
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
The Randomness Floor: Measuring Intrinsic Non-Randomness in Language Model Token Distributions
by: Hryszko, Jarosław
Published: (2026)
by: Hryszko, Jarosław
Published: (2026)
Token Distillation: Attention-aware Input Embeddings For New Tokens
by: Dobler, Konstantin, et al.
Published: (2025)
by: Dobler, Konstantin, et al.
Published: (2025)
Attention with Trained Embeddings Provably Selects Important Tokens
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions
by: Yoon, Jinsung, et al.
Published: (2024)
by: Yoon, Jinsung, et al.
Published: (2024)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
LabellessFace: Fair Metric Learning for Face Recognition without Attribute Labels
by: Ohki, Tetsushi, et al.
Published: (2024)
by: Ohki, Tetsushi, et al.
Published: (2024)
Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representations
by: Basile, Lorenzo, et al.
Published: (2024)
by: Basile, Lorenzo, et al.
Published: (2024)
Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story
by: Pedashenko, Vladislav, et al.
Published: (2025)
by: Pedashenko, Vladislav, et al.
Published: (2025)
Enhancing Remote Adversarial Patch Attacks on Face Detectors with Tiling and Scaling
by: Okano, Masora, et al.
Published: (2024)
by: Okano, Masora, et al.
Published: (2024)
Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
by: Ding, Xueying, et al.
Published: (2025)
by: Ding, Xueying, et al.
Published: (2025)
FoNE: Precise Single-Token Number Embeddings via Fourier Features
by: Zhou, Tianyi, et al.
Published: (2025)
by: Zhou, Tianyi, et al.
Published: (2025)
StaICC: Standardized Evaluation for Classification Task in In-context Learning
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
by: Chiang, Ting-Rui, et al.
Published: (2025)
by: Chiang, Ting-Rui, et al.
Published: (2025)
Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence
by: Işık, İlker, et al.
Published: (2024)
by: Işık, İlker, et al.
Published: (2024)
TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
by: Altıntaş, Gül Sena, et al.
Published: (2025)
by: Altıntaş, Gül Sena, et al.
Published: (2025)
Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information
by: Hu, Zhengmian, et al.
Published: (2023)
by: Hu, Zhengmian, et al.
Published: (2023)
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
by: Kuratov, Yuri, et al.
Published: (2025)
by: Kuratov, Yuri, et al.
Published: (2025)
Harmonic Token Projection (HTP): A Vocabulary-Free, Training-Free, Deterministic, and Reversible Embedding Methodology
by: Schmitz, Tcharlies
Published: (2025)
by: Schmitz, Tcharlies
Published: (2025)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
by: Kumar, Gokul Karthik, et al.
Published: (2026)
by: Kumar, Gokul Karthik, et al.
Published: (2026)
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
Divergent Token Metrics: Measuring degradation to prune away LLM components -- and optimize quantization
by: Deiseroth, Björn, et al.
Published: (2023)
by: Deiseroth, Björn, et al.
Published: (2023)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
by: Samragh, Mohammad, et al.
Published: (2025)
by: Samragh, Mohammad, et al.
Published: (2025)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
by: Berezin, Sergei, et al.
Published: (2025)
by: Berezin, Sergei, et al.
Published: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Memory Tokens: Large Language Models Can Generate Reversible Sentence Embeddings
by: Sastre, Ignacio, et al.
Published: (2025)
by: Sastre, Ignacio, et al.
Published: (2025)
IRIS: Intrinsic Reward Image Synthesis
by: Chen, Yihang, et al.
Published: (2025)
by: Chen, Yihang, et al.
Published: (2025)
The Proxy Presumption: From Semantic Embeddings to Valid Social Measures
by: Li, Baishi, et al.
Published: (2026)
by: Li, Baishi, et al.
Published: (2026)
From Wide to Deep: Dimension Lifting Network for Parameter-efficient Knowledge Graph Embedding
by: Cai, Borui, et al.
Published: (2023)
by: Cai, Borui, et al.
Published: (2023)
Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Mixture of Experts Made Intrinsically Interpretable
by: Yang, Xingyi, et al.
Published: (2025)
by: Yang, Xingyi, et al.
Published: (2025)
TokenShapley: Token Level Context Attribution with Shapley Value
by: Xiao, Yingtai, et al.
Published: (2025)
by: Xiao, Yingtai, et al.
Published: (2025)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2026)
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2026)
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
by: Razzhigaev, Anton, et al.
Published: (2023)
by: Razzhigaev, Anton, et al.
Published: (2023)
Similar Items
-
Understanding Token Probability Encoding in Output Embeddings
by: Cho, Hakaze, et al.
Published: (2024) -
Token-based Decision Criteria Are Suboptimal in In-context Learning
by: Cho, Hakaze, et al.
Published: (2024) -
Affinity and Diversity: A Unified Metric for Demonstration Selection via Internal Representations
by: Kato, Mariko, et al.
Published: (2025) -
Mechanism of Task-oriented Information Removal in In-context Learning
by: Cho, Hakaze, et al.
Published: (2025) -
Revisiting In-context Learning Inference Circuit in Large Language Models
by: Cho, Hakaze, et al.
Published: (2024)