FoNE: Precise Single-Token Number Embeddings via Fourier Features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Tianyi, Fu, Deqing, Soltanolkotabi, Mahdi, Jia, Robin, Sharan, Vatsal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pre-trained Large Language Models Use Fourier Features to Compute Addition
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
Transformers Learn Low Sensitivity Functions: Investigations and Implications
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
von: Ye, Qilin, et al.
Veröffentlicht: (2025)
von: Ye, Qilin, et al.
Veröffentlicht: (2025)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
TLDR: Token-Level Detective Reward Model for Large Vision Language Models
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
von: Li, Ziyue, et al.
Veröffentlicht: (2024)
von: Li, Ziyue, et al.
Veröffentlicht: (2024)
Luna-2: Scalable Single-Token Evaluation with Small Language Models
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
von: Goel, Vatsal, et al.
Veröffentlicht: (2026)
Are LLM Decisions Faithful to Verbal Confidence?
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
Latent Concept Disentanglement in Transformer-based Language Models
von: Hong, Guan Zhe, et al.
Veröffentlicht: (2025)
von: Hong, Guan Zhe, et al.
Veröffentlicht: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
von: Dobler, Konstantin, et al.
Veröffentlicht: (2025)
Limitations on Accurate, Trusted, Human-level Reasoning
von: Panigrahy, Rina, et al.
Veröffentlicht: (2025)
von: Panigrahy, Rina, et al.
Veröffentlicht: (2025)
FoQA: A Faroese Question-Answering Dataset
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
von: Simonsen, Annika, et al.
Veröffentlicht: (2025)
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
von: Kuratov, Yuri, et al.
Veröffentlicht: (2025)
von: Kuratov, Yuri, et al.
Veröffentlicht: (2025)
Measuring Intrinsic Dimension of Token Embeddings
von: Kataiwa, Takuya, et al.
Veröffentlicht: (2025)
von: Kataiwa, Takuya, et al.
Veröffentlicht: (2025)
Attention with Trained Embeddings Provably Selects Important Tokens
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
von: Ramjee, Sharan
Veröffentlicht: (2026)
von: Ramjee, Sharan
Veröffentlicht: (2026)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
von: Dadgarnia, Alireza, et al.
Veröffentlicht: (2026)
von: Dadgarnia, Alireza, et al.
Veröffentlicht: (2026)
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
Simultaneous Swap Regret Minimization via KL-Calibration
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
Unsupervised Text Segmentation via Kernel Change-Point Detection on Sentence Embeddings
von: Jia, Mumin, et al.
Veröffentlicht: (2026)
von: Jia, Mumin, et al.
Veröffentlicht: (2026)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
von: Ding, Xueying, et al.
Veröffentlicht: (2025)
von: Ding, Xueying, et al.
Veröffentlicht: (2025)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
Can Public LLMs be used for Self-Diagnosis of Medical Conditions ?
von: Balasubramanian, Nikil Sharan Prabahar, et al.
Veröffentlicht: (2024)
von: Balasubramanian, Nikil Sharan Prabahar, et al.
Veröffentlicht: (2024)
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
von: Godbole, Ameya, et al.
Veröffentlicht: (2025)
von: Godbole, Ameya, et al.
Veröffentlicht: (2025)
Language Model Training Paradigms for Clinical Feature Embeddings
von: Hu, Yurong, et al.
Veröffentlicht: (2023)
von: Hu, Yurong, et al.
Veröffentlicht: (2023)
Understanding Token Probability Encoding in Output Embeddings
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
von: Cho, Hakaze, et al.
Veröffentlicht: (2024)
Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence
von: Işık, İlker, et al.
Veröffentlicht: (2024)
von: Işık, İlker, et al.
Veröffentlicht: (2024)
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
von: Feng, Weitao, et al.
Veröffentlicht: (2025)
von: Feng, Weitao, et al.
Veröffentlicht: (2025)
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
Learning to Recall with Transformers Beyond Orthogonal Embeddings
von: Vural, Nuri Mert, et al.
Veröffentlicht: (2026)
von: Vural, Nuri Mert, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Pre-trained Large Language Models Use Fourier Features to Compute Addition
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024) -
Convergent Evolution: How Different Language Models Learn Similar Number Representations
von: Fu, Deqing, et al.
Veröffentlicht: (2026) -
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
von: Fu, Deqing, et al.
Veröffentlicht: (2023) -
Transformers Learn Low Sensitivity Functions: Investigations and Implications
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024) -
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
von: Ye, Qilin, et al.
Veröffentlicht: (2025)