A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Nie, Zhijie, Zhang, Richong, Wu, Zhanyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
by: An, Ruize, et al.
Published: (2025)
by: An, Ruize, et al.
Published: (2025)
When Text Embedding Meets Large Language Model: A Comprehensive Survey
by: Nie, Zhijie, et al.
Published: (2024)
by: Nie, Zhijie, et al.
Published: (2024)
Learning to Select: Query-Aware Adaptive Dimension Selection for Dense Retrieval
by: Wu, Zhanyu, et al.
Published: (2026)
by: Wu, Zhanyu, et al.
Published: (2026)
Rethinking the Privacy of Text Embeddings: A Reproducibility Study of "Text Embeddings Reveal (Almost) As Much As Text"
by: Seputis, Dominykas, et al.
Published: (2025)
by: Seputis, Dominykas, et al.
Published: (2025)
LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
by: Vujanic, Robin, et al.
Published: (2025)
by: Vujanic, Robin, et al.
Published: (2025)
HeceTokenizer: A Syllable-Based Tokenization Approach for Turkish Retrieval
by: Gulgonul, Senol
Published: (2026)
by: Gulgonul, Senol
Published: (2026)
Text Clustering as Classification with LLMs
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
by: Mohr, Isabelle, et al.
Published: (2024)
by: Mohr, Isabelle, et al.
Published: (2024)
Improving Text Embeddings with Large Language Models
by: Wang, Liang, et al.
Published: (2023)
by: Wang, Liang, et al.
Published: (2023)
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
by: Chen, Jianlyu, et al.
Published: (2025)
by: Chen, Jianlyu, et al.
Published: (2025)
Multilingual E5 Text Embeddings: A Technical Report
by: Wang, Liang, et al.
Published: (2024)
by: Wang, Liang, et al.
Published: (2024)
JFinTEB: Japanese Financial Text Embedding Benchmark
by: Suzuki, Masahiro, et al.
Published: (2026)
by: Suzuki, Masahiro, et al.
Published: (2026)
Text Embeddings by Weakly-Supervised Contrastive Pre-training
by: Wang, Liang, et al.
Published: (2022)
by: Wang, Liang, et al.
Published: (2022)
FinMTEB: Finance Massive Text Embedding Benchmark
by: Tang, Yixuan, et al.
Published: (2025)
by: Tang, Yixuan, et al.
Published: (2025)
Interpretable Text Embeddings and Text Similarity Explanation: A Survey
by: Opitz, Juri, et al.
Published: (2025)
by: Opitz, Juri, et al.
Published: (2025)
ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings
by: Kasmaee, Ali Shiraee, et al.
Published: (2025)
by: Kasmaee, Ali Shiraee, et al.
Published: (2025)
Enhancing Lexicon-Based Text Embeddings with Large Language Models
by: Lei, Yibin, et al.
Published: (2025)
by: Lei, Yibin, et al.
Published: (2025)
Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
by: Babakhin, Yauhen, et al.
Published: (2025)
by: Babakhin, Yauhen, et al.
Published: (2025)
Applying Text Embedding Models for Efficient Analysis in Labeled Property Graphs
by: Podstawski, Michal
Published: (2025)
by: Podstawski, Michal
Published: (2025)
Pooling and Semantic Shift: The Fundamental Challenges in Long Text Embedding and Retrieval
by: Gao, Hang, et al.
Published: (2026)
by: Gao, Hang, et al.
Published: (2026)
Training LLMs to be Better Text Embedders through Bidirectional Reconstruction
by: Su, Chang, et al.
Published: (2025)
by: Su, Chang, et al.
Published: (2025)
Text2Cypher Across Languages: Evaluating and Finetuning LLMs
by: Ozsoy, Makbule Gulcin, et al.
Published: (2025)
by: Ozsoy, Makbule Gulcin, et al.
Published: (2025)
From Tokens to Concepts: Leveraging SAE for SPLADE
by: Zong, Yuxuan, et al.
Published: (2026)
by: Zong, Yuxuan, et al.
Published: (2026)
FIT-RAG: Black-Box RAG with Factual Information and Token Reduction
by: Mao, Yuren, et al.
Published: (2024)
by: Mao, Yuren, et al.
Published: (2024)
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
by: Merrick, Luke, et al.
Published: (2024)
by: Merrick, Luke, et al.
Published: (2024)
Rethinking the Role of Token Retrieval in Multi-Vector Retrieval
by: Lee, Jinhyuk, et al.
Published: (2023)
by: Lee, Jinhyuk, et al.
Published: (2023)
Is Semantic Chunking Worth the Computational Cost?
by: Qu, Renyi, et al.
Published: (2024)
by: Qu, Renyi, et al.
Published: (2024)
Bagging-Based Model Merging for Robust General Text Embeddings
by: Zhang, Hengran, et al.
Published: (2026)
by: Zhang, Hengran, et al.
Published: (2026)
MMTEB: Massive Multilingual Text Embedding Benchmark
by: Enevoldsen, Kenneth, et al.
Published: (2025)
by: Enevoldsen, Kenneth, et al.
Published: (2025)
TokenRec: Learning to Tokenize ID for LLM-based Generative Recommendation
by: Qu, Haohao, et al.
Published: (2024)
by: Qu, Haohao, et al.
Published: (2024)
Quantifying Positional Biases in Text Embedding Models
by: Lee, Reagan J., et al.
Published: (2024)
by: Lee, Reagan J., et al.
Published: (2024)
mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
Token and Span Classification for Entity Recognition in French Historical Encyclopedias
by: Moncla, Ludovic, et al.
Published: (2025)
by: Moncla, Ludovic, et al.
Published: (2025)
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
by: Tavakoli, Mohammad, et al.
Published: (2025)
by: Tavakoli, Mohammad, et al.
Published: (2025)
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
by: Xing, Tiancheng, et al.
Published: (2025)
by: Xing, Tiancheng, et al.
Published: (2025)
Do We Really Need Specialization? Evaluating Generalist Text Embeddings for Zero-Shot Recommendation and Search
by: Attimonelli, Matteo, et al.
Published: (2025)
by: Attimonelli, Matteo, et al.
Published: (2025)
Less is More: Adapting Text Embeddings for Low-Resource Languages with Small Scale Noisy Synthetic Data
by: Navasardyan, Zaruhi, et al.
Published: (2026)
by: Navasardyan, Zaruhi, et al.
Published: (2026)
From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications
by: Ma, Yongqiang, et al.
Published: (2024)
by: Ma, Yongqiang, et al.
Published: (2024)
Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
by: Albarede, Lucas, et al.
Published: (2025)
by: Albarede, Lucas, et al.
Published: (2025)
Training Sparse Mixture Of Experts Text Embedding Models
by: Nussbaum, Zach, et al.
Published: (2025)
by: Nussbaum, Zach, et al.
Published: (2025)
Similar Items
-
Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
by: An, Ruize, et al.
Published: (2025) -
When Text Embedding Meets Large Language Model: A Comprehensive Survey
by: Nie, Zhijie, et al.
Published: (2024) -
Learning to Select: Query-Aware Adaptive Dimension Selection for Dense Retrieval
by: Wu, Zhanyu, et al.
Published: (2026) -
Rethinking the Privacy of Text Embeddings: A Reproducibility Study of "Text Embeddings Reveal (Almost) As Much As Text"
by: Seputis, Dominykas, et al.
Published: (2025) -
LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
by: Vujanic, Robin, et al.
Published: (2025)