Length-Induced Embedding Collapse in PLM-based Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yuqi, Dai, Sunhao, Cao, Zhanshuo, Zhang, Xiao, Xu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Learning to Retrieve from Agent Trajectories
by: Zhou, Yuqi, et al.
Published: (2026)
by: Zhou, Yuqi, et al.
Published: (2026)
Neural Retrievers are Biased Towards LLM-Generated Content
by: Dai, Sunhao, et al.
Published: (2023)
by: Dai, Sunhao, et al.
Published: (2023)
Exploring the Escalation of Source Bias in User, Data, and Recommender System Feedback Loop
by: Zhou, Yuqi, et al.
Published: (2024)
by: Zhou, Yuqi, et al.
Published: (2024)
Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
Think Before Recommend: Unleashing the Latent Reasoning Power for Sequential Recommendation
by: Tang, Jiakai, et al.
Published: (2025)
by: Tang, Jiakai, et al.
Published: (2025)
ReCODE: Modeling Repeat Consumption with Neural ODE
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
by: Merrick, Luke, et al.
Published: (2024)
by: Merrick, Luke, et al.
Published: (2024)
NExT-Search: Rebuilding User Feedback Ecosystem for Generative AI Search
by: Dai, Sunhao, et al.
Published: (2025)
by: Dai, Sunhao, et al.
Published: (2025)
OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
by: Dai, Sunhao, et al.
Published: (2025)
by: Dai, Sunhao, et al.
Published: (2025)
DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data
by: Venus Team, et al.
Published: (2026)
by: Venus Team, et al.
Published: (2026)
UOEP: User-Oriented Exploration Policy for Enhancing Long-Term User Experiences in Recommender Systems
by: Zhang, Changshuo, et al.
Published: (2024)
by: Zhang, Changshuo, et al.
Published: (2024)
Tabular Embedding Model (TEM): Finetuning Embedding Models For Tabular RAG Applications
by: Khanna, Sujit, et al.
Published: (2024)
by: Khanna, Sujit, et al.
Published: (2024)
C-Pack: Packed Resources For General Chinese Embeddings
by: Xiao, Shitao, et al.
Published: (2023)
by: Xiao, Shitao, et al.
Published: (2023)
Bagging-Based Model Merging for Robust General Text Embeddings
by: Zhang, Hengran, et al.
Published: (2026)
by: Zhang, Hengran, et al.
Published: (2026)
When Text Embedding Meets Large Language Model: A Comprehensive Survey
by: Nie, Zhijie, et al.
Published: (2024)
by: Nie, Zhijie, et al.
Published: (2024)
A Span-based Model for Extracting Overlapping PICO Entities from RCT Publications
by: Zhang, Gongbo, et al.
Published: (2024)
by: Zhang, Gongbo, et al.
Published: (2024)
Quantifying Positional Biases in Text Embedding Models
by: Lee, Reagan J., et al.
Published: (2024)
by: Lee, Reagan J., et al.
Published: (2024)
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
by: Sun, Yiqun, et al.
Published: (2025)
by: Sun, Yiqun, et al.
Published: (2025)
Training Sparse Mixture Of Experts Text Embedding Models
by: Nussbaum, Zach, et al.
Published: (2025)
by: Nussbaum, Zach, et al.
Published: (2025)
Tug-of-War Between Knowledge: Exploring and Resolving Knowledge Conflicts in Retrieval-Augmented Language Models
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
Reasoning on Efficient Knowledge Paths:Knowledge Graph Guides Large Language Model for Domain Question Answering
by: Wang, Yuqi, et al.
Published: (2024)
by: Wang, Yuqi, et al.
Published: (2024)
List-aware Reranking-Truncation Joint Model for Search and Retrieval-augmented Generation
by: Xu, Shicheng, et al.
Published: (2024)
by: Xu, Shicheng, et al.
Published: (2024)
Cutting Off the Head Ends the Conflict: A Mechanism for Interpreting and Mitigating Knowledge Conflicts in Language Models
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
Don't Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space Transformation
by: Feng, Yingchaojie, et al.
Published: (2025)
by: Feng, Yingchaojie, et al.
Published: (2025)
A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models
by: Zhang, Qinggang, et al.
Published: (2025)
by: Zhang, Qinggang, et al.
Published: (2025)
An Open-Source Dual-Loss Embedding Model for Semantic Retrieval in Higher Education
by: Sajja, Ramteja, et al.
Published: (2025)
by: Sajja, Ramteja, et al.
Published: (2025)
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
by: Wang, Guangzhi, et al.
Published: (2025)
by: Wang, Guangzhi, et al.
Published: (2025)
Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
MMTEB: Massive Multilingual Text Embedding Benchmark
by: Enevoldsen, Kenneth, et al.
Published: (2025)
by: Enevoldsen, Kenneth, et al.
Published: (2025)
LEMUR: A Corpus for Robust Fine-Tuning of Multilingual Law Embedding Models for Retrieval
by: Ahmadi, Narges Baba, et al.
Published: (2026)
by: Ahmadi, Narges Baba, et al.
Published: (2026)
One Model Is Enough: Native Retrieval Embeddings from LLM Agent Hidden States
by: Jiang, Bo
Published: (2026)
by: Jiang, Bo
Published: (2026)
ConceptFormer: Towards Efficient Use of Knowledge-Graph Embeddings in Large Language Models
by: Barmettler, Joel, et al.
Published: (2025)
by: Barmettler, Joel, et al.
Published: (2025)
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
by: Lee, Chankyu, et al.
Published: (2024)
by: Lee, Chankyu, et al.
Published: (2024)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
Robustness Risk of Conversational Retrieval: Identifying and Mitigating Noise Sensitivity in Qwen3-Embedding Model
by: Chen, Weishu, et al.
Published: (2026)
by: Chen, Weishu, et al.
Published: (2026)
The Massive Legal Embedding Benchmark (MLEB)
by: Butler, Umar, et al.
Published: (2025)
by: Butler, Umar, et al.
Published: (2025)
A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
Improving Embedding Accuracy for Document Retrieval Using Entity Relationship Maps and Model-Aware Contrastive Sampling
by: Aviss, Thea
Published: (2024)
by: Aviss, Thea
Published: (2024)
Predicting Oscar-Nominated Screenplays with Sentence Embeddings
by: Gross, Francis
Published: (2025)
by: Gross, Francis
Published: (2025)
Similar Items
-
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
by: Wang, Haoyu, et al.
Published: (2025) -
Learning to Retrieve from Agent Trajectories
by: Zhou, Yuqi, et al.
Published: (2026) -
Neural Retrievers are Biased Towards LLM-Generated Content
by: Dai, Sunhao, et al.
Published: (2023) -
Exploring the Escalation of Source Bias in User, Data, and Recommender System Feedback Loop
by: Zhou, Yuqi, et al.
Published: (2024) -
Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era
by: Dai, Sunhao, et al.
Published: (2024)