Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Xueying, Huang, Xingyue, Ju, Mingxuan, Collins, Liam, Liu, Yozen, Akoglu, Leman, Shah, Neil, Zhao, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
by: Huang, Xingyue, et al.
Published: (2026)
by: Huang, Xingyue, et al.
Published: (2026)
Toward Privileged Foundation Models:LUPI for Accelerated and Improved Learning
by: Ding, Xueying, et al.
Published: (2026)
by: Ding, Xueying, et al.
Published: (2026)
Fast Unsupervised Deep Outlier Model Selection with Hypernetworks
by: Ding, Xueying, et al.
Published: (2023)
by: Ding, Xueying, et al.
Published: (2023)
Pard: Permutation-Invariant Autoregressive Diffusion for Graph Generation
by: Zhao, Lingxiao, et al.
Published: (2024)
by: Zhao, Lingxiao, et al.
Published: (2024)
VIP-COP: Context Optimization for Tabular Foundation Models
by: Chen, Yilong, et al.
Published: (2026)
by: Chen, Yilong, et al.
Published: (2026)
Outlier Detection Bias Busted: Understanding Sources of Algorithmic Bias through Data-centric Factors
by: Ding, Xueying, et al.
Published: (2024)
by: Ding, Xueying, et al.
Published: (2024)
Unified Discrete Diffusion for Categorical Data
by: Zhao, Lingxiao, et al.
Published: (2024)
by: Zhao, Lingxiao, et al.
Published: (2024)
Robust Training Objectives Improve Embedding-based Retrieval in Industrial Recommendation Systems
by: Kolodner, Matthew, et al.
Published: (2024)
by: Kolodner, Matthew, et al.
Published: (2024)
From Zero to Hero: Advancing Zero-Shot Foundation Models for Tabular Outlier Detection
by: Ding, Xueying, et al.
Published: (2026)
by: Ding, Xueying, et al.
Published: (2026)
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
by: Zhu, Jing, et al.
Published: (2025)
by: Zhu, Jing, et al.
Published: (2025)
How Does Message Passing Improve Collaborative Filtering?
by: Ju, Mingxuan, et al.
Published: (2024)
by: Ju, Mingxuan, et al.
Published: (2024)
MacrOData: New Benchmarks of Thousands of Datasets for Tabular Outlier Detection
by: Ding, Xueying, et al.
Published: (2026)
by: Ding, Xueying, et al.
Published: (2026)
Understanding Generative Recommendation with Semantic IDs from a Model-scaling View
by: Liu, Jingzhe, et al.
Published: (2025)
by: Liu, Jingzhe, et al.
Published: (2025)
Improving Out-of-Vocabulary Handling in Recommendation Systems
by: Shiao, William, et al.
Published: (2024)
by: Shiao, William, et al.
Published: (2024)
Self-Tuning Self-Supervised Image Anomaly Detection
by: Yoo, Jaemin, et al.
Published: (2023)
by: Yoo, Jaemin, et al.
Published: (2023)
Exploiting ID-Text Complementarity via Ensembling for Sequential Recommendation
by: Collins, Liam, et al.
Published: (2025)
by: Collins, Liam, et al.
Published: (2025)
Node Duplication Improves Cold-start Link Prediction
by: Guo, Zhichun, et al.
Published: (2024)
by: Guo, Zhichun, et al.
Published: (2024)
Descriptive Kernel Convolution Network with Improved Random Walk Kernel
by: Lee, Meng-Chieh, et al.
Published: (2024)
by: Lee, Meng-Chieh, et al.
Published: (2024)
Generative Recommendation with Semantic IDs: A Practitioner's Handbook
by: Ju, Clark Mingxuan, et al.
Published: (2025)
by: Ju, Clark Mingxuan, et al.
Published: (2025)
On the Detection of Reviewer-Author Collusion Rings From Paper Bidding
by: Jecmen, Steven, et al.
Published: (2024)
by: Jecmen, Steven, et al.
Published: (2024)
CoBAD: Modeling Collective Behaviors for Human Mobility Anomaly Detection
by: Wen, Haomin, et al.
Published: (2025)
by: Wen, Haomin, et al.
Published: (2025)
FoMo-0D: A Foundation Model for Zero-shot Tabular Outlier Detection
by: Shen, Yuchen, et al.
Published: (2024)
by: Shen, Yuchen, et al.
Published: (2024)
Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights
by: Chen, Wenbo, et al.
Published: (2026)
by: Chen, Wenbo, et al.
Published: (2026)
FlexRec: Adapting LLM-based Recommenders for Flexible Needs via Reinforcement Learning
by: Pan, Yijun, et al.
Published: (2026)
by: Pan, Yijun, et al.
Published: (2026)
Token Prepending: A Training-Free Approach for Eliciting Better Sentence Embeddings from LLMs
by: Fu, Yuchen, et al.
Published: (2024)
by: Fu, Yuchen, et al.
Published: (2024)
Enhancing Item Tokenization for Generative Recommendation through Self-Improvement
by: Chen, Runjin, et al.
Published: (2024)
by: Chen, Runjin, et al.
Published: (2024)
Can Machine Learning Target Health Care Fraud? Evidence From Medicare Hospitalizations
by: Shubhranshu Shekhar, et al.
Published: (2025)
by: Shubhranshu Shekhar, et al.
Published: (2025)
A Pre-training Framework for Relational Data with Information-theoretic Principles
by: Truong, Quang, et al.
Published: (2025)
by: Truong, Quang, et al.
Published: (2025)
GraphHash: Graph Clustering Enables Parameter Efficiency in Recommender Systems
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
On the Role of Weight Decay in Collaborative Filtering: A Popularity Perspective
by: Loveland, Donald, et al.
Published: (2025)
by: Loveland, Donald, et al.
Published: (2025)
Masked Diffusion for Generative Recommendation
by: Shah, Kulin, et al.
Published: (2025)
by: Shah, Kulin, et al.
Published: (2025)
Revisiting Self-attention for Cross-domain Sequential Recommendation
by: Ju, Clark Mingxuan, et al.
Published: (2025)
by: Ju, Clark Mingxuan, et al.
Published: (2025)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
by: Zhou, Tianyang, et al.
Published: (2026)
by: Zhou, Tianyang, et al.
Published: (2026)
Heuristic Methods are Good Teachers to Distill MLPs for Graph Link Prediction
by: Qin, Zongyue, et al.
Published: (2025)
by: Qin, Zongyue, et al.
Published: (2025)
Understanding and Scaling Collaborative Filtering Optimization from the Perspective of Matrix Rank
by: Loveland, Donald, et al.
Published: (2024)
by: Loveland, Donald, et al.
Published: (2024)
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
by: Deshpande, Vijeta, et al.
Published: (2026)
by: Deshpande, Vijeta, et al.
Published: (2026)
Uncertainty-aware Human Mobility Modeling and Anomaly Detection
by: Wen, Haomin, et al.
Published: (2024)
by: Wen, Haomin, et al.
Published: (2024)
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
by: Lam, Tsz Kin, et al.
Published: (2025)
by: Lam, Tsz Kin, et al.
Published: (2025)
AnyPro: Preference-Preserving Anycast Optimization based on Strategic AS-Path Prepending
by: Zhou, Minyuan, et al.
Published: (2026)
by: Zhou, Minyuan, et al.
Published: (2026)
CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search
by: Zeng, Hansi, et al.
Published: (2026)
by: Zeng, Hansi, et al.
Published: (2026)
Similar Items
-
Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
by: Huang, Xingyue, et al.
Published: (2026) -
Toward Privileged Foundation Models:LUPI for Accelerated and Improved Learning
by: Ding, Xueying, et al.
Published: (2026) -
Fast Unsupervised Deep Outlier Model Selection with Hypernetworks
by: Ding, Xueying, et al.
Published: (2023) -
Pard: Permutation-Invariant Autoregressive Diffusion for Graph Generation
by: Zhao, Lingxiao, et al.
Published: (2024) -
VIP-COP: Context Optimization for Tabular Foundation Models
by: Chen, Yilong, et al.
Published: (2026)