Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Zhuo-Yang, Li, Zeyu, Cao, Qing-Hong, Luo, Ming-xing, Zhu, Hua Xing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explainable AI-assisted Optimization for Feynman Integral Reduction
by: Song, Zhuo-Yang, et al.
Published: (2025)
by: Song, Zhuo-Yang, et al.
Published: (2025)
Detailed balance in large language model-driven agents
by: Song, Zhuo-Yang, et al.
Published: (2025)
by: Song, Zhuo-Yang, et al.
Published: (2025)
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
A Theory of LLM Information Susceptibility
by: Song, Zhuo-Yang, et al.
Published: (2026)
by: Song, Zhuo-Yang, et al.
Published: (2026)
WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
by: Li, Kuan, et al.
Published: (2025)
by: Li, Kuan, et al.
Published: (2025)
Improve Decoding Factuality by Token-wise Cross Layer Entropy of Large Language Models
by: Wu, Jialiang, et al.
Published: (2025)
by: Wu, Jialiang, et al.
Published: (2025)
Dimensionality Reduction in Sentence Transformer Vector Databases with Fast Fourier Transform
by: Bulgakov, Vitaly, et al.
Published: (2024)
by: Bulgakov, Vitaly, et al.
Published: (2024)
Improving Clustering on Occupational Text Data through Dimensionality Reduction
by: García, Iago Xabier Vázquez, et al.
Published: (2025)
by: García, Iago Xabier Vázquez, et al.
Published: (2025)
ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing
by: An, Yongqi, et al.
Published: (2026)
by: An, Yongqi, et al.
Published: (2026)
LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation
by: Yang, Ruixiao, et al.
Published: (2026)
by: Yang, Ruixiao, et al.
Published: (2026)
FTP: A Fine-grained Token-wise Pruner for Large Language Models via Token Routing
by: Li, Zekai, et al.
Published: (2024)
by: Li, Zekai, et al.
Published: (2024)
Navigating Cultural Chasms: Exploring and Unlocking the Cultural POV of Text-To-Image Models
by: Ventura, Mor, et al.
Published: (2023)
by: Ventura, Mor, et al.
Published: (2023)
Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
by: Li, Shiwei, et al.
Published: (2025)
by: Li, Shiwei, et al.
Published: (2025)
An End-to-end Architecture for Collider Physics and Beyond
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
Iterated Agent for Symbolic Regression
by: Song, Zhuo-Yang, et al.
Published: (2025)
by: Song, Zhuo-Yang, et al.
Published: (2025)
Uncertainty Quantification of Large Language Models through Multi-Dimensional Responses
by: Chen, Tiejin, et al.
Published: (2025)
by: Chen, Tiejin, et al.
Published: (2025)
Evaluating Unsupervised Dimensionality Reduction Methods for Pretrained Sentence Embeddings
by: Zhang, Gaifan, et al.
Published: (2024)
by: Zhang, Gaifan, et al.
Published: (2024)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Uncovering Limitations of Large Language Models in Information Seeking from Tables
by: Pang, Chaoxu, et al.
Published: (2024)
by: Pang, Chaoxu, et al.
Published: (2024)
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
by: Bae, Sangmin, et al.
Published: (2024)
by: Bae, Sangmin, et al.
Published: (2024)
Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate
by: Bochkov, A.
Published: (2025)
by: Bochkov, A.
Published: (2025)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
by: Shi, Zhenmei, et al.
Published: (2024)
by: Shi, Zhenmei, et al.
Published: (2024)
Soft Theorem to Three Loops in QCD and ${\cal N} = 4$ Super Yang-Mills Theory
by: Chen, Wen, et al.
Published: (2023)
by: Chen, Wen, et al.
Published: (2023)
Layer-wise Regularized Dropout for Neural Language Models
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Layer-wise Swapping for Generalizable Multilingual Safety
by: Shin, Hyunseo, et al.
Published: (2026)
by: Shin, Hyunseo, et al.
Published: (2026)
Skip-Layer Attention: Bridging Abstract and Detailed Dependencies in Transformers
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
Head-wise Shareable Attention for Large Language Models
by: Cao, Zouying, et al.
Published: (2024)
by: Cao, Zouying, et al.
Published: (2024)
Towards Token-Level Text Anomaly Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
Emergence of a High-Dimensional Abstraction Phase in Language Transformers
by: Cheng, Emily, et al.
Published: (2024)
by: Cheng, Emily, et al.
Published: (2024)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
Empirical Study on Updating Key-Value Memories in Transformer Feed-forward Layers
by: Qiu, Zihan, et al.
Published: (2024)
by: Qiu, Zihan, et al.
Published: (2024)
On the Effect of Uncertainty on Layer-wise Inference Dynamics
by: Kim, Sunwoo, et al.
Published: (2025)
by: Kim, Sunwoo, et al.
Published: (2025)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
by: Song, Guanghui, et al.
Published: (2025)
by: Song, Guanghui, et al.
Published: (2025)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
by: Zhou, Enshuai, et al.
Published: (2026)
by: Zhou, Enshuai, et al.
Published: (2026)
Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization
by: Wang, Dixuan, et al.
Published: (2024)
by: Wang, Dixuan, et al.
Published: (2024)
Layer by Layer: Uncovering Hidden Representations in Language Models
by: Skean, Oscar, et al.
Published: (2025)
by: Skean, Oscar, et al.
Published: (2025)
Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding
by: Zhong, Lin, et al.
Published: (2026)
by: Zhong, Lin, et al.
Published: (2026)
ToDi: Token-wise Distillation via Fine-Grained Divergence Control
by: Jung, Seongryong, et al.
Published: (2025)
by: Jung, Seongryong, et al.
Published: (2025)
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
by: Jiang, Jingzhou, et al.
Published: (2026)
by: Jiang, Jingzhou, et al.
Published: (2026)
Similar Items
-
Explainable AI-assisted Optimization for Feynman Integral Reduction
by: Song, Zhuo-Yang, et al.
Published: (2025) -
Detailed balance in large language model-driven agents
by: Song, Zhuo-Yang, et al.
Published: (2025) -
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
by: Waldis, Andreas, et al.
Published: (2024) -
A Theory of LLM Information Susceptibility
by: Song, Zhuo-Yang, et al.
Published: (2026) -
WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
by: Li, Kuan, et al.
Published: (2025)